Instructions to use Wafflebyte/Wiss_2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Wafflebyte/Wiss_2b with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Wafflebyte/Wiss_2b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Wiss 2B
Wiss 2B is a small (2B) assistant fine-tuned from Qwen3.5-2B for coding, multi-step tool use and short, finished reasoning. It keeps the base model's image understanding (the image encoder and projector are unchanged).
It is a narrow fine-tune, not a general upgrade: on standard benchmarks it is behind the base model on instruction following, hard math and code, and ahead on BBH-style logic puzzles. The tables below show every number we measured.
A note from the team: we were stumped by these results. We aimed for a model that was better at coding and tool use without giving up general ability, and on standard benchmarks it fell short of the base model in several places. We are releasing it anyway, openly and with all the numbers, because it is still useful for what it was built for. Our next model, Wiss2, will try our best to fix this: more general-purpose training data, replay of earlier data to stop forgetting, and release checks measured against the base model. Due to some technical issues, we are unable to release the model right now.
Benchmarks
All models were run by us with the same settings: lm-evaluation-harness on vLLM, each model's own chat template, thinking mode off, greedy decoding, bf16. IFEval through MMLU-Pro are the Open LLM Leaderboard v2 tasks (raw scores, not the leaderboard's normalized ones). GSM8K is 5-shot strict match; HumanEval is pass@1 (instruct prompt). IFEval, MATH-hard, GSM8K and HumanEval use the full test sets. BBH, MuSR and MMLU-Pro use only the first 60 questions of every subtask / subject (the same questions for both models), so differences of a few points on those are noise.
| Benchmark | Wiss 2B | Qwen3.5-2B (base) | Difference |
|---|---|---|---|
| IFEval, prompt-level strict | 49.5 | 67.3 | β17.8 |
| IFEval, instruction-level strict | 60.4 | 76.0 | β15.6 |
| IFEval, prompt-level loose | 54.0 | 71.2 | β17.2 |
| IFEval, instruction-level loose | 64.4 | 79.4 | β15.0 |
| MATH-hard (exact match) | 11.9 | 31.2 | β19.3 |
| GSM8K | 61.8 | 63.8 | β2.0 |
| HumanEval (pass@1) | 42.7 | 50.0 | β7.3 |
| BBH (acc_norm) | 48.4 | 39.4 | +9.0 |
| MuSR (acc_norm) | 41.7 | 38.3 | +3.3 |
| MMLU-Pro (acc) | 25.0 | 26.7 | β1.7 (within noise) |
With thinking on (GSM8K, first 500 test questions)
Standard prompt ("reason step by step, and put your final answer within \boxed{}"), greedy decoding, a 2,048-token budget for both models. An answer counts only if the model finished thinking and its final answer is right.
| Model | GSM8K |
|---|---|
| Wiss 2B | 66.8 |
| Qwen3.5-2B (base) | 17.8 |
The base model's low score here is mostly a budget effect: it thinks at such length that most answers are cut off at 2,048 tokens before a final answer. Wiss thinks briefly and finishes inside the budget. This is not evidence that Wiss is a better mathematician than the base model with a larger budget (see the non-thinking table above, where the base model is ahead on math).
Usage
Needs transformers >= 5.
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "Wiss-2B" # path or Hub id of this model
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="auto", device_map="auto")
messages = [{"role": "user", "content": [{"type": "text", "text": "Write a Python function that merges two sorted lists."}]}]
# pictures: add {"type": "image", "image": "photo.jpg"} to the content list
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True, return_dict=True,
return_tensors="pt", enable_thinking=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=1024, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Recommended sampling: thinking on temperature=0.6, top_p=0.95, top_k=20; thinking off
temperature=0.7, top_p=0.8, top_k=20. Tools are passed in the standard OpenAI format through the chat template
(tools=[...]).
For llama.cpp (text and images), use the GGUF build with its image projector:
llama-server -m Wiss-2B-Q5_K_M.gguf --mmproj mmproj-Wiss-2B-f16.gguf --jinja --image-max-tokens 2048
Training
Method: LoRA (r=32, Ξ±=64) on every linear projection of the language model, including the Gated DeltaNet linear-attention layers, merged into the weights. The image encoder and projector are unchanged from Qwen3.5-2B.
Schedule: 2 epochs, peak LR 1e-4 with cosine decay, max length 2,048 tokens, loss only on assistant turns.
Compute: one RTX 4090, about 22 minutes, with Unsloth.
Data: 2,974 verified synthetic conversations (2,362 with reasoning, 697 without), plus 85 held out; multi-step conversations train every assistant turn, giving 4,442 training examples. Every row passed an automatic check (unit tests, exact tool-call matching or answer matching) before it was kept.
Family Conversations Code 800 Reasoning 697 Agent (multi-step tool use) 678 Instruction following 346 Grounded answers 253 Single-call tools 200 Exact and near duplicates were removed, and the set was checked for 13-gram overlap with GSM8K, MATH-500, HumanEval, MBPP and IFEval (0 overlaps found). It was not checked against MMLU-Pro, GPQA, BBH or MuSR.
Not trained on: general-knowledge text, competition-level math, or any replay of the base model's own outputs.
Limitations
- Weaker than the base model at instruction following and hard math (see the benchmark table), and somewhat weaker at code on HumanEval. If you need strict format-following or competition math, use Qwen3.5-2B.
- Tool use and vision were the training goals, but the benchmarks above do not measure them.
- A 2B model: it makes small mistakes, especially in code details. Test what it writes.
- Vision was not fine-tuned. It describes pictures well but can invent details in dense screenshots; give it a high image-token budget or crop to the part that matters.
- English only in training; other languages come from the base model.
- Single runs; the multiple-choice sets are small, so differences of a few points are within noise.
License
Apache 2.0, the license of the base model Qwen3.5-2B. Synthetic training data was generated with gpt-oss-120b (Apache 2.0), Pixel Canary (a model offered on the Vercel AI Gateway) and Ling 3.1 Flash (inclusionAI).