Wiss 2B

Wiss 2B is a small (2B) assistant fine-tuned from Qwen3.5-2B for coding, multi-step tool use and short, finished reasoning. It keeps the base model's image understanding (the image encoder and projector are unchanged).

It is a narrow fine-tune, not a general upgrade: on standard benchmarks it is behind the base model on instruction following, hard math and code, and ahead on BBH-style logic puzzles. The tables below show every number we measured.

A note from the team: we were stumped by these results. We aimed for a model that was better at coding and tool use without giving up general ability, and on standard benchmarks it fell short of the base model in several places. We are releasing it anyway, openly and with all the numbers, because it is still useful for what it was built for. Our next model, Wiss2, will try our best to fix this: more general-purpose training data, replay of earlier data to stop forgetting, and release checks measured against the base model. Due to some technical issues, we are unable to release the model right now.

Benchmarks

All models were run by us with the same settings: lm-evaluation-harness on vLLM, each model's own chat template, thinking mode off, greedy decoding, bf16. IFEval through MMLU-Pro are the Open LLM Leaderboard v2 tasks (raw scores, not the leaderboard's normalized ones). GSM8K is 5-shot strict match; HumanEval is pass@1 (instruct prompt). IFEval, MATH-hard, GSM8K and HumanEval use the full test sets. BBH, MuSR and MMLU-Pro use only the first 60 questions of every subtask / subject (the same questions for both models), so differences of a few points on those are noise.

Benchmark Wiss 2B Qwen3.5-2B (base) Difference
IFEval, prompt-level strict 49.5 67.3 βˆ’17.8
IFEval, instruction-level strict 60.4 76.0 βˆ’15.6
IFEval, prompt-level loose 54.0 71.2 βˆ’17.2
IFEval, instruction-level loose 64.4 79.4 βˆ’15.0
MATH-hard (exact match) 11.9 31.2 βˆ’19.3
GSM8K 61.8 63.8 βˆ’2.0
HumanEval (pass@1) 42.7 50.0 βˆ’7.3
BBH (acc_norm) 48.4 39.4 +9.0
MuSR (acc_norm) 41.7 38.3 +3.3
MMLU-Pro (acc) 25.0 26.7 βˆ’1.7 (within noise)

With thinking on (GSM8K, first 500 test questions)

Standard prompt ("reason step by step, and put your final answer within \boxed{}"), greedy decoding, a 2,048-token budget for both models. An answer counts only if the model finished thinking and its final answer is right.

Model GSM8K
Wiss 2B 66.8
Qwen3.5-2B (base) 17.8

The base model's low score here is mostly a budget effect: it thinks at such length that most answers are cut off at 2,048 tokens before a final answer. Wiss thinks briefly and finishes inside the budget. This is not evidence that Wiss is a better mathematician than the base model with a larger budget (see the non-thinking table above, where the base model is ahead on math).

Usage

Needs transformers >= 5.

from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "Wiss-2B"  # path or Hub id of this model
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="auto", device_map="auto")

messages = [{"role": "user", "content": [{"type": "text", "text": "Write a Python function that merges two sorted lists."}]}]
# pictures: add {"type": "image", "image": "photo.jpg"} to the content list
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=True, return_dict=True,
                                       return_tensors="pt", enable_thinking=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=1024, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Recommended sampling: thinking on temperature=0.6, top_p=0.95, top_k=20; thinking off temperature=0.7, top_p=0.8, top_k=20. Tools are passed in the standard OpenAI format through the chat template (tools=[...]).

For llama.cpp (text and images), use the GGUF build with its image projector:

llama-server -m Wiss-2B-Q5_K_M.gguf --mmproj mmproj-Wiss-2B-f16.gguf --jinja --image-max-tokens 2048

Training

  • Method: LoRA (r=32, Ξ±=64) on every linear projection of the language model, including the Gated DeltaNet linear-attention layers, merged into the weights. The image encoder and projector are unchanged from Qwen3.5-2B.

  • Schedule: 2 epochs, peak LR 1e-4 with cosine decay, max length 2,048 tokens, loss only on assistant turns.

  • Compute: one RTX 4090, about 22 minutes, with Unsloth.

  • Data: 2,974 verified synthetic conversations (2,362 with reasoning, 697 without), plus 85 held out; multi-step conversations train every assistant turn, giving 4,442 training examples. Every row passed an automatic check (unit tests, exact tool-call matching or answer matching) before it was kept.

    Family Conversations
    Code 800
    Reasoning 697
    Agent (multi-step tool use) 678
    Instruction following 346
    Grounded answers 253
    Single-call tools 200

    Exact and near duplicates were removed, and the set was checked for 13-gram overlap with GSM8K, MATH-500, HumanEval, MBPP and IFEval (0 overlaps found). It was not checked against MMLU-Pro, GPQA, BBH or MuSR.

  • Not trained on: general-knowledge text, competition-level math, or any replay of the base model's own outputs.

Limitations

  • Weaker than the base model at instruction following and hard math (see the benchmark table), and somewhat weaker at code on HumanEval. If you need strict format-following or competition math, use Qwen3.5-2B.
  • Tool use and vision were the training goals, but the benchmarks above do not measure them.
  • A 2B model: it makes small mistakes, especially in code details. Test what it writes.
  • Vision was not fine-tuned. It describes pictures well but can invent details in dense screenshots; give it a high image-token budget or crop to the part that matters.
  • English only in training; other languages come from the base model.
  • Single runs; the multiple-choice sets are small, so differences of a few points are within noise.

License

Apache 2.0, the license of the base model Qwen3.5-2B. Synthetic training data was generated with gpt-oss-120b (Apache 2.0), Pixel Canary (a model offered on the Vercel AI Gateway) and Ling 3.1 Flash (inclusionAI).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Wafflebyte/Wiss_2b

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(453)
this model