Request update model to google updates of july 2026

#3
by bagwani - opened

Google made huge improvements to Gemma 4 tool-calling and chat accuracy, reliability + speed

Near complete list of Google's fixes/improvements:

  • Flash Attention 4: Uniform FA4 support on NVIDIA Hopper GPUs.
  • Speed gains: 25–70% higher prefill throughput and up to 31% lower time-to-first-token.
  • Chat formatting: Fixed null handling, input validation and unbalanced/missing turn tags.
  • Reasoning preservation: Corrected how thinking/reasoning content is retained and rendered.
  • Tool responses: Restored the assistant turn and thinking cue after tool outputs.
  • Tool calling: Improved execution accuracy, consistency and tool-call-only turn closure.
  • Continuation turns: Removed duplicate <turn|> tags and unwanted extra newlines.
  • Generation prompts: Reverted an add_generation_prompt regression and restored expected defaults.
  • Conversation history: Removed the obsolete APC thought primer and corrected historical-turn handling.
  • Template standardization: Added the canonical chat-template header and updated stale comments.
  • Vision controls: Default remains max_soft_tokens=280; use 1120 for sharper OCR and up to 2.51MP detail.
  • 31B benchmark gains: BFCL +0.4, TB2 +4.5, Retail +3.1, Airline +2.0 and Telecom +10.1 points.
QuantTrio org

Thanks for the reminder. I’ve updated tokenizer_config.json from the upstream model repository.

JunHowie changed discussion status to closed

Sign up or log in to comment