Remove max_new_tokens and add default tool_call_parser, reasoning_parser, enable_thinking values c226526 verified tapu1125 commited on 1 day ago
restore reasoning: enable_thinking default true + preserve_thinking fc72eb9 verified joerowell commited on 9 days ago
spinquantless FP8 (no SpinQuant rotation), 256K context - fixes agentic looping 17cacdc verified joerowell commited on 9 days ago
Enable thinking by default, preserve reasoning; drop max_new_tokens cap 610e625 verified joerowell commited on 10 days ago
Mark </assistant> (token 24) as special in tokenizer.json 20594d0 verified joerowell commited on 10 days ago
Mark </assistant> (token 24) as special to match internal serving aaf93c7 verified joerowell commited on 10 days ago
Fix chat template: preserve reasoning across turns (llama.cpp --reasoning-preserve) 2bea609 verified joerowell commited on 10 days ago