A/B: raise max_tokens 256->512 so POCKET's reasoning + final answer complete (not cut) 336526e verified SeaWolf-AI commited on 7 days ago
A/B: show reasoning again (remove no_think prefill); raise A/B max_tokens 96->256 so thinking completes 8bfa67a verified SeaWolf-AI commited on 7 days ago
A/B: one-sentence example prompts (stable tok/s, still fast) 7bf27a9 verified SeaWolf-AI commited on 7 days ago
A/B: shorter-answer example prompts + cap max_tokens 200->96 for snappy races 87a05d4 verified SeaWolf-AI commited on 7 days ago
A/B is now the landing tab (swap order); add prism-ml/Bonsai-27B-gguf to models relation bb234a8 verified SeaWolf-AI commited on 7 days ago
add Tab2 live A/B vs Bonsai (sequential: Bonsai first, then POCKET); 2nd llama-server for Bonsai Q1_0; README=BONSAI vs POCKET e77d0bf verified SeaWolf-AI commited on 7 days ago
show live tok/s: backend streams predicted_per_second; UI displays live speed + final tok/s under each reply 4df8486 verified SeaWolf-AI commited on 7 days ago