feat: refuse low-information single-token extractive spans (confidently-wrong QA pattern) ae6a269 sathvik commited on about 3 hours ago
fix: raise fast-tier budget gate to 140ms so extractive answers always finish inside the 200ms cap 8e47e87 sathvik commited on about 3 hours ago
feat: strict 200ms latency mode - deadline on by default, pre-fast-tier budget gate, post-STT deadline scope for voice 78ae586 sathvik commited on about 3 hours ago
fix: nest chain budget inside outer generation deadline to bound recovery double-pass 9ba0946 sathvik commited on about 8 hours ago
fix: script-aware grounding - cross-lingual answers grounded on number conservation; Indic-aware tokenization f048f66 sathvik commited on about 9 hours ago
fix: retry truncated Sarvam JSON within attempt budget and raise its token ceiling 1da3c1a sathvik commited on about 9 hours ago
fix: tolerate proxy-injected request ids; document generation chain and measured latency 0b0830a sathvik commited on about 9 hours ago
feat: sarvam-first generation chain (groq throttles under sustained load) 4d7aa62 sathvik commited on about 9 hours ago
fix: groq connection pool starvation - multiple pool slots and bounded attempts cbdd8cd sathvik commited on about 9 hours ago
feat: chain-wide generation budget so total latency stays bounded across tiers 24e1af1 sathvik commited on about 9 hours ago
feat: borderline evidence goes to grounded generative review; bounded API client timeouts 88096c2 sathvik commited on about 10 hours ago
feat: route English queries straight to generative chain (corpus is Indic-only) 4c91f5c sathvik commited on about 10 hours ago
fix: self-healing Groq payload on HTTP 400 and detailed chain attempt diagnostics 44c9dc7 sathvik commited on about 10 hours ago
feat: ordered generation chain (groq -> sarvam -> local GGUF) with per-tier deadlines and bounded llama.cpp threads a730df3 sathvik commited on about 10 hours ago
fix: compile GBNF answer grammar to LlamaGrammar for llama-cpp-python 0.3.35 98f1ee0 sathvik commited on about 10 hours ago
fix: resilient generation - ZeroGPU circuit breaker, resident CPU model, API fallback net ce9639d sathvik commited on about 10 hours ago
feat: strict-latency deadline guard, grammar-constrained GGUF decoding, hardened guardrails (#28) 2d46a89 Sathvik0101 sankalphs commited on about 12 hours ago
fix: hard 30s generation timeout and slow-first-attempt recovery guard (#27) fd95fc0 Sathvik0101 sankalphs commited on about 12 hours ago
fix: drop obsolete ambient reset rule for fixed background (#26) 6f9dc01 Sathvik0101 sankalphs commited on about 13 hours ago
fix: generate multilingual answers instead of extracts (#23) ecf287b Sathvik0101 sankalphs commited on about 22 hours ago
fix: export Qwen model for Space startup (#21) 7bb8bb0 Sathvik0101 sankalphs commited on about 23 hours ago
Fix multilingual retrieval and Qwen fallback (#20) f090a67 Sathvik0101 sankalphs commited on about 23 hours ago
Rebase retrieval latency optimization onto latest Space (#18) d752bbf Sathvik0101 sankalphs commited on about 23 hours ago
Fix selected output language end to end (#17) 7647874 Sathvik0101 sankalphs commited on about 24 hours ago
Make the UI visibly clearer and more useful (#13) 5d8249b Sathvik0101 sankalphs commited on about 24 hours ago
Improve answer UX and provenance transparency (#11) 6d29100 Sathvik0101 sankalphs commited on 1 day ago
Remove slow GGUF JSON grammar and shrink multilingual prompt (#8) 2b7af1b Sathvik0101 sankalphs commited on 1 day ago
Use Blackwell cu130 Q4 runtime and short multilingual generation path (#7) 86f73d2 Sathvik0101 sankalphs commited on 1 day ago
Fallback to CPU llama.cpp when ZeroGPU CUDA GGUF kernels fail (#6) d3b7203 Sathvik0101 sankalphs commited on 1 day ago
Fix fixed background circles and stable input toggle (#2) 1d89aa4 Sathvik0101 sankalphs commited on 1 day ago