| # SVARASETU performance |
|
|
| These are measured results, not estimated or hard-coded values. |
|
|
| **Run:** 2026-08-22 on Windows 11, Python 3.13.3, 16 logical CPUs, 15.63 GB RAM |
|
|
| **Indexes:** native 367,832 · sentence-window 44,489 · semantic 8,828 · total 421,149 |
|
|
| **Window:** transcript to extractive answer. Sarvam STT and optional Gemini are outside the Task 2 clock. |
|
|
| ## Task 2 SLA (`py -3.13 -m svarasetu.measure.sla 20`) |
|
|
| | Layer | n | P50 | P70 | P99 | P100 | Budget | Pass | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | --- | |
| | Exact cache lookup | 20 | 0.043 | 0.046 | 0.056 | 0.057 | 0.1 ms | yes | |
| | FAISS HNSW only | 20 | 0.439 | 0.567 | 0.951 | 0.955 | 0.1 ms P50 stretch | no | |
| | Embed + FAISS | 20 | 7.68 | 8.06 | 9.44 | 9.46 | 50 ms | yes | |
| | Full RAG cache-on | 20 | 0.243 | 0.283 | 0.350 | 0.352 | 200 ms | yes | |
| | **Full RAG cache-off** | 20 | **11.235** | **12.022** | **12.855** | **12.860** | 200 ms | **yes** | |
|
|
| The full path includes dynamic IDF lexical rescue over the loaded corpus. This prevents high dense scores from overriding missing subject evidence. Unsupported questions decline rather than receiving a static answer. |
|
|
| Guardrails: unsafe 2/2 blocked; off-topic/intent 2/2 blocked. |
|
|
| ## Retrieval and answer quality (`py -3.13 -m svarasetu.measure.quality_eval`) |
| |
| | Metric | Value | |
| | --- | ---: | |
| | Recall@1 | 0.4033 | |
| | Recall@5 | 0.9167 | |
| | Recall@10 | 0.9867 | |
| | Precision@5 | 0.1833 | |
| | MRR | 0.6096 | |
| |
| The strict multilingual answer gate passed 6/6: correctness 1.0, groundedness 1.0, refusal correctness 1.0, and hallucination rate 0.0. It covers grounded English/Hindi/Marathi answers plus unsupported, unsafe, and creative-query refusals. The evaluator exits nonzero if quality gates fail. |
| |
| ## Chunking sample |
| |
| | Strategy | Mean chunks | Hashed chunks | |
| | --- | ---: | ---: | |
| | passage-native | 1.00 | 12 | |
| | sentence-window | 5.17 | 62 | |
| | semantic | 1.33 | 16 | |
| | sliding | 1.08 | 13 | |
| | adaptive | 5.00 | 60 | |
| |
| ## Live voice |
| |
| `py -3.13 -m svarasetu.measure.live_voice` passed 15/15 checks with real Sarvam calls on 2026-08-21: |
|
|
| - TTS: English, Hindi, Marathi |
| - STT → RAG: English and Hindi |
| - RAG-core during live API checks: 59–135 ms |
| - observed Sarvam STT: about 5.7–8.3 seconds (reported separately) |
| - cached TTS: 8.96 ms |
|
|
| The deployed 2026-08-22 smoke test also passed a complete Bulbul v3 TTS → Saaras v3 STT → grounded RAG round trip. Browser voice mode uses voice-activity detection to submit after natural silence, pauses its microphone during playback, and resumes listening automatically. |
|
|
| Current machine-readable artifacts are in `svarasetu/measure/reports/svarasetu_sla.json` and `svarasetu_quality_eval.json`. |
|
|