Spaces:
Paused
Evaluation
make eval is the only source of the retrieval/answer-quality numbers that
appear anywhere in this repo's docs β nothing is hand-typed (see
CONTRIBUTING.md: "Never hand-write benchmark numbers").
Running it
make eval # full run, writes reports/eval_report.json
auralynq eval --smoke # tiny 3-item run (what CI runs)
auralynq eval --report # same as `make eval`
Works fully offline at $0: with no embedding/LLM extras installed, it uses the hash embedder and extractive answerer (ADR-0003) β the numbers will be lower quality than a real provider, but the harness itself, the report schema, and the drift check all still run and pass.
What it measures
For a frozen golden question set (auralynq/eval/datasets.py, --smoke
uses the first 3 items):
- Retrieval comparison across
naive,hybrid, andpathragretrievers:recall_at_k,precision_at_k,ndcg_at_10,mrr,latency_p50_ms. - Agentic (full pipeline): the same retrieval metrics plus Ragas (or a
proxy) faithfulness/answer-relevancy/context-precision scores, run through
the actual
auralynq_ragagent β not a bare retriever. - ASR: word error rate against
data/golden/asr_refs.json, when audio golden items exist. - Drift check: compares
agentic_recall,agentic_faithfulness, andhybrid_ndcgagainstreports/golden_baseline.json. The first run on a machine creates the baseline; subsequent runs flag"status": "regressed"if any of those three metrics drops by more thanDRIFT_TOLERANCE(0.05) versus the stored baseline. This is a local regression trip-wire, not a cross-machine comparison β a baseline created on one machine/config isn't meaningful compared against a run on a different one.
Report shape
Every report (eval_report.json and every bench-* report β see
docs/benchmarks.md) includes a provenance block
(auralynq/eval/provenance.py):
{
"provenance": {
"git_commit": "β¦",
"generated_at": "2026-07-01T20:38:42+00:00",
"hardware": { "cpu": {...}, "ram_gb": ..., "gpus": [...], "best_backend": "..." },
"dataset_version": "golden_qa.json n=5 (smoke=False)"
}
}
hardware reuses the ModelFit hardware probe (auralynq/modelfit/hardware.py)
rather than duplicating detection logic β the same numbers you'd see from
auralynq-modelfit hardware.
Where reports live
reports/ is git-ignored except .gitkeep/README.md β every report is
regenerated locally, never committed. If you want to cite a number in a PR
or doc, regenerate it and quote the provenance block alongside it so
reviewers can reproduce it.
CI
.github/workflows/ci.yml runs python -m auralynq.cli eval --smoke on
every push β this is a correctness gate (the harness runs end-to-end with no
keys or extras installed), not a quality gate. A green CI run says nothing
about answer quality; it says the pipeline didn't break.