Spaces:
Paused
Benchmarks
Every number under this page's targets is produced by running the command β
never hand-written, never copied from a prior run without regenerating it.
See docs/evaluation.md for make eval (retrieval/answer-quality); this
page covers the rest.
Commands
| Command | What it measures | Requires | Report |
|---|---|---|---|
make bench |
Vector-index recall/latency/memory across quantization modes (none/scalar/binary) | Nothing β runs on whatever's indexed, in-process | reports/bench_report.json |
make bench-modelfit |
ModelFit Index rankings for this machine's hardware | Nothing β pure hardware probe + scoring formula, no network | reports/modelfit_bench_report.json |
make bench-visual-grounding |
Grounding resolution rates (span/segment/page/unavailable) over the golden set |
Nothing β offline-safe | reports/visual_grounding_report.json |
make bench-rag |
RAG-quality benchmark: groundedness, citation coverage, abstention accuracy | A local Ollama server with the target model pulled β degrades to a report noting "not reachable" if absent, never fabricates numbers | reports/rag_bench_report.json |
make export-paper-tables |
Renders every report above into one Markdown file | Whichever reports already exist; missing ones are listed, not backfilled | reports/paper_tables.md |
make bench
make bench-modelfit # or: TASK=coding make bench-modelfit
make bench-visual-grounding
make bench-rag # or: MODEL=ollama:llama3.1:8b make bench-rag
make export-paper-tables
Estimated vs. measured β always separated
Every ModelFit number carries an explicit is_estimate/estimate_used flag
(auralynq/modelfit/scoring.py, resource_estimator.py,
benchmark_runner.py) β a formula-based VRAM/speed estimate is never
presented the same way as a number that came from an actual timed run. The
same discipline extends to the RAG-quality benchmark
(RAGBenchMetrics.is_measured) and community-submitted results
(auralynq/modelfit/community.py, which additionally rejects implausible
submissions β tok_per_sec > 10000 β and strips PII-shaped hardware fields
before accepting a result; see docs/modelfit/community-contributions.md).
Report provenance
Every report β eval_report.json and all five above β embeds the same
provenance block: git commit, UTC timestamp, a full hardware summary, and
a short description of exactly which dataset produced it. See
docs/evaluation.md's "Report shape" section for the exact fields. This is
what make export-paper-tables surfaces under each table so a reader can
tell when, on what hardware, and against which commit a number was
produced β not just what the number was.
The published README numbers
The retrieval-comparison and quantization tables in the main README.md's
Benchmarks section came from real make eval / make bench runs β if you
regenerate them on your own machine, expect different absolute numbers
(different hardware, different corpus state) but the same report shape and
the same provenance fields. If a number in a doc doesn't have a
provenance block backing it somewhere, don't trust it β flag it as a docs
bug.
Offline / CI-safe subset
make bench, make bench-modelfit, and make bench-visual-grounding all
run with zero external services and zero paid keys β they're safe to run in
CI or a Hugging Face Space build. make bench-rag is the one exception
(needs a real local Ollama + model) and is not part of any CI gate; run it
manually when you want real RAG-quality numbers for a specific model.