alto-llm-corrector / scripts /benchmark.py

Commit History

bench: ECE/Brier calibration harness + blocking ceilings on the real corpus
befdc2b
unverified

Claude Claude Fable 5 commited on

Versioned benchmark + ground-truth corpus seed (P4.2)
baf1728
unverified

Claude Claude Fable 5 commited on