nba-viz / eval /run_eval.py

Commit History

Question-level checkpointing in the eval (--checkpoint-every, default 10)
65ca914

HenriLD Claude Opus 4.8 commited on

Make the eval harness resume-safe (--out)
5b23e34

HenriLD Claude Opus 4.8 commited on

Log generation/request ids on eval failures for trace debugging
0a88727

HenriLD Claude Opus 4.8 commited on

Add AGENT_MAX_RPM global throttle for free-tier benchmarking
091d09a

HenriLD Claude Opus 4.8 commited on

Experiment: AGENT_SQL_ONLY mode (no templates, SQL for everything)
c9d0871

HenriLD Claude Opus 4.8 commited on

Parallelize the eval harness + add per-model health table
8114f20

HenriLD Claude Opus 4.8 commited on

Consolidate all eval questions into one document
2b7e74b

HenriLD Claude Opus 4.8 commited on

Steer interpretive questions toward insightful metrics
efba11e

HenriLD Claude Opus 4.8 commited on

Eval: provider/model health tracking, --runs, tougher interpretive questions
625627e

HenriLD Claude Opus 4.8 commited on

Render multiple charts side by side + harden against model API errors
7e84d02

HenriLD Claude Fable 5 commited on

Add flexible SQL query path + new templates + safe executor
c3b1803

HenriLD Claude Fable 5 commited on

Initial deploy: NL-to-chart NBA visualization app
7592718

HenriLD Claude Fable 5 commited on