nba-viz / eval

Commit History

Question-level checkpointing in the eval (--checkpoint-every, default 10)
65ca914

HenriLD Claude Opus 4.8 commited on

Make the eval harness resume-safe (--out)
5b23e34

HenriLD Claude Opus 4.8 commited on

Log generation/request ids on eval failures for trace debugging
0a88727

HenriLD Claude Opus 4.8 commited on

Add AGENT_MAX_RPM global throttle for free-tier benchmarking
091d09a

HenriLD Claude Opus 4.8 commited on

Experiment: AGENT_SQL_ONLY mode (no templates, SQL for everything)
c9d0871

HenriLD Claude Opus 4.8 commited on

Add opponent/won to v_shots; fix benchmark label
fa27abb

HenriLD Claude Opus 4.8 commited on

Parallelize the eval harness + add per-model health table
8114f20

HenriLD Claude Opus 4.8 commited on

Consolidate all eval questions into one document
2b7e74b

HenriLD Claude Opus 4.8 commited on

Expand end-to-end benchmark to 109 Q/A pairs
e3443d1

HenriLD Claude Opus 4.8 commited on

Advanced analytics layer: make deep/compound questions first-class
8fd83e2

HenriLD Claude Opus 4.8 commited on

eval: widen want_metric keywords to match rendered axis labels
807b566

HenriLD Claude Opus 4.8 commited on

Steer interpretive questions toward insightful metrics
efba11e

HenriLD Claude Opus 4.8 commited on

Shot chart: color by any category (quarter, zone, opponent), not just made/miss
c32c794

HenriLD Claude Opus 4.8 commited on

Expose court renderers as chart types for the flexible SQL path
ecaae24

HenriLD Claude Opus 4.8 commited on

Eval: correct two stale 'decline' expectations
d7660b1

HenriLD Claude Opus 4.8 commited on

Eval: provider/model health tracking, --runs, tougher interpretive questions
625627e

HenriLD Claude Opus 4.8 commited on

Distributions: drop mean line, fix theming, support player/team comparisons
ba3ab08

HenriLD Claude Opus 4.8 commited on

Bias toward distributions: violin/box/histogram + stat_distribution template
855bc7f

HenriLD Claude Opus 4.8 commited on

Render multiple charts side by side + harden against model API errors
7e84d02

HenriLD Claude Fable 5 commited on

Make benchmark run() merge into existing results
4b5505b

HenriLD Claude Fable 5 commited on

Fix benchmark/eval to use auto tool_choice (forced function unsupported)
d36a3cf

HenriLD Claude Fable 5 commited on

Pretty field labels + multi-model benchmark harness
8ba3cab

HenriLD Claude Fable 5 commited on

Add flexible SQL query path + new templates + safe executor
c3b1803

HenriLD Claude Fable 5 commited on

Initial deploy: NL-to-chart NBA visualization app
7592718

HenriLD Claude Fable 5 commited on