Commit History

evals: tunable reasoning budget, so reasoning models score fairly
4546097

dakshtaneja Claude Opus 5 commited on

evals: benchmark candidate models for the t=0 web-search gate
a7b9c56

dakshtaneja Claude Opus 5 commited on

eval harness + env-swappable frontier model
6d1ccea

dakshtaneja commited on