Ferrell Synthetic Intelligence
Record live model pack benchmark results
8e0c3e6
|
Raw
History Blame Contribute Delete
499 Bytes

AIDE Model Benchmarks

Run the common suite with npm run benchmarks. Run one installed model at a time by setting AIDE_BENCH_MODEL; each model must be served on the endpoint recorded in models/manifest.json.

The first live scorecard is live-results-2026-08-10.json. It is intentionally small and honest: all three public packs passed function and plan tasks, and all three failed the strict unified-diff task. That failure keeps raw model output behind the patch parser and Veritas gate.