Update model references: Llama 3.3 70B β Llama 3.1 8B Instant, runtime 19.8min β ~10min 1e308ec RohitChandramouli6618 commited on Apr 10
Update greedy baseline numbers from latest test_local run; recalculate lifts a251104 RohitChandramouli6618 commited on Apr 10
Fix duplicate WebSocket log output; update benchmark numbers to latest run (avg 75.9%, ~10min) bfbbcd1 RohitChandramouli6618 commited on Apr 10
Reduce rollouts easy=2 medium=3 hard=3 β target ~15min runtime with buffer for judge overhead 219555f RohitChandramouli6618 commited on Apr 9
Fix 8 issues spotted in code review: Dockerfile path, task descriptions, duplicate function, fallback score consistency, twin constants comment, has_data_lag docs 74f461a RohitChandramouli6618 commited on Apr 8
Update easy rollouts 3β2 everywhere: app.py scores, README 026ea26 RohitChandramouli6618 commited on Apr 8
Update all dashboard tabs and README with actual test run numbers 6e28a1b RohitChandramouli6618 commited on Apr 8
Fix restriction auto-lift, density-weighted penalty, /info endpoint accuracy, updated README 3d5cf7f RohitChandramouli6618 commited on Apr 7
Fix inference.py: mandatory [START][STEP][END] logs, runtime < 20min, fix README hard steps 666127d RohitChandramouli6618 commited on Apr 6