Expose data_lag_days in CityObservation; remove fragile num_districts==6 inference b43f9e6 RohitChandramouli6618 commited on Apr 9
Reduce rollouts easy=2 medium=3 hard=3 β target ~15min runtime with buffer for judge overhead 219555f RohitChandramouli6618 commited on Apr 9
Fix 8 issues spotted in code review: Dockerfile path, task descriptions, duplicate function, fallback score consistency, twin constants comment, has_data_lag docs 74f461a RohitChandramouli6618 commited on Apr 8
Cut easy rollouts 3β2 to stay under 20min; update GRPO scores; fill config/tasks.yaml 3e4a196 RohitChandramouli6618 commited on Apr 8
Easy rollouts 2β3 for better score stability, still under 20min a263356 RohitChandramouli6618 commited on Apr 7
Per-task rollouts: easy=2, medium=4, hard=4 β reduces runtime to ~16min 1422c62 RohitChandramouli6618 commited on Apr 7
Reduce max_tokens 120β60 β agent uses ~32 tokens, no content affected ddcd303 RohitChandramouli6618 commited on Apr 7
Fix memory threshold, pre-compute lag estimates in hard prompt c454091 RohitChandramouli6618 commited on Apr 6
Fix memory threshold, pre-compute lag estimates in hard prompt 7f110fa RohitChandramouli6618 commited on Apr 6
Fix evaluator:Increase number of rollouts to 8 for better performance. c568c32 RohitChandramouli6618 commited on Apr 6
Improve agent: fix strategy rule, chain-of-thought, 8 rollouts, better memory f27fd6a RohitChandramouli6618 commited on Apr 5
Fix treatment/recovery calibration, hospital breach reward, efficiency grader, recalibrate seeds 4d57df8 RohitChandramouli6618 commited on Apr 5
Fix realism: natural recovery in spread, hospital breach at 10%, linear spillover, correct prompt fe22c22 RohitChandramouli6618 commited on Apr 5
Increase rollouts to 5 for better GRPO learning convergence 7060479 RohitChandramouli6618 commited on Apr 4
Swap grader weights: hospital=45% primary constraint, containment=30%, calibrate environment 1f18d1f RohitChandramouli6618 commited on Apr 4
Lower spread rate to 0.09, fix medium seeds, teach focused allocation strategy d3b81d6 RohitChandramouli6618 commited on Apr 4
Raise treatment reduction, lower medium/hard seeds, focus prompt strategy 0954bcf RohitChandramouli6618 commited on Apr 4
Fix resource replenishment, grader efficiency, seed calibration, prompt strategy ec747a2 RohitChandramouli6618 commited on Apr 4
Add /grade endpoint, use real grader scores in evaluator 9b98195 RohitChandramouli6618 commited on Apr 2