feat(grpo): enhance action response parsing by removing reasoning blocks and refining regex handling 5202cdc Bemohit commited on Apr 26
feat(grpo): update max sequence length and refine prompt formatting in training scripts 79bced7 Bemohit commited on Apr 25
refactor: remove SakhaEnvWrapper class and streamline reward function in GRPO training script 097c9e4 Bemohit commited on Apr 25
test(rubric): add golden parity tests for rubric migration validation ce083e4 unverified atharva-again commited on Apr 25
data(fixtures): capture pre-migration golden reward fixtures for parity testing 827cfe7 unverified atharva-again commited on Apr 25
test(inference): add prompt profile and rank_candidates tests f38de94 unverified atharva-again commited on Apr 7
fix: heuristic takeover when llm fails, step reward/penalty derived from grader, some other fixes 52b4770 unverified atharva-again commited on Apr 4