Pm-ops / training /rewards.py

Commit History

fix(grpo): fresh reward design — runbook-compliance scoring
1c6aad2

SavK1 Claude Sonnet 4.6 commited on

fix(training): reward signal flow + new v2 notebook
5e6d53a

SavK1 Claude Sonnet 4.6 commited on

modifying the training script to be more memory efficient: Unsloth, also reduced the number of steps to 15 instead of 40
860d7e4

SavK1 commited on

writing the GRPO script to train qwen1.7 on A100 GRPO
5bce7ac

SavK1 commited on