Text Generation
Transformers
Safetensors
PEFT
gemma-3
continued-pretraining
sft
lora
synthetic-data
alignment
midtraining
scimt
sidbaines commited on
Commit
14fa133
·
verified ·
1 Parent(s): 1bbdc16

add rl_grpo/results (copied from sidbaines/scimt-prior-coins-dispatch-sdf-aft-v1)

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. rl_grpo/results/charter_real_4x_direct-step128/eval_holdout_agreement.jsonl +0 -0
  2. rl_grpo/results/charter_real_4x_direct-step128/eval_holdout_conflict.jsonl +0 -0
  3. rl_grpo/results/charter_real_4x_direct-step128/eval_trained_agreement.jsonl +0 -0
  4. rl_grpo/results/charter_real_4x_direct-step128/eval_trained_conflict.jsonl +0 -0
  5. rl_grpo/results/charter_real_4x_direct-step128/sanity_prompts.jsonl +0 -0
  6. rl_grpo/results/charter_real_4x_direct-step16/eval_holdout_agreement.jsonl +0 -0
  7. rl_grpo/results/charter_real_4x_direct-step16/eval_holdout_conflict.jsonl +0 -0
  8. rl_grpo/results/charter_real_4x_direct-step16/eval_trained_agreement.jsonl +0 -0
  9. rl_grpo/results/charter_real_4x_direct-step16/eval_trained_conflict.jsonl +0 -0
  10. rl_grpo/results/charter_real_4x_direct-step16/sanity_prompts.jsonl +0 -0
  11. rl_grpo/results/charter_real_4x_direct-step256/eval_holdout_agreement.jsonl +0 -0
  12. rl_grpo/results/charter_real_4x_direct-step256/eval_holdout_conflict.jsonl +0 -0
  13. rl_grpo/results/charter_real_4x_direct-step256/eval_trained_agreement.jsonl +0 -0
  14. rl_grpo/results/charter_real_4x_direct-step256/eval_trained_conflict.jsonl +0 -0
  15. rl_grpo/results/charter_real_4x_direct-step256/sanity_prompts.jsonl +0 -0
  16. rl_grpo/results/charter_real_4x_direct-step32/eval_holdout_agreement.jsonl +0 -0
  17. rl_grpo/results/charter_real_4x_direct-step32/eval_holdout_conflict.jsonl +0 -0
  18. rl_grpo/results/charter_real_4x_direct-step32/eval_trained_agreement.jsonl +0 -0
  19. rl_grpo/results/charter_real_4x_direct-step32/eval_trained_conflict.jsonl +0 -0
  20. rl_grpo/results/charter_real_4x_direct-step32/sanity_prompts.jsonl +0 -0
  21. rl_grpo/results/charter_real_4x_direct-step64/eval_holdout_agreement.jsonl +0 -0
  22. rl_grpo/results/charter_real_4x_direct-step64/eval_holdout_conflict.jsonl +0 -0
  23. rl_grpo/results/charter_real_4x_direct-step64/eval_trained_agreement.jsonl +0 -0
  24. rl_grpo/results/charter_real_4x_direct-step64/eval_trained_conflict.jsonl +0 -0
  25. rl_grpo/results/charter_real_4x_direct-step64/sanity_prompts.jsonl +0 -0
  26. rl_grpo/results/charter_real_4x_direct/CELL_DONE.json +1 -0
  27. rl_grpo/results/charter_real_4x_thinking-step128/eval_holdout_agreement.jsonl +0 -0
  28. rl_grpo/results/charter_real_4x_thinking-step128/eval_holdout_conflict.jsonl +0 -0
  29. rl_grpo/results/charter_real_4x_thinking-step128/eval_trained_agreement.jsonl +0 -0
  30. rl_grpo/results/charter_real_4x_thinking-step128/eval_trained_conflict.jsonl +0 -0
  31. rl_grpo/results/charter_real_4x_thinking-step16/eval_holdout_agreement.jsonl +0 -0
  32. rl_grpo/results/charter_real_4x_thinking-step16/eval_holdout_conflict.jsonl +0 -0
  33. rl_grpo/results/charter_real_4x_thinking-step16/eval_trained_agreement.jsonl +0 -0
  34. rl_grpo/results/charter_real_4x_thinking-step16/eval_trained_conflict.jsonl +0 -0
  35. rl_grpo/results/charter_real_4x_thinking-step256/eval_holdout_agreement.jsonl +0 -0
  36. rl_grpo/results/charter_real_4x_thinking-step256/eval_holdout_conflict.jsonl +0 -0
  37. rl_grpo/results/charter_real_4x_thinking-step256/eval_trained_agreement.jsonl +0 -0
  38. rl_grpo/results/charter_real_4x_thinking-step256/eval_trained_conflict.jsonl +0 -0
  39. rl_grpo/results/charter_real_4x_thinking-step32/eval_holdout_agreement.jsonl +0 -0
  40. rl_grpo/results/charter_real_4x_thinking-step32/eval_holdout_conflict.jsonl +0 -0
  41. rl_grpo/results/charter_real_4x_thinking-step32/eval_trained_agreement.jsonl +0 -0
  42. rl_grpo/results/charter_real_4x_thinking-step32/eval_trained_conflict.jsonl +0 -0
  43. rl_grpo/results/charter_real_4x_thinking-step64/eval_holdout_agreement.jsonl +0 -0
  44. rl_grpo/results/charter_real_4x_thinking-step64/eval_holdout_conflict.jsonl +0 -0
  45. rl_grpo/results/charter_real_4x_thinking-step64/eval_trained_agreement.jsonl +0 -0
  46. rl_grpo/results/charter_real_4x_thinking-step64/eval_trained_conflict.jsonl +0 -0
  47. rl_grpo/results/charter_real_4x_thinking/CELL_DONE.json +1 -0
  48. rl_grpo/results/charter_real_4x_thinking/sanity_prompts.jsonl +0 -0
  49. rl_grpo/results/charter_real_4x_thinking__base/eval_holdout_agreement.jsonl +0 -0
  50. rl_grpo/results/charter_real_4x_thinking__base/eval_holdout_conflict.jsonl +0 -0
rl_grpo/results/charter_real_4x_direct-step128/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step128/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step128/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step128/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step128/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step16/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step16/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step16/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step16/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step16/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step256/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step256/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step256/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step256/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step256/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step32/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step32/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step32/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step32/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step32/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step64/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step64/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step64/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step64/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct-step64/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_direct/CELL_DONE.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"label": "charter_real_4x_direct", "mode": "direct", "endpoints": [16, 32, 64, 128, 256], "at": "2026-08-11T14:26:41Z"}
rl_grpo/results/charter_real_4x_thinking-step128/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step128/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step128/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step128/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step16/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step16/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step16/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step16/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step256/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step256/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step256/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step256/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step32/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step32/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step32/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step32/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step64/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step64/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step64/eval_trained_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking-step64/eval_trained_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking/CELL_DONE.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"label": "charter_real_4x_thinking", "mode": "thinking", "endpoints": [16, 32, 64, 128, 256], "at": "2026-08-11T20:50:06Z"}
rl_grpo/results/charter_real_4x_thinking/sanity_prompts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking__base/eval_holdout_agreement.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
rl_grpo/results/charter_real_4x_thinking__base/eval_holdout_conflict.jsonl ADDED
The diff for this file is too large to render. See raw diff