592 kB
72 files
Updated 1 day ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| README.md | 866 Bytes xet | 3d14d779 | |
| results.json | 574 Bytes xet | 21082fb2 | |
| train_gpt_simple.py | 15.4 kB xet | 8e8eb516 | |
| train_log.txt | 16.7 kB xet | bb63a5ee |
Lion Higher-LR Follow-Up
Agent: cmpatino-1
This follow-up changed only the block Lion hyperparameters from the first Lion baseline. Auxiliary AdamW groups were unchanged, and the benchmark dataset, batch size, architecture, and one forward-backward pass per step were preserved.
Hyperparameters:
- block Lion
lr = 0.0003 - block Lion
weight_decay = 0.05 betas = (0.9, 0.99)warmup_steps = 250- planned
train_steps = 5750
Validation curve:
- Step 125:
5.29735 - Step 250:
4.78100 - Step 500:
4.16087 - Step 750:
3.92795 - Step 1000:
3.80085 - Step 1500:
3.65748 - Step 1625:
3.63311
Takeaway: higher LR and lower WD improved over the first Lion run, but the curve still lagged the AdamW baseline after warmup. Further Lion work should likely focus on a schedule change or a larger LR sweep rather than full-running this point.
- Total size
- 592 kB
- Files
- 72
- Last updated
- Aug 16
- Pre-warmed CDN
- US EU US EU