Random-depth SFT and depth-as-action GRPO checkpoints for Ouro-1.4B-Thinking.
Omar Abul-Hassan
omar81939
AI & ML interests
None yet
Recent Activity
upvoted a collection 6 days ago
DIAL: GRPO on looped models updated a collection 7 days ago
DIAL: GRPO on looped models updated a collection 7 days ago
DIAL: GRPO on looped modelsOrganizations
None yet