Inference Providers
Active filters: GRPO
Text Generation
• 0.1B • Updated • 4
Sarahpa/spGRPO-135M-readability
Text Generation
• 0.1B • Updated • 6
alfredcs/gemma-3-27b-grpo-med-merged
Image-Text-to-Text
• Updated • 8
alfredcs/gemma-3-27b-firstaid-icd10-merged
Image-Text-to-Text
• Updated • 3
mradermacher/gemma-3-27b-firstaid-icd10-merged-GGUF
28B • Updated • 26
jinlovespho/SmolGRPO-135M
Text Generation
• 0.1B • Updated • 6
Sarahpa/spGRPO-135M-readability-2
Text Generation
• 0.1B • Updated • 10
Text Generation
• 0.1B • Updated • 9
tariktuna/Summarizer-Demo-SmolGRPO-135M
Text Generation
• 0.1B • Updated • 6
Text Generation
• 0.1B • Updated • 5
supermodelresearch/VAR-d16-GRPO-Aesthetic
Text-to-Image
• Updated supermodelresearch/VAR-d30-GRPO-Aesthetic
Text-to-Image
• Updated dzungever/SmolLM-135M-Instruct-GRPO
Text Generation
• 0.1B • Updated • 6
ritwik098/SmolGRPO-360M-Ritwik
Text Generation
• 0.4B • Updated • 9
Text Generation
• 0.1B • Updated • 6
alfredcs/torchrun-medgemma-27b-grpo-merged
Image-Text-to-Text
• 27B • Updated • 5
KhushalM/Qwen2.5-1.5B-GRPO-Complete
Text Generation
• 2B • Updated • 6
Text Generation
• 0.1B • Updated • 8
Mhammad2023/SmolGRPO-135M
Text Generation
• 0.1B • Updated • 8
Text Generation
• 0.1B • Updated • 7
Text Generation
• 0.1B • Updated • 5
Text Generation
• 0.1B • Updated • 8
mlx-community/VisualQuality-R1-7B-bf16
Reinforcement Learning
• 8B • Updated • 17
mlx-community/VisualQuality-R1-7B-6bit
Reinforcement Learning
• Updated • 13
mlx-community/VisualQuality-R1-7B-8bit
Reinforcement Learning
• Updated • 11
mlx-community/VisualQuality-R1-7B-4bit
Reinforcement Learning
• Updated • 20
• 1
Text Generation
• 0.1B • Updated • 6
kavanmevada/SmolGRPO-135M
Text Generation
• 0.6B • Updated • 7
kavanmevada/SmolGRPO-135M-adapter
Updated
Text Generation
• 0.1B • Updated • 8