Inference Providers
Active filters: RL
SII-Enigma/Qwen2.5-7B-Ins-AMPO
Text Generation
• 8B • Updated • 23
SII-Enigma/Qwen2.5-7B-Ins-SFT-GRPO
Text Generation
• 8B • Updated • 16
SII-Enigma/Llama3.2-8B-Ins-GRPO
Text Generation
• 2B • Updated • 7
• 1
mradermacher/Llama3.2-8B-Ins-GRPO-GGUF
8B • Updated • 101
• 1
SII-Enigma/Qwen2.5-7B-Ins-GRPO
Text Generation
• 2B • Updated • 5
SII-Enigma/Qwen2.5-1.5B-Ins-AMPO
Text Generation
• 2B • Updated • 8
SII-Enigma/Llama3.2-8B-Ins-AMPO
Text Generation
• 8B • Updated • 9
SII-Enigma/Qwen2.5-1.5B-Ins-GRPO
Text Generation
• 2B • Updated • 3
Text Generation
• 2B • Updated • 5
mradermacher/GCPO-R1-1.5B-GGUF
2B • Updated • 95
mradermacher/GCPO-R1-1.5B-i1-GGUF
2B • Updated • 53
mradermacher/DeepHermes-Egregore-8B-131K-GGUF
Reinforcement Learning
• 8B • Updated • 194
• 1
mradermacher/DeepHermes-Egregore-8B-131K-i1-GGUF
Reinforcement Learning
• 8B • Updated • 135
• 1
stephenchungmh/thinker_r1_5b
2B • Updated • 7
• 1
stephenchungmh/thinker_q1_5b
2B • Updated • 2
• 1
stephenchungmh/thinker_r7b
8B • Updated • 1
• 1
8B • Updated • 7
• 1
mradermacher/RENT-Qwen-7B-GGUF
8B • Updated • 238
• 1
mradermacher/RENT-Qwen-7B-i1-GGUF
8B • Updated • 728
• 1
beyoru/MinCoder-4B-Expert
Text Generation
• 4B • Updated • 12
• • 1
mradermacher/MinCoder-4B-Expert-GGUF
4B • Updated • 76
• 2
mradermacher/MinCoder-4B-Expert-i1-GGUF
4B • Updated • 71
• 1
Text Generation
• 4B • Updated • 1
aryan-kolapkar/MathReasoner-Mini-1.5b
Text Generation
• 2B • Updated • 380
• 1
mradermacher/MathReasoner-Mini-1.5b-GGUF
2B • Updated • 41
ryota39/Qwen3-8B-math-RL-ja
8B • Updated • 6
nvidia/Nemotron-Cascade-8B-Thinking
Text Generation
• 8B • Updated • 348
• • 41
nvidia/Nemotron-Cascade-14B-Thinking
Text Generation
• 15B • Updated • 555
• • 80
8B • Updated • 112
Reinforcement Learning
• Updated • 1