Inference Providers
Active filters: rlvr
AIML-TUDA/OlmoLogic-7B-Think
Text Generation
• 528k • Updated • 87
• 4
lastmass/Qwen3.5-Medical-GSPO
Image-Text-to-Text
• 5B • Updated • 4.98k
• 12
mradermacher/Qwen3.5-Medical-GSPO-GGUF
4B • Updated • 449
• 1
Text Generation
• 3B • Updated • 218
• 1
Chinzhu/BirdAgent-Qwen3VL-4B
Image-Text-to-Text
• Updated • 40
• 1
SultanR/SmolTulu-1.7b-Reinforced-GGUF
Text Generation
• 2B • Updated • 9
• 1
thuml/rt1-world-model-multi-step-rlvr
Updated • 32
thuml/rt1-world-model-single-step-rlvr
Updated • 22
thuml/webarena-world-model-rlvr
2B • Updated • 8
thuml/bytesized32-world-model-rlvr-binary-reward
2B • Updated • 10
thuml/bytesized32-world-model-rlvr-task-specific-reward
2B • Updated • 3
DebateLabKIT/Llama-3.1-Argunaut-1-8B-HIRPO
Text Generation
• 8B • Updated • 10
• 1
Question Answering
• 4B • Updated • 6
• 2
thinkwee/NOVER1-Qwen2.5-7B
Question Answering
• 8B • Updated • 4
• 2
mradermacher/NOVER1-Qwen3-4B-GGUF
4B • Updated • 198
• 1
mradermacher/NOVER1-Qwen2.5-7B-GGUF
8B • Updated • 68
• 1
mradermacher/NOVER1-Qwen3-4B-i1-GGUF
4B • Updated • 110
• 1
mradermacher/NOVER1-Qwen2.5-7B-i1-GGUF
8B • Updated • 421
• 1
DebateLabKIT/Phi-4-Argunaut-1-HIRPO
Text Generation
• 415k • Updated • 24
mradermacher/Llama-3.1-Argunaut-1-8B-HIRPO-GGUF
8B • Updated • 113
• 1
mradermacher/Llama-3.1-Argunaut-1-8B-HIRPO-i1-GGUF
8B • Updated • 166
• 1
Text Generation
• 2B • Updated • 29
• 9
Text Generation
• 4B • Updated • 5
• 1
mradermacher/airesupdated-v2-GGUF
Reinforcement Learning
• 4B • Updated • 72
ABaroian/Apertus-8B-RLVR-GSM
Text Generation
• Updated • 2
Anonymouslolol/qwen3-8B-hanabi-step110
Reinforcement Learning
• Updated • 2
Text Generation
• 4B • Updated • 1
anonymousatom/IntelliAsk-Qwen3-32B-450-Merged
Text Generation
• 33B • Updated • 5
mradermacher/IntelliAsk-Qwen3-32B-450-Merged-GGUF
Reinforcement Learning
• 33B • Updated • 110
mradermacher/Phi-4-Argunaut-1-HIRPO-GGUF
15B • Updated • 54