MinCoder
Collection
RL with verify reward • 3 items • Updated • 1
# Install vLLM from pip:
pip install vllm# Start the vLLM server:
vllm serve "beyoru/MaxCoder-4B"# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "beyoru/MaxCoder-4B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/beyoru/MaxCoder-4BThis model is fine-tuned Qwen model using a custom reinforcement learning (RL) framework that rewards the model for producing solutions passing automated test cases — similar to the process of programming task evaluation on LeetCode.
Instead of relying on labeled ground truth answers, the model learns through test-case-based rewards, promoting generalization and reasoning ability in algorithmic problem-
Base model
beyoru/EvolLLM
# Gated model: Login with a HF token with gated access permission hf auth login