Instructions to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="krogoldAI/QueryRefiner-0.5B-v0.1-GRPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("krogoldAI/QueryRefiner-0.5B-v0.1-GRPO") model = AutoModelForCausalLM.from_pretrained("krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/krogoldAI/QueryRefiner-0.5B-v0.1-GRPO
- SGLang
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with Docker Model Runner:
docker model run hf.co/krogoldAI/QueryRefiner-0.5B-v0.1-GRPO
Update README.md
Browse files
README.md
CHANGED
|
@@ -152,9 +152,10 @@ This stratified distribution ensures the model learns to handle the full spectru
|
|
| 152 |
|
| 153 |
## Training Procedure
|
| 154 |
|
| 155 |
-
QueryRefiner-0.5B-v0.1-GRPO underwent a
|
| 156 |
|
| 157 |
### Phase 1: Reinforcement Learning with GRPO
|
|
|
|
| 158 |
The model was first trained using Group Relative Policy Optimization (GRPO) on 3,000 examples from [krogoldAI/rag-ambiguous-queries](https://huggingface.co/datasets/krogoldAI/rag-ambiguous-queries) for 1 epoch with a learning rate of 5e-6. This phase focused on learning correct XML structure and formatting before semantic refinement.
|
| 159 |
|
| 160 |
The GRPO reward function evaluated outputs through a weighted combination of five components: tag structure (30%), XML validity (25%), element ordering (25%), confidence formatting (18%), and confidence distribution (2%). The structure component verified the presence of all required XML elements, while validity ensured parseability. The ordering component checked that tags appeared in the correct sequence, and the confidence component validated that confidence values were properly formatted and summed to `1.0` for ambiguous cases.
|
|
|
|
| 152 |
|
| 153 |
## Training Procedure
|
| 154 |
|
| 155 |
+
QueryRefiner-0.5B-v0.1-GRPO underwent a three-phase training procedure combining reinforcement learning with two-stage supervised fine-tuning, all using full parameter updates (not parameter-efficient methods like LoRA) on [Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct).
|
| 156 |
|
| 157 |
### Phase 1: Reinforcement Learning with GRPO
|
| 158 |
+
|
| 159 |
The model was first trained using Group Relative Policy Optimization (GRPO) on 3,000 examples from [krogoldAI/rag-ambiguous-queries](https://huggingface.co/datasets/krogoldAI/rag-ambiguous-queries) for 1 epoch with a learning rate of 5e-6. This phase focused on learning correct XML structure and formatting before semantic refinement.
|
| 160 |
|
| 161 |
The GRPO reward function evaluated outputs through a weighted combination of five components: tag structure (30%), XML validity (25%), element ordering (25%), confidence formatting (18%), and confidence distribution (2%). The structure component verified the presence of all required XML elements, while validity ensured parseability. The ordering component checked that tags appeared in the correct sequence, and the confidence component validated that confidence values were properly formatted and summed to `1.0` for ambiguous cases.
|