Instructions to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="krogoldAI/QueryRefiner-0.5B-v0.1-GRPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("krogoldAI/QueryRefiner-0.5B-v0.1-GRPO") model = AutoModelForCausalLM.from_pretrained("krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/krogoldAI/QueryRefiner-0.5B-v0.1-GRPO
- SGLang
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "krogoldAI/QueryRefiner-0.5B-v0.1-GRPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use krogoldAI/QueryRefiner-0.5B-v0.1-GRPO with Docker Model Runner:
docker model run hf.co/krogoldAI/QueryRefiner-0.5B-v0.1-GRPO
Update README.md
Browse files
README.md
CHANGED
|
@@ -454,7 +454,7 @@ Note: The [ambiguous] tag indicates the analyzer determined the query has multip
|
|
| 454 |
Output only valid JSON. Do not include any explanations, comments, or text outside the JSON structure.
|
| 455 |
"""
|
| 456 |
```
|
| 457 |
-
|
| 458 |
</details>
|
| 459 |
|
| 460 |
To further examine the model's robustness across varying query complexity, we analyzed performance stratified by ambiguity level. The following table shows results broken down by the four ambiguity categories in the test set (NONE: 45%, LOW: 30%, MEDIUM: 20%, HIGH: 5%), which reflect the distribution used during the GRPO training phase.
|
|
@@ -473,18 +473,18 @@ To further examine the model's robustness across varying query complexity, we an
|
|
| 473 |
|
| 474 |
<details>
|
| 475 |
<summary><i>Expand for further discussion of results</i></summary>
|
| 476 |
-
|
| 477 |
The evaluation results reveal a clear task difficulty hierarchy that aligns with the inherent complexity of each component. Structural and classification metrics (domain accuracy, intent accuracy, ambiguity assessment) achieve 98-99% performance, while the generative rephrasing task scores lower at ~90%. This gap reflects the fundamental difference in task complexity rather than a training deficiency.
|
| 478 |
|
| 479 |
Domain and intent classification are essentially pattern recognition tasks where the model must map queries to learned categories—a task well-suited to the model's 0.5B parameter capacity. Similarly, ambiguity assessment and guideline adherence involve rule-following and structural analysis, which the three-phase training procedure was explicitly designed to optimize.
|
| 480 |
|
| 481 |
Rephrasing quality, by contrast, requires the model to make nuanced judgments about when to intervene (versus preserving already-optimal queries), how to balance specificity against over-constraint, and whether to expand acronyms or add disambiguating context. These decisions demand deeper semantic understanding and generation capabilities that push the limits of a 0.5B model. The 90.33% score with 17.61% standard deviation represents strong performance on this challenging task, particularly given the diversity of the test set.
|
| 482 |
|
| 483 |
-
The per-ambiguity breakdown supports this interpretation: rephrasing quality remains relatively stable across ambiguity levels (87.5-90.8%), with the expected slight degradation on highly ambiguous queries. This consistency across query types suggests the model has learned generalizable rephrasing strategies rather than memorizing domain-specific patterns. The ~18% standard deviation across all metrics indicates natural variance in query difficulty rather than systematic failures on particular query types.
|
| 484 |
|
| 485 |
For production deployments, users should anticipate that the model will perform most reliably on structural conformance and classification tasks, while rephrasing decisions may occasionally require human review, particularly for edge cases or highly specialized domains underrepresented in the training data.
|
| 486 |
|
| 487 |
-
</details>
|
| 488 |
|
| 489 |
<!-- Performance analysis across four ambiguity levels shows that QueryRefiner-0.5B-v0.1-GRPO sustains near-saturated accuracy for domain and intent identification (~99%) and achieves perfect structural and ambiguity assessment at the highest ambiguity tier. These results confirm that the model effectively internalized the XML schema and ambiguity-recognition criteria learned during GRPO pretraining.
|
| 490 |
|
|
|
|
| 454 |
Output only valid JSON. Do not include any explanations, comments, or text outside the JSON structure.
|
| 455 |
"""
|
| 456 |
```
|
| 457 |
+
<!--
|
| 458 |
</details>
|
| 459 |
|
| 460 |
To further examine the model's robustness across varying query complexity, we analyzed performance stratified by ambiguity level. The following table shows results broken down by the four ambiguity categories in the test set (NONE: 45%, LOW: 30%, MEDIUM: 20%, HIGH: 5%), which reflect the distribution used during the GRPO training phase.
|
|
|
|
| 473 |
|
| 474 |
<details>
|
| 475 |
<summary><i>Expand for further discussion of results</i></summary>
|
| 476 |
+
-->
|
| 477 |
The evaluation results reveal a clear task difficulty hierarchy that aligns with the inherent complexity of each component. Structural and classification metrics (domain accuracy, intent accuracy, ambiguity assessment) achieve 98-99% performance, while the generative rephrasing task scores lower at ~90%. This gap reflects the fundamental difference in task complexity rather than a training deficiency.
|
| 478 |
|
| 479 |
Domain and intent classification are essentially pattern recognition tasks where the model must map queries to learned categories—a task well-suited to the model's 0.5B parameter capacity. Similarly, ambiguity assessment and guideline adherence involve rule-following and structural analysis, which the three-phase training procedure was explicitly designed to optimize.
|
| 480 |
|
| 481 |
Rephrasing quality, by contrast, requires the model to make nuanced judgments about when to intervene (versus preserving already-optimal queries), how to balance specificity against over-constraint, and whether to expand acronyms or add disambiguating context. These decisions demand deeper semantic understanding and generation capabilities that push the limits of a 0.5B model. The 90.33% score with 17.61% standard deviation represents strong performance on this challenging task, particularly given the diversity of the test set.
|
| 482 |
|
| 483 |
+
<!-- The per-ambiguity breakdown supports this interpretation: rephrasing quality remains relatively stable across ambiguity levels (87.5-90.8%), with the expected slight degradation on highly ambiguous queries. This consistency across query types suggests the model has learned generalizable rephrasing strategies rather than memorizing domain-specific patterns. The ~18% standard deviation across all metrics indicates natural variance in query difficulty rather than systematic failures on particular query types. -->
|
| 484 |
|
| 485 |
For production deployments, users should anticipate that the model will perform most reliably on structural conformance and classification tasks, while rephrasing decisions may occasionally require human review, particularly for edge cases or highly specialized domains underrepresented in the training data.
|
| 486 |
|
| 487 |
+
<!--</details>-->
|
| 488 |
|
| 489 |
<!-- Performance analysis across four ambiguity levels shows that QueryRefiner-0.5B-v0.1-GRPO sustains near-saturated accuracy for domain and intent identification (~99%) and achieves perfect structural and ambiguity assessment at the highest ambiguity tier. These results confirm that the model effectively internalized the XML schema and ambiguity-recognition criteria learned during GRPO pretraining.
|
| 490 |
|