Text Generation
Transformers
Safetensors
English
qwen2
conversational
text-generation-inference
krogoldAI commited on
Commit
ba78e12
·
verified ·
1 Parent(s): 684f5db

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +4 -4
README.md CHANGED
@@ -454,7 +454,7 @@ Note: The [ambiguous] tag indicates the analyzer determined the query has multip
454
  Output only valid JSON. Do not include any explanations, comments, or text outside the JSON structure.
455
  """
456
  ```
457
-
458
  </details>
459
 
460
  To further examine the model's robustness across varying query complexity, we analyzed performance stratified by ambiguity level. The following table shows results broken down by the four ambiguity categories in the test set (NONE: 45%, LOW: 30%, MEDIUM: 20%, HIGH: 5%), which reflect the distribution used during the GRPO training phase.
@@ -473,18 +473,18 @@ To further examine the model's robustness across varying query complexity, we an
473
 
474
  <details>
475
  <summary><i>Expand for further discussion of results</i></summary>
476
-
477
  The evaluation results reveal a clear task difficulty hierarchy that aligns with the inherent complexity of each component. Structural and classification metrics (domain accuracy, intent accuracy, ambiguity assessment) achieve 98-99% performance, while the generative rephrasing task scores lower at ~90%. This gap reflects the fundamental difference in task complexity rather than a training deficiency.
478
 
479
  Domain and intent classification are essentially pattern recognition tasks where the model must map queries to learned categories—a task well-suited to the model's 0.5B parameter capacity. Similarly, ambiguity assessment and guideline adherence involve rule-following and structural analysis, which the three-phase training procedure was explicitly designed to optimize.
480
 
481
  Rephrasing quality, by contrast, requires the model to make nuanced judgments about when to intervene (versus preserving already-optimal queries), how to balance specificity against over-constraint, and whether to expand acronyms or add disambiguating context. These decisions demand deeper semantic understanding and generation capabilities that push the limits of a 0.5B model. The 90.33% score with 17.61% standard deviation represents strong performance on this challenging task, particularly given the diversity of the test set.
482
 
483
- The per-ambiguity breakdown supports this interpretation: rephrasing quality remains relatively stable across ambiguity levels (87.5-90.8%), with the expected slight degradation on highly ambiguous queries. This consistency across query types suggests the model has learned generalizable rephrasing strategies rather than memorizing domain-specific patterns. The ~18% standard deviation across all metrics indicates natural variance in query difficulty rather than systematic failures on particular query types.
484
 
485
  For production deployments, users should anticipate that the model will perform most reliably on structural conformance and classification tasks, while rephrasing decisions may occasionally require human review, particularly for edge cases or highly specialized domains underrepresented in the training data.
486
 
487
- </details>
488
 
489
  <!-- Performance analysis across four ambiguity levels shows that QueryRefiner-0.5B-v0.1-GRPO sustains near-saturated accuracy for domain and intent identification (~99%) and achieves perfect structural and ambiguity assessment at the highest ambiguity tier. These results confirm that the model effectively internalized the XML schema and ambiguity-recognition criteria learned during GRPO pretraining.
490
 
 
454
  Output only valid JSON. Do not include any explanations, comments, or text outside the JSON structure.
455
  """
456
  ```
457
+ <!--
458
  </details>
459
 
460
  To further examine the model's robustness across varying query complexity, we analyzed performance stratified by ambiguity level. The following table shows results broken down by the four ambiguity categories in the test set (NONE: 45%, LOW: 30%, MEDIUM: 20%, HIGH: 5%), which reflect the distribution used during the GRPO training phase.
 
473
 
474
  <details>
475
  <summary><i>Expand for further discussion of results</i></summary>
476
+ -->
477
  The evaluation results reveal a clear task difficulty hierarchy that aligns with the inherent complexity of each component. Structural and classification metrics (domain accuracy, intent accuracy, ambiguity assessment) achieve 98-99% performance, while the generative rephrasing task scores lower at ~90%. This gap reflects the fundamental difference in task complexity rather than a training deficiency.
478
 
479
  Domain and intent classification are essentially pattern recognition tasks where the model must map queries to learned categories—a task well-suited to the model's 0.5B parameter capacity. Similarly, ambiguity assessment and guideline adherence involve rule-following and structural analysis, which the three-phase training procedure was explicitly designed to optimize.
480
 
481
  Rephrasing quality, by contrast, requires the model to make nuanced judgments about when to intervene (versus preserving already-optimal queries), how to balance specificity against over-constraint, and whether to expand acronyms or add disambiguating context. These decisions demand deeper semantic understanding and generation capabilities that push the limits of a 0.5B model. The 90.33% score with 17.61% standard deviation represents strong performance on this challenging task, particularly given the diversity of the test set.
482
 
483
+ <!-- The per-ambiguity breakdown supports this interpretation: rephrasing quality remains relatively stable across ambiguity levels (87.5-90.8%), with the expected slight degradation on highly ambiguous queries. This consistency across query types suggests the model has learned generalizable rephrasing strategies rather than memorizing domain-specific patterns. The ~18% standard deviation across all metrics indicates natural variance in query difficulty rather than systematic failures on particular query types. -->
484
 
485
  For production deployments, users should anticipate that the model will perform most reliably on structural conformance and classification tasks, while rephrasing decisions may occasionally require human review, particularly for edge cases or highly specialized domains underrepresented in the training data.
486
 
487
+ <!--</details>-->
488
 
489
  <!-- Performance analysis across four ambiguity levels shows that QueryRefiner-0.5B-v0.1-GRPO sustains near-saturated accuracy for domain and intent identification (~99%) and achieves perfect structural and ambiguity assessment at the highest ambiguity tier. These results confirm that the model effectively internalized the XML schema and ambiguity-recognition criteria learned during GRPO pretraining.
490