support-ops-env / eval_compare.md
raj921
feat: default GRPO/eval to Qwen/Qwen3-1.7B + LoRA; disable Qwen3 thinking in chat template
eff6196
|
Raw
History Blame Contribute Delete
1.32 kB

DriftShield — eval snapshot (eval_compare.md)

This file is generated by scripts/generate_proof_artifacts.py for hackathon / README evidence. Baseline = two-step hand rollouts per task (opens case then submits a weak resolution). Trained column uses the deterministic scripted strong policy in inference.fallback_action (same trajectories as tests/test_driftshield.py), standing in for a GRPO-tuned LoRA until you attach a real adapter under outputs/.


Baseline vs Trained — DriftShield

  • Base model: Qwen/Qwen3-1.7B
  • Adapter: inference.fallback_action (scripted strong baseline, not LoRA)
  • Curriculum: driftshield
  • Episodes: 1 per run per task

Component means (higher is better except penalty)

Metric Baseline Trained Δ
Total (progress_score) +0.347 +0.962 ▲ +0.615
Investigation +0.050 +1.000 ▲ +0.950
Routing +0.750 +1.000 ▲ +0.250
Reply quality +0.000 +1.000 ▲ +1.000
Groundedness +0.750 +1.000 ▲ +0.250
Submission +0.750 +1.000 ▲ +0.250
Pass rate (total≥0.5) +0.000 +1.000 ▲ +1.000
Done rate +1.000 +1.000 · +0.000
Penalty total (↓) +0.280 +0.037 ▼ -0.243
Turns (mean) +2.500 +4.250 ▲ +1.750