Spaces:
Paused
Paused
raj921
feat: default GRPO/eval to Qwen/Qwen3-1.7B + LoRA; disable Qwen3 thinking in chat template
eff6196
DriftShield — eval snapshot (eval_compare.md)
This file is generated by scripts/generate_proof_artifacts.py for hackathon / README evidence. Baseline = two-step hand rollouts per task (opens case then submits a weak resolution). Trained column uses the deterministic scripted strong policy in inference.fallback_action (same trajectories as tests/test_driftshield.py), standing in for a GRPO-tuned LoRA until you attach a real adapter under outputs/.
Baseline vs Trained — DriftShield
- Base model:
Qwen/Qwen3-1.7B - Adapter:
inference.fallback_action (scripted strong baseline, not LoRA) - Curriculum:
driftshield - Episodes: 1 per run per task
Component means (higher is better except penalty)
| Metric | Baseline | Trained | Δ |
|---|---|---|---|
| Total (progress_score) | +0.347 | +0.962 | ▲ +0.615 |
| Investigation | +0.050 | +1.000 | ▲ +0.950 |
| Routing | +0.750 | +1.000 | ▲ +0.250 |
| Reply quality | +0.000 | +1.000 | ▲ +1.000 |
| Groundedness | +0.750 | +1.000 | ▲ +0.250 |
| Submission | +0.750 | +1.000 | ▲ +0.250 |
| Pass rate (total≥0.5) | +0.000 | +1.000 | ▲ +1.000 |
| Done rate | +1.000 | +1.000 | · +0.000 |
| Penalty total (↓) | +0.280 | +0.037 | ▼ -0.243 |
| Turns (mean) | +2.500 | +4.250 | ▲ +1.750 |