Spaces:
Runtime error
A newer version of the Gradio SDK is available: 6.22.0
Calibration report β simulator vs. real observations
Date: 2026-06-12 Β· Harness: uv run python -m sim.calibrate --data <obs> Β·
Model under test: sim/outcome.py (BANDS + penalty constants).
Result: data cleaned, structural bottleneck identified
| Dataset | n | Outcome accuracy | Quality MAE | fan=0 rows |
|---|---|---|---|---|
observations.sample.jsonl (hand-crafted) |
12 | 66.7% | 0.228 | 2 |
observations.modal.jsonl (FDM error detection, cleaned) |
178 | 32.6% | 0.227 | 0 |
Confusion on the 178-row cleaned FDM set:
β failed_sag β success Γ61
β failed_stringing β success Γ59
β success β success Γ58
Phase 2 fixes applied (per RESEARCH-NEEDS.md)
- β M106 parsed across full file + slicer footer β fan=0 rows: 178β0
- β M109/M190 final temps β nozzle temps now 210/215/220 (was preheat 150)
- β Retraction from G1 E- moves + PrusaSlicer footer β real values 1-8.5mm
- β Rows with undetermined fan skipped β no more fan=0 defaults
What the clean data reveals
The simulator correctly identifies all 58 success cases but misses all 120 failures. This is NOT a constant-tuning problem β it's a structural sensitivity gap:
Sag (61 misses): At fan=67%, the sag formula produces penalty=0 because
need=60for overhang. Raisingneedto 75 produces sag=0.056, but the quality threshold (0.7) requires total penalty >0.3 to flip outcome. The penalty formulas are too weak for moderate conditions.Stringing (59 misses): At humidity=45%, nozzle=210Β°C (in-band), retraction=3mm, the stringing formula produces 0. A base penalty for stringing geometry would need to be >0.5 to overcome the 0.7 threshold β implausibly strong. The formula needs additional features (travel speed, filament moisture, part complexity).
Decision: no constant changes
The 0.7 quality threshold + weak penalty scaling make the simulator insensitive to moderate failure conditions. This is a structural design choice β the simulator was calibrated for the sample data's more extreme conditions. Fixing it requires either:
- Lowering the quality threshold (affects all failure modes β risky)
- Adding features to penalty formulas (travel speed, extrusion multiplier, filament moisture)
- Per-geometry quality thresholds
Path forward: Add extrusion_multiplier and travel_speed to the Lane C
schema and the simulator inputs. These are the missing features that distinguish
success from failure in the FDM data.
What we learned from the bad data (the interesting part)
- Quantity β quality. The ingestion produced 1,304 reference facts, but the
correct, prompt-useful material baselines all came from 4 tiny curated config
files (Prusa/Klipper/Marlin) β not the 825 MB dataset.
reference_blocknow prefers those curated sources and drops the bulk model-metadata. - One unparsed field silently flipped 112 predictions.
fanβ0looked like real data and would have corrupted the simulator had we tuned to it blind. The harness caught it β calibration is also a data-quality check, not just a model check. - Honest-by-design, demonstrated. We did not manufacture a "67%β92%" win. We measured 34.2%, found the dataset wasn't trustworthy, and said so. That is the same integrity rule the whole project rests on, applied to our own results.
Reproduce
uv run python -m sim.calibrate --data sim/calibration/observations.modal.jsonl
# machine-readable metrics: sim/calibration/results/baseline.modal.json