"""Labeled before/after accuracy evaluation runner. Runs the real pipeline twice over an address list — a BEFORE leg (base env flags only) and an AFTER leg (base + toggled flags) — and saves everything under a labeled, dated directory so results are reusable and comparable across sessions: data/evals/