# Accuracy Experiments The estimation is accurate enough to ship (and now drives pricing), but the partner's visual review surfaced a recurring **over-count** that inflates prices, plus a seasonal-imagery weakness. This brief scopes three experiments to close that gap. Read `CLAUDE.md` (architecture) and `ROADMAP.md` first. ## Why this matters Our lawn estimate (`lawn_sqft`) is the pricing basis (Phase 6, live). The partner reviewed production (Google-imagery) measurements against his own job sheet and confirmed they're solid overall but **over-count** in three recurring ways, each of which pushes the price up: - **Sidewalk in the right-of-way** counted as lawn. - **Backyard concrete** (patios/pads) counted as lawn. - **Roof edges** bleeding into the lawn mask. He also flagged a season issue: current Google imagery is sometimes **leaf-off (bare trees) with dormant, dry grass**, which the green-keyed color baseline and the segmentation model read poorly. Reconciliation context (see the `lawn-estimator-pricing` memory): our measured lawn runs a median ~1.27× his recorded sqft, so pricing off it is ~+22% vs his historical charges. He accepts that (he under-measured), but tightening the over-count directly improves price accuracy — so accuracy is the current highest-value work. ## Hard constraints — read before touching anything > **HISTORICAL NOTE (2026-07):** the "SAM is LOCAL ONLY" constraint below was written while > SAM was an experiment. It is **superseded** — Experiment 2 shipped: SAM (`facebook/sam-vit-base`) > is now a lazy-loaded prod dependency, pre-baked into the Docker image, behind the `SAM_RESTRICT` > flag (ON in prod). It is imported by `restrict.py` and reached from `strategies.py`. Prod > segmentation still runs on `mfaytin/mask2former-satellite`; SAM is a not-lawn *restriction*, not > the segmenter. The rest of this brief is preserved as-written for provenance. - **Do not perturb production or the byte-identical Douglas gate.** The estimation path is guarded by `data/regression/douglas_qa.csv`: run `python -m lawn_estimator.cli --csv data/regression/douglas_qa.csv --imagery naip` and diff the numeric columns against `data/regression/phase35_douglas.csv`. A change that *intentionally* shifts numbers (Experiment 1 will) is a deliberate re-baseline — gate it behind a config flag and review, don't silently move prod output. - **SAM work is LOCAL and EXPERIMENTAL only.** Do NOT import SAM into `pipeline.py` / `segmentation.py` or the prod dependency set / Docker image. Use `notebooks/` or a scratch script and a separate optional dependency. Prod stays on `mfaytin/mask2former-satellite`. - **Branch/PR flow:** work on a feature branch → PR to `dev` (auto-deploys the staging Space) → PR `dev` → `main` (auto-deploys prod). CI (ruff + pytest) must be green. Never push to the `space` remote by hand. Details in the `lawn-estimator-cicd` memory. - **Env:** always `"c:/Users/sergi/miniconda3/Scripts/conda.exe" run -p ./conda-env --no-capture-output python …` (invoking `conda-env/python.exe` directly crashes numpy's LAPACK with no traceback). `conda run` cannot execute `python -c` with newlines — write a temp `.py` file. - **NAIP ≠ prod.** The QA baseline uses NAIP; production uses **Google** (`imagery=auto` + billed key), which reads ~10–15% lower. Evaluate accuracy on Google imagery, never the NAIP QA CSV. ## Evaluation — the owner judges, visually **The partner's square footage is a loose reference, NOT ground truth.** He under-measures (reconciliation showed our lawn ~1.27× his recorded sqft), so do not tune to match `data/lawn_care_schedule.csv` — treat it only as a rough sanity check. **The real check is visual, and the OWNER makes the call.** Generate the 4-panel visualization for each test address (panel 4 = the measured lawn, cyan outline) and surface the images to the owner so he can verify accuracy himself — the same "View" experience as the live app. - Produce the PNGs by running the pipeline: `python -m lawn_estimator.cli --csv --imagery google` (or `--address "…"`) → PNGs land in `data/outputs/google/`; panel 4 is the measured lawn. - **Show them where the owner can actually see them.** He is often away / on mobile, so the editor side panel isn't enough — publish a **shareable page that embeds the PNGs** (data-URI ``s, openable on a phone) or point him to the staging/live app to run the addresses and hit **View**. For a geometry change (Experiment 1) show **before/after** of the same properties so the effect is obvious. - `data/regression/phase35_douglas.csv` is the NAIP byte-identical gate (regression only, not an accuracy reference). ## Experiment 1 — Carve the sidewalk out of the right-of-way **Idea:** public sidewalks sit at a fairly standard position/width inside the ROW (commonly a ~4–5 ft walk set back a foot or two from the property line). The ROW extension currently adds the full `ROW_BUFFER_FT` (12 ft) strip as candidate lawn on street-facing edges; a standard sidewalk band within that strip is hardscape, not turf. **Where:** `geometry.py` — `parallel_offset_zone` builds the outward quads on street-facing edges (within `STREET_MAX_DIST_M`); `build_estimation_geometry` assembles the estimation geometry + the per-frontage `extensions`. The sidewalk band is a parallel sub-strip to subtract from the ROW quad (or to exclude from the lawn mask within the ROW zone). **Notes / risks:** offsets/widths vary; a fixed assumption is an approximation — validate against the viz. The RGB mask already excludes *some* pavement, so measure the **incremental** effect and don't double-subtract. This touches estimation geometry, so it WILL move Douglas QA numbers → treat as a deliberate, reviewed re-baseline, ideally behind a config flag (off by default) until proven. Highest-value, most tractable, pure geometry — **start here.** ## Experiment 2 — SAM vs mask2former **Idea:** Segment Anything (SAM/SAM2) is class-agnostic and may give tighter turf/hardscape edges than `mfaytin/mask2former-satellite`. The roadmap frames it as boundary *refinement* (prompt SAM with our existing lawn mask), not a wholesale replacement. **Where (reference only — do NOT modify for prod):** `segmentation.py` (`run_lawn_area_model`, `_run_segmentation`, `MODEL_LAWN_CLASS_IDS`). **How:** local notebook/script; install SAM in a separate/optional env (likely needs GPU). Run SAM on the same Google tiles for the partner addresses; compare resulting lawn area/edges to mask2former, against the viz and his sqft. Decide whether SAM as a refinement pass earns its weight. If it graduates, it returns as a proposal — never a silent import into prod. **Findings (2026-07 — tested locally on 10 Douglas lots through the real LiDAR+RGB math):** - SAM runs on CPU with **no new dependency** (`transformers` ships `SamModel`). Naive point-prompt refinement *expands* the mask (wrong direction). "Segment-and-classify" (keep SAM pieces overlapping the mask2former lawn) drops real turf — **82% of what pure-SAM removes is grass** across the sample. - The only safe design is the **hybrid**: mask2former for coverage, SAM subtracts only its *confirmed hardscape* pieces. Effect is small (**~−2%**) because the LiDAR path already ignores roof bleed (roofs return no ground points) and the sidewalk carve (Exp 1) handles the ROW. - A SAM-first **per-segment classifier** (transparent feature vote per piece) *agrees* with mask2former within **±6%** on confident pieces; the only disagreement is shadowed/ambiguous strips, which no signal we have resolves — **LiDAR intensity is a coin flip (AUC 0.51)**. - **Net:** no free accuracy win — the billed number is LiDAR-dominated and every segmentation swap agrees or worsens. See the `lawn-estimator-accuracy-exp2` memory + the review pages. **The promising direction — SAM as a hardscape *restriction* mask (not a lawn finder):** SAM segments house and concrete *beautifully* (crisp boundaries) even without a single "building" class, and both LiDAR and mask2former sometimes eat into the roof. Use SAM's precise house/driveway/patio segments to **restrict the whole pipeline** — exclude any LiDAR ground point (and mask pixel) that falls inside them — to stop that roof/edge bleed. This is the near-term lever, and it also makes SAM the natural **auto-labeler** for Experiment 4. ## Experiment 3 — Season / imagery vintage **Idea:** current Google imagery can be leaf-off with dormant/dry grass. The color baseline `run_vegetation_color_threshold` (`segmentation.py`) keys on green (hue + excess-green), so dry turf fails it, and the model may weaken on dormant grass. **Where:** imagery adapters `sources/imagery.py` (NAIP, county orthos) + `sources/google.py`; `_fetch_analysis_image` in `pipeline.py` selects the source; the model + color baseline in `segmentation.py`. Roadmap "Later" already flags "Winter/leaf-off imagery." **How (measurement-only unless a change proves out):** compare sources/seasons for the same addresses — Google (current, maybe dormant) vs NAIP (leaf-on ~2022) — and measure how `lawn_sqft` and the masks shift. Quantify what dry-grass season costs and whether a source preference or threshold/model tweak helps. **Finding (2026-07): premise NOT reproducible — CLOSED.** Current Google *analysis* imagery is leaf-ON; the dormant-looking tiles were the county-ortho **DISPLAY** layer (never analyzed). An initial "signal" was an artifact of that conflation plus a hardscape-confounded greenness proxy. NAIP is out (resolution too low). No code. Revisit only if Google actually serves leaf-off imagery. See `lawn-estimator-accuracy-exp3`. ## Experiment 4 — Train our own segmentation model (LATER — the real ceiling-raise) **Why:** every off-the-shelf swap (mask2former variants, SAM, per-segment) either agrees with the current number or worsens it — the billed number is LiDAR-dominated and the disputed area is small + genuinely ambiguous (shadowed strips, flat hardscape). The only way to move the ceiling is a model trained on *our own* metro imagery with *our* classes (house / driveway / sidewalk / deck-patio / lawn / canopy). This is the roadmap's "Fine-tuned segmentation" item, promoted to a first-class experiment. **The bottleneck is labeling, not architecture — so this is about AUTO-LABELING / data-engine techniques, not hand-labeling:** - **SAM as the auto-annotator** (Meta's Segment Anything **data engine** — how SA-1B was built: model-assisted masks, human-in-the-loop on the hard cases, progressing toward fully automatic). We've shown SAM segments house/concrete crisply → it proposes object masks we auto-classify (mask2former vote + color + LiDAR ground/roof) into training labels. - **Karpathy-style "data engine" flywheel** (Tesla Autopilot): model proposes labels → focus human review on the *uncertain / failure* cases (our shadowed strips) → retrain → repeat. Active learning, not bulk labeling. - **Free supervision we already have:** LiDAR classification (building=6 vs ground=2) auto-labels roof vs at-grade; county building footprints (Microsoft/Google Open Buildings — a roadmap item) auto-label the house; parcel − footprint − SAM-hardscape ≈ lawn candidate. Combine these weak/auto labels to bootstrap a training set at near-zero labeling cost. **Model + validation:** fine-tune mask2former or train a DeepLabV3+/UNet on the auto-labeled metro set; validate against LiDAR + owner review before it ever replaces `mfaytin/mask2former-satellite`. **Constraints:** training needs a GPU (this box is CPU-only) → cloud GPU, out of band. Strictly experimental; prod stays on the current model until a validated replacement earns it. Method references to evaluate: **Karpathy's auto-labeling / data-engine ("auto-research")** and **Meta's AutoData / SAM data-engine** work. ## Experiment 5 — Green reclaim (turf the model mislabels as water/grass) **Trigger (2026-07-16, owner report):** 8571 Young St, Omaha 68122 — a mostly-lawn new-construction lot measured **1,215 sqft (16.4% of the estimation area)** while the color baseline alone saw 4,399 sqft (59.5%). Obvious under-count, the opposite of the over-counts above. **Root cause (debug_hardscape.py + pinned-revision check — identical either way):** the model splits **large uniform flat turf** into classes we don't count. On the Young St tile, class 6 (published "water") covered **44% of the tile — the entire contiguous backyard lawn — and 98.7% of those pixels pass the green color threshold**. The empirical `[1, 4]` lawn mapping was derived on established, tree-heavy lots where class 1 dominates; on treeless uniform turf (new construction, big open backyards) the model reads the texture as water/grass, the mask excludes it, the LiDAR ground points there are dropped, and the estimate comes out several× low. **Tile survey (23 addresses: 5 QA + 18 schedule, Google imagery):** - Class 6 is **absent (0%) on 20/23 tiles**. Where present it is either turf (Young St 44%, 657 J E George Blvd 3.3% @ 98.8% green, 17531 Madison St 2.4% @ 97% green) or **actual water that fails the green gate** (708 Kountze Memorial Dr, Bellevue lakefront: class 6 only 3.6% green → correctly not reclaimed). - Class 2 (published "grass") is a mixed street-edge band on every tile (7–32%) but only 0–8% green — **blanket-adding class 2 would wreck good tiles**; the green gate keeps its reclaim to +0.2–1.8% of a tile. **Fix (SHIPPED to branch `feat/green-reclaim`, flag `GREEN_RECLAIM`, off by default = byte-identical):** `reclaim_green_turf` in `segmentation.py` — pixels the model assigned to `GREEN_RECLAIM_CLASS_IDS` ([2, 6]) that also pass `run_vegetation_color_threshold` are added back to the lawn mask before it gates the LiDAR points. Reuses the veg mask the pipeline already computes (no extra pass). Zero-network unit tests in `tests/test_segmentation_reclaim.py`. **Limitation:** the green gate can't recover *dormant/dry* mislabeled turf (brown grass fails the color threshold — Young St's unestablished front strips stay uncounted). That remaining gap is Experiment 4 territory. ## Status / order (2026-07) 1. **Experiment 1 (sidewalk)** — **RETIRED (FAILED, owner verdict 2026-07).** Built behind `SIDEWALK_CARVE` on `feat/sidewalk-carve` (−4%, byte-identical off) but the fixed-geometry ROW sub-band carve did not hold up and will NOT ship. Superseded by the SAM/labeler sidewalk-road line-projection direction and the leaf-on/leaf-off fusion carve (empirical, not a fixed band). 2. **Experiment 3 (season)** — **CLOSED**: premise not reproducible (analysis imagery is leaf-on). 3. **Experiment 2 (SAM)** — **EXPLORED**: no free accuracy win (LiDAR-dominated). Safe use is the hybrid / hardscape-restriction mask (~−2%); real value is interpretability + auto-labeling for #4. 4. **Experiment 4 (train our own model)** — IN PROGRESS: Phase 1 bake-off + the shipped cascade (`LAWN_CASCADE`) live in `docs/model-upgrade-plan.md` (the plan of record for Phases 2–3, incl. per-step licensing); research facts in `docs/phase2-findings.md`. Exp 5 (green reclaim, below) shipped as v0.2.