lawn-estimator-dev / docs /accuracy-experiments.md
TempuraML's picture
feat(tuning): leaf-on/leaf-off fusion β€” carve hidden under-canopy hardscape
84e2326
|
Raw
History Blame Contribute Delete
15.3 kB
# Accuracy Experiments
The estimation is accurate enough to ship (and now drives pricing), but the partner's
visual review surfaced a recurring **over-count** that inflates prices, plus a
seasonal-imagery weakness. This brief scopes three experiments to close that gap. Read
`CLAUDE.md` (architecture) and `ROADMAP.md` first.
## Why this matters
Our lawn estimate (`lawn_sqft`) is the pricing basis (Phase 6, live). The partner
reviewed production (Google-imagery) measurements against his own job sheet and confirmed
they're solid overall but **over-count** in three recurring ways, each of which pushes the
price up:
- **Sidewalk in the right-of-way** counted as lawn.
- **Backyard concrete** (patios/pads) counted as lawn.
- **Roof edges** bleeding into the lawn mask.
He also flagged a season issue: current Google imagery is sometimes **leaf-off (bare
trees) with dormant, dry grass**, which the green-keyed color baseline and the
segmentation model read poorly.
Reconciliation context (see the `lawn-estimator-pricing` memory): our measured lawn runs a
median ~1.27Γ— his recorded sqft, so pricing off it is ~+22% vs his historical charges. He
accepts that (he under-measured), but tightening the over-count directly improves price
accuracy β€” so accuracy is the current highest-value work.
## Hard constraints β€” read before touching anything
> **HISTORICAL NOTE (2026-07):** the "SAM is LOCAL ONLY" constraint below was written while
> SAM was an experiment. It is **superseded** β€” Experiment 2 shipped: SAM (`facebook/sam-vit-base`)
> is now a lazy-loaded prod dependency, pre-baked into the Docker image, behind the `SAM_RESTRICT`
> flag (ON in prod). It is imported by `restrict.py` and reached from `strategies.py`. Prod
> segmentation still runs on `mfaytin/mask2former-satellite`; SAM is a not-lawn *restriction*, not
> the segmenter. The rest of this brief is preserved as-written for provenance.
- **Do not perturb production or the byte-identical Douglas gate.** The estimation path is
guarded by `data/regression/douglas_qa.csv`: run
`python -m lawn_estimator.cli --csv data/regression/douglas_qa.csv --imagery naip` and
diff the numeric columns against `data/regression/phase35_douglas.csv`. A change that
*intentionally* shifts numbers (Experiment 1 will) is a deliberate re-baseline β€” gate it
behind a config flag and review, don't silently move prod output.
- **SAM work is LOCAL and EXPERIMENTAL only.** Do NOT import SAM into `pipeline.py` /
`segmentation.py` or the prod dependency set / Docker image. Use `notebooks/` or a scratch
script and a separate optional dependency. Prod stays on `mfaytin/mask2former-satellite`.
- **Branch/PR flow:** work on a feature branch β†’ PR to `dev` (auto-deploys the staging
Space) β†’ PR `dev` β†’ `main` (auto-deploys prod). CI (ruff + pytest) must be green. Never
push to the `space` remote by hand. Details in the `lawn-estimator-cicd` memory.
- **Env:** always `"c:/Users/sergi/miniconda3/Scripts/conda.exe" run -p ./conda-env
--no-capture-output python …` (invoking `conda-env/python.exe` directly crashes numpy's
LAPACK with no traceback). `conda run` cannot execute `python -c` with newlines β€” write a
temp `.py` file.
- **NAIP β‰  prod.** The QA baseline uses NAIP; production uses **Google** (`imagery=auto` +
billed key), which reads ~10–15% lower. Evaluate accuracy on Google imagery, never the
NAIP QA CSV.
## Evaluation β€” the owner judges, visually
**The partner's square footage is a loose reference, NOT ground truth.** He under-measures
(reconciliation showed our lawn ~1.27Γ— his recorded sqft), so do not tune to match
`data/lawn_care_schedule.csv` β€” treat it only as a rough sanity check.
**The real check is visual, and the OWNER makes the call.** Generate the 4-panel
visualization for each test address (panel 4 = the measured lawn, cyan outline) and surface
the images to the owner so he can verify accuracy himself β€” the same "View" experience as
the live app.
- Produce the PNGs by running the pipeline: `python -m lawn_estimator.cli --csv <addresses>
--imagery google` (or `--address "…"`) β†’ PNGs land in `data/outputs/google/`; panel 4 is
the measured lawn.
- **Show them where the owner can actually see them.** He is often away / on mobile, so the
editor side panel isn't enough β€” publish a **shareable page that embeds the PNGs** (data-URI
`<img>`s, openable on a phone) or point him to the staging/live app to run the addresses and
hit **View**. For a geometry change (Experiment 1) show **before/after** of the same
properties so the effect is obvious.
- `data/regression/phase35_douglas.csv` is the NAIP byte-identical gate (regression only, not
an accuracy reference).
## Experiment 1 β€” Carve the sidewalk out of the right-of-way
**Idea:** public sidewalks sit at a fairly standard position/width inside the ROW (commonly
a ~4–5 ft walk set back a foot or two from the property line). The ROW extension currently
adds the full `ROW_BUFFER_FT` (12 ft) strip as candidate lawn on street-facing edges; a
standard sidewalk band within that strip is hardscape, not turf.
**Where:** `geometry.py` β€” `parallel_offset_zone` builds the outward quads on street-facing
edges (within `STREET_MAX_DIST_M`); `build_estimation_geometry` assembles the estimation
geometry + the per-frontage `extensions`. The sidewalk band is a parallel sub-strip to
subtract from the ROW quad (or to exclude from the lawn mask within the ROW zone).
**Notes / risks:** offsets/widths vary; a fixed assumption is an approximation β€” validate
against the viz. The RGB mask already excludes *some* pavement, so measure the
**incremental** effect and don't double-subtract. This touches estimation geometry, so it
WILL move Douglas QA numbers β†’ treat as a deliberate, reviewed re-baseline, ideally behind
a config flag (off by default) until proven. Highest-value, most tractable, pure geometry
β€” **start here.**
## Experiment 2 β€” SAM vs mask2former
**Idea:** Segment Anything (SAM/SAM2) is class-agnostic and may give tighter turf/hardscape
edges than `mfaytin/mask2former-satellite`. The roadmap frames it as boundary *refinement*
(prompt SAM with our existing lawn mask), not a wholesale replacement.
**Where (reference only β€” do NOT modify for prod):** `segmentation.py`
(`run_lawn_area_model`, `_run_segmentation`, `MODEL_LAWN_CLASS_IDS`).
**How:** local notebook/script; install SAM in a separate/optional env (likely needs GPU).
Run SAM on the same Google tiles for the partner addresses; compare resulting lawn
area/edges to mask2former, against the viz and his sqft. Decide whether SAM as a refinement
pass earns its weight. If it graduates, it returns as a proposal β€” never a silent import
into prod.
**Findings (2026-07 β€” tested locally on 10 Douglas lots through the real LiDAR+RGB math):**
- SAM runs on CPU with **no new dependency** (`transformers` ships `SamModel`). Naive
point-prompt refinement *expands* the mask (wrong direction). "Segment-and-classify" (keep SAM
pieces overlapping the mask2former lawn) drops real turf β€” **82% of what pure-SAM removes is
grass** across the sample.
- The only safe design is the **hybrid**: mask2former for coverage, SAM subtracts only its
*confirmed hardscape* pieces. Effect is small (**~βˆ’2%**) because the LiDAR path already ignores
roof bleed (roofs return no ground points) and the sidewalk carve (Exp 1) handles the ROW.
- A SAM-first **per-segment classifier** (transparent feature vote per piece) *agrees* with
mask2former within **Β±6%** on confident pieces; the only disagreement is shadowed/ambiguous
strips, which no signal we have resolves β€” **LiDAR intensity is a coin flip (AUC 0.51)**.
- **Net:** no free accuracy win β€” the billed number is LiDAR-dominated and every segmentation
swap agrees or worsens. See the `lawn-estimator-accuracy-exp2` memory + the review pages.
**The promising direction β€” SAM as a hardscape *restriction* mask (not a lawn finder):** SAM
segments house and concrete *beautifully* (crisp boundaries) even without a single "building"
class, and both LiDAR and mask2former sometimes eat into the roof. Use SAM's precise
house/driveway/patio segments to **restrict the whole pipeline** β€” exclude any LiDAR ground point
(and mask pixel) that falls inside them β€” to stop that roof/edge bleed. This is the near-term
lever, and it also makes SAM the natural **auto-labeler** for Experiment 4.
## Experiment 3 β€” Season / imagery vintage
**Idea:** current Google imagery can be leaf-off with dormant/dry grass. The color baseline
`run_vegetation_color_threshold` (`segmentation.py`) keys on green (hue + excess-green), so
dry turf fails it, and the model may weaken on dormant grass.
**Where:** imagery adapters `sources/imagery.py` (NAIP, county orthos) + `sources/google.py`;
`_fetch_analysis_image` in `pipeline.py` selects the source; the model + color baseline in
`segmentation.py`. Roadmap "Later" already flags "Winter/leaf-off imagery."
**How (measurement-only unless a change proves out):** compare sources/seasons for the same
addresses β€” Google (current, maybe dormant) vs NAIP (leaf-on ~2022) β€” and measure how
`lawn_sqft` and the masks shift. Quantify what dry-grass season costs and whether a source
preference or threshold/model tweak helps.
**Finding (2026-07): premise NOT reproducible β€” CLOSED.** Current Google *analysis* imagery is
leaf-ON; the dormant-looking tiles were the county-ortho **DISPLAY** layer (never analyzed). An
initial "signal" was an artifact of that conflation plus a hardscape-confounded greenness proxy.
NAIP is out (resolution too low). No code. Revisit only if Google actually serves leaf-off
imagery. See `lawn-estimator-accuracy-exp3`.
## Experiment 4 β€” Train our own segmentation model (LATER β€” the real ceiling-raise)
**Why:** every off-the-shelf swap (mask2former variants, SAM, per-segment) either agrees with the
current number or worsens it β€” the billed number is LiDAR-dominated and the disputed area is small
+ genuinely ambiguous (shadowed strips, flat hardscape). The only way to move the ceiling is a
model trained on *our own* metro imagery with *our* classes (house / driveway / sidewalk /
deck-patio / lawn / canopy). This is the roadmap's "Fine-tuned segmentation" item, promoted to a
first-class experiment.
**The bottleneck is labeling, not architecture β€” so this is about AUTO-LABELING / data-engine
techniques, not hand-labeling:**
- **SAM as the auto-annotator** (Meta's Segment Anything **data engine** β€” how SA-1B was built:
model-assisted masks, human-in-the-loop on the hard cases, progressing toward fully automatic).
We've shown SAM segments house/concrete crisply β†’ it proposes object masks we auto-classify
(mask2former vote + color + LiDAR ground/roof) into training labels.
- **Karpathy-style "data engine" flywheel** (Tesla Autopilot): model proposes labels β†’ focus human
review on the *uncertain / failure* cases (our shadowed strips) β†’ retrain β†’ repeat. Active
learning, not bulk labeling.
- **Free supervision we already have:** LiDAR classification (building=6 vs ground=2) auto-labels
roof vs at-grade; county building footprints (Microsoft/Google Open Buildings β€” a roadmap item)
auto-label the house; parcel βˆ’ footprint βˆ’ SAM-hardscape β‰ˆ lawn candidate. Combine these weak/auto
labels to bootstrap a training set at near-zero labeling cost.
**Model + validation:** fine-tune mask2former or train a DeepLabV3+/UNet on the auto-labeled metro
set; validate against LiDAR + owner review before it ever replaces `mfaytin/mask2former-satellite`.
**Constraints:** training needs a GPU (this box is CPU-only) β†’ cloud GPU, out of band. Strictly
experimental; prod stays on the current model until a validated replacement earns it. Method
references to evaluate: **Karpathy's auto-labeling / data-engine ("auto-research")** and **Meta's
AutoData / SAM data-engine** work.
## Experiment 5 β€” Green reclaim (turf the model mislabels as water/grass)
**Trigger (2026-07-16, owner report):** 8571 Young St, Omaha 68122 β€” a mostly-lawn
new-construction lot measured **1,215 sqft (16.4% of the estimation area)** while the
color baseline alone saw 4,399 sqft (59.5%). Obvious under-count, the opposite of the
over-counts above.
**Root cause (debug_hardscape.py + pinned-revision check β€” identical either way):** the
model splits **large uniform flat turf** into classes we don't count. On the Young St tile,
class 6 (published "water") covered **44% of the tile β€” the entire contiguous backyard
lawn β€” and 98.7% of those pixels pass the green color threshold**. The empirical `[1, 4]`
lawn mapping was derived on established, tree-heavy lots where class 1 dominates; on
treeless uniform turf (new construction, big open backyards) the model reads the texture
as water/grass, the mask excludes it, the LiDAR ground points there are dropped, and the
estimate comes out severalΓ— low.
**Tile survey (23 addresses: 5 QA + 18 schedule, Google imagery):**
- Class 6 is **absent (0%) on 20/23 tiles**. Where present it is either turf
(Young St 44%, 657 J E George Blvd 3.3% @ 98.8% green, 17531 Madison St 2.4% @ 97%
green) or **actual water that fails the green gate** (708 Kountze Memorial Dr,
Bellevue lakefront: class 6 only 3.6% green β†’ correctly not reclaimed).
- Class 2 (published "grass") is a mixed street-edge band on every tile (7–32%) but only
0–8% green β€” **blanket-adding class 2 would wreck good tiles**; the green gate keeps
its reclaim to +0.2–1.8% of a tile.
**Fix (SHIPPED to branch `feat/green-reclaim`, flag `GREEN_RECLAIM`, off by default =
byte-identical):** `reclaim_green_turf` in `segmentation.py` β€” pixels the model assigned
to `GREEN_RECLAIM_CLASS_IDS` ([2, 6]) that also pass `run_vegetation_color_threshold`
are added back to the lawn mask before it gates the LiDAR points. Reuses the veg mask the
pipeline already computes (no extra pass). Zero-network unit tests in
`tests/test_segmentation_reclaim.py`.
**Limitation:** the green gate can't recover *dormant/dry* mislabeled turf (brown grass
fails the color threshold β€” Young St's unestablished front strips stay uncounted). That
remaining gap is Experiment 4 territory.
## Status / order (2026-07)
1. **Experiment 1 (sidewalk)** β€” **RETIRED (FAILED, owner verdict 2026-07).** Built behind
`SIDEWALK_CARVE` on `feat/sidewalk-carve` (βˆ’4%, byte-identical off) but the fixed-geometry ROW
sub-band carve did not hold up and will NOT ship. Superseded by the SAM/labeler sidewalk-road
line-projection direction and the leaf-on/leaf-off fusion carve (empirical, not a fixed band).
2. **Experiment 3 (season)** β€” **CLOSED**: premise not reproducible (analysis imagery is leaf-on).
3. **Experiment 2 (SAM)** β€” **EXPLORED**: no free accuracy win (LiDAR-dominated). Safe use is the
hybrid / hardscape-restriction mask (~βˆ’2%); real value is interpretability + auto-labeling for #4.
4. **Experiment 4 (train our own model)** β€” IN PROGRESS: Phase 1 bake-off + the
shipped cascade (`LAWN_CASCADE`) live in `docs/model-upgrade-plan.md` (the
plan of record for Phases 2–3, incl. per-step licensing); research facts in
`docs/phase2-findings.md`. Exp 5 (green reclaim, below) shipped as v0.2.