Spaces:
Running
Running
File size: 15,337 Bytes
cebc4ea 7ee8be4 cebc4ea 22e96c3 cebc4ea 7d33a76 cebc4ea 7d33a76 466c0c5 7d33a76 84e2326 7d33a76 31ebc97 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | # Accuracy Experiments
The estimation is accurate enough to ship (and now drives pricing), but the partner's
visual review surfaced a recurring **over-count** that inflates prices, plus a
seasonal-imagery weakness. This brief scopes three experiments to close that gap. Read
`CLAUDE.md` (architecture) and `ROADMAP.md` first.
## Why this matters
Our lawn estimate (`lawn_sqft`) is the pricing basis (Phase 6, live). The partner
reviewed production (Google-imagery) measurements against his own job sheet and confirmed
they're solid overall but **over-count** in three recurring ways, each of which pushes the
price up:
- **Sidewalk in the right-of-way** counted as lawn.
- **Backyard concrete** (patios/pads) counted as lawn.
- **Roof edges** bleeding into the lawn mask.
He also flagged a season issue: current Google imagery is sometimes **leaf-off (bare
trees) with dormant, dry grass**, which the green-keyed color baseline and the
segmentation model read poorly.
Reconciliation context (see the `lawn-estimator-pricing` memory): our measured lawn runs a
median ~1.27Γ his recorded sqft, so pricing off it is ~+22% vs his historical charges. He
accepts that (he under-measured), but tightening the over-count directly improves price
accuracy β so accuracy is the current highest-value work.
## Hard constraints β read before touching anything
> **HISTORICAL NOTE (2026-07):** the "SAM is LOCAL ONLY" constraint below was written while
> SAM was an experiment. It is **superseded** β Experiment 2 shipped: SAM (`facebook/sam-vit-base`)
> is now a lazy-loaded prod dependency, pre-baked into the Docker image, behind the `SAM_RESTRICT`
> flag (ON in prod). It is imported by `restrict.py` and reached from `strategies.py`. Prod
> segmentation still runs on `mfaytin/mask2former-satellite`; SAM is a not-lawn *restriction*, not
> the segmenter. The rest of this brief is preserved as-written for provenance.
- **Do not perturb production or the byte-identical Douglas gate.** The estimation path is
guarded by `data/regression/douglas_qa.csv`: run
`python -m lawn_estimator.cli --csv data/regression/douglas_qa.csv --imagery naip` and
diff the numeric columns against `data/regression/phase35_douglas.csv`. A change that
*intentionally* shifts numbers (Experiment 1 will) is a deliberate re-baseline β gate it
behind a config flag and review, don't silently move prod output.
- **SAM work is LOCAL and EXPERIMENTAL only.** Do NOT import SAM into `pipeline.py` /
`segmentation.py` or the prod dependency set / Docker image. Use `notebooks/` or a scratch
script and a separate optional dependency. Prod stays on `mfaytin/mask2former-satellite`.
- **Branch/PR flow:** work on a feature branch β PR to `dev` (auto-deploys the staging
Space) β PR `dev` β `main` (auto-deploys prod). CI (ruff + pytest) must be green. Never
push to the `space` remote by hand. Details in the `lawn-estimator-cicd` memory.
- **Env:** always `"c:/Users/sergi/miniconda3/Scripts/conda.exe" run -p ./conda-env
--no-capture-output python β¦` (invoking `conda-env/python.exe` directly crashes numpy's
LAPACK with no traceback). `conda run` cannot execute `python -c` with newlines β write a
temp `.py` file.
- **NAIP β prod.** The QA baseline uses NAIP; production uses **Google** (`imagery=auto` +
billed key), which reads ~10β15% lower. Evaluate accuracy on Google imagery, never the
NAIP QA CSV.
## Evaluation β the owner judges, visually
**The partner's square footage is a loose reference, NOT ground truth.** He under-measures
(reconciliation showed our lawn ~1.27Γ his recorded sqft), so do not tune to match
`data/lawn_care_schedule.csv` β treat it only as a rough sanity check.
**The real check is visual, and the OWNER makes the call.** Generate the 4-panel
visualization for each test address (panel 4 = the measured lawn, cyan outline) and surface
the images to the owner so he can verify accuracy himself β the same "View" experience as
the live app.
- Produce the PNGs by running the pipeline: `python -m lawn_estimator.cli --csv <addresses>
--imagery google` (or `--address "β¦"`) β PNGs land in `data/outputs/google/`; panel 4 is
the measured lawn.
- **Show them where the owner can actually see them.** He is often away / on mobile, so the
editor side panel isn't enough β publish a **shareable page that embeds the PNGs** (data-URI
`<img>`s, openable on a phone) or point him to the staging/live app to run the addresses and
hit **View**. For a geometry change (Experiment 1) show **before/after** of the same
properties so the effect is obvious.
- `data/regression/phase35_douglas.csv` is the NAIP byte-identical gate (regression only, not
an accuracy reference).
## Experiment 1 β Carve the sidewalk out of the right-of-way
**Idea:** public sidewalks sit at a fairly standard position/width inside the ROW (commonly
a ~4β5 ft walk set back a foot or two from the property line). The ROW extension currently
adds the full `ROW_BUFFER_FT` (12 ft) strip as candidate lawn on street-facing edges; a
standard sidewalk band within that strip is hardscape, not turf.
**Where:** `geometry.py` β `parallel_offset_zone` builds the outward quads on street-facing
edges (within `STREET_MAX_DIST_M`); `build_estimation_geometry` assembles the estimation
geometry + the per-frontage `extensions`. The sidewalk band is a parallel sub-strip to
subtract from the ROW quad (or to exclude from the lawn mask within the ROW zone).
**Notes / risks:** offsets/widths vary; a fixed assumption is an approximation β validate
against the viz. The RGB mask already excludes *some* pavement, so measure the
**incremental** effect and don't double-subtract. This touches estimation geometry, so it
WILL move Douglas QA numbers β treat as a deliberate, reviewed re-baseline, ideally behind
a config flag (off by default) until proven. Highest-value, most tractable, pure geometry
β **start here.**
## Experiment 2 β SAM vs mask2former
**Idea:** Segment Anything (SAM/SAM2) is class-agnostic and may give tighter turf/hardscape
edges than `mfaytin/mask2former-satellite`. The roadmap frames it as boundary *refinement*
(prompt SAM with our existing lawn mask), not a wholesale replacement.
**Where (reference only β do NOT modify for prod):** `segmentation.py`
(`run_lawn_area_model`, `_run_segmentation`, `MODEL_LAWN_CLASS_IDS`).
**How:** local notebook/script; install SAM in a separate/optional env (likely needs GPU).
Run SAM on the same Google tiles for the partner addresses; compare resulting lawn
area/edges to mask2former, against the viz and his sqft. Decide whether SAM as a refinement
pass earns its weight. If it graduates, it returns as a proposal β never a silent import
into prod.
**Findings (2026-07 β tested locally on 10 Douglas lots through the real LiDAR+RGB math):**
- SAM runs on CPU with **no new dependency** (`transformers` ships `SamModel`). Naive
point-prompt refinement *expands* the mask (wrong direction). "Segment-and-classify" (keep SAM
pieces overlapping the mask2former lawn) drops real turf β **82% of what pure-SAM removes is
grass** across the sample.
- The only safe design is the **hybrid**: mask2former for coverage, SAM subtracts only its
*confirmed hardscape* pieces. Effect is small (**~β2%**) because the LiDAR path already ignores
roof bleed (roofs return no ground points) and the sidewalk carve (Exp 1) handles the ROW.
- A SAM-first **per-segment classifier** (transparent feature vote per piece) *agrees* with
mask2former within **Β±6%** on confident pieces; the only disagreement is shadowed/ambiguous
strips, which no signal we have resolves β **LiDAR intensity is a coin flip (AUC 0.51)**.
- **Net:** no free accuracy win β the billed number is LiDAR-dominated and every segmentation
swap agrees or worsens. See the `lawn-estimator-accuracy-exp2` memory + the review pages.
**The promising direction β SAM as a hardscape *restriction* mask (not a lawn finder):** SAM
segments house and concrete *beautifully* (crisp boundaries) even without a single "building"
class, and both LiDAR and mask2former sometimes eat into the roof. Use SAM's precise
house/driveway/patio segments to **restrict the whole pipeline** β exclude any LiDAR ground point
(and mask pixel) that falls inside them β to stop that roof/edge bleed. This is the near-term
lever, and it also makes SAM the natural **auto-labeler** for Experiment 4.
## Experiment 3 β Season / imagery vintage
**Idea:** current Google imagery can be leaf-off with dormant/dry grass. The color baseline
`run_vegetation_color_threshold` (`segmentation.py`) keys on green (hue + excess-green), so
dry turf fails it, and the model may weaken on dormant grass.
**Where:** imagery adapters `sources/imagery.py` (NAIP, county orthos) + `sources/google.py`;
`_fetch_analysis_image` in `pipeline.py` selects the source; the model + color baseline in
`segmentation.py`. Roadmap "Later" already flags "Winter/leaf-off imagery."
**How (measurement-only unless a change proves out):** compare sources/seasons for the same
addresses β Google (current, maybe dormant) vs NAIP (leaf-on ~2022) β and measure how
`lawn_sqft` and the masks shift. Quantify what dry-grass season costs and whether a source
preference or threshold/model tweak helps.
**Finding (2026-07): premise NOT reproducible β CLOSED.** Current Google *analysis* imagery is
leaf-ON; the dormant-looking tiles were the county-ortho **DISPLAY** layer (never analyzed). An
initial "signal" was an artifact of that conflation plus a hardscape-confounded greenness proxy.
NAIP is out (resolution too low). No code. Revisit only if Google actually serves leaf-off
imagery. See `lawn-estimator-accuracy-exp3`.
## Experiment 4 β Train our own segmentation model (LATER β the real ceiling-raise)
**Why:** every off-the-shelf swap (mask2former variants, SAM, per-segment) either agrees with the
current number or worsens it β the billed number is LiDAR-dominated and the disputed area is small
+ genuinely ambiguous (shadowed strips, flat hardscape). The only way to move the ceiling is a
model trained on *our own* metro imagery with *our* classes (house / driveway / sidewalk /
deck-patio / lawn / canopy). This is the roadmap's "Fine-tuned segmentation" item, promoted to a
first-class experiment.
**The bottleneck is labeling, not architecture β so this is about AUTO-LABELING / data-engine
techniques, not hand-labeling:**
- **SAM as the auto-annotator** (Meta's Segment Anything **data engine** β how SA-1B was built:
model-assisted masks, human-in-the-loop on the hard cases, progressing toward fully automatic).
We've shown SAM segments house/concrete crisply β it proposes object masks we auto-classify
(mask2former vote + color + LiDAR ground/roof) into training labels.
- **Karpathy-style "data engine" flywheel** (Tesla Autopilot): model proposes labels β focus human
review on the *uncertain / failure* cases (our shadowed strips) β retrain β repeat. Active
learning, not bulk labeling.
- **Free supervision we already have:** LiDAR classification (building=6 vs ground=2) auto-labels
roof vs at-grade; county building footprints (Microsoft/Google Open Buildings β a roadmap item)
auto-label the house; parcel β footprint β SAM-hardscape β lawn candidate. Combine these weak/auto
labels to bootstrap a training set at near-zero labeling cost.
**Model + validation:** fine-tune mask2former or train a DeepLabV3+/UNet on the auto-labeled metro
set; validate against LiDAR + owner review before it ever replaces `mfaytin/mask2former-satellite`.
**Constraints:** training needs a GPU (this box is CPU-only) β cloud GPU, out of band. Strictly
experimental; prod stays on the current model until a validated replacement earns it. Method
references to evaluate: **Karpathy's auto-labeling / data-engine ("auto-research")** and **Meta's
AutoData / SAM data-engine** work.
## Experiment 5 β Green reclaim (turf the model mislabels as water/grass)
**Trigger (2026-07-16, owner report):** 8571 Young St, Omaha 68122 β a mostly-lawn
new-construction lot measured **1,215 sqft (16.4% of the estimation area)** while the
color baseline alone saw 4,399 sqft (59.5%). Obvious under-count, the opposite of the
over-counts above.
**Root cause (debug_hardscape.py + pinned-revision check β identical either way):** the
model splits **large uniform flat turf** into classes we don't count. On the Young St tile,
class 6 (published "water") covered **44% of the tile β the entire contiguous backyard
lawn β and 98.7% of those pixels pass the green color threshold**. The empirical `[1, 4]`
lawn mapping was derived on established, tree-heavy lots where class 1 dominates; on
treeless uniform turf (new construction, big open backyards) the model reads the texture
as water/grass, the mask excludes it, the LiDAR ground points there are dropped, and the
estimate comes out severalΓ low.
**Tile survey (23 addresses: 5 QA + 18 schedule, Google imagery):**
- Class 6 is **absent (0%) on 20/23 tiles**. Where present it is either turf
(Young St 44%, 657 J E George Blvd 3.3% @ 98.8% green, 17531 Madison St 2.4% @ 97%
green) or **actual water that fails the green gate** (708 Kountze Memorial Dr,
Bellevue lakefront: class 6 only 3.6% green β correctly not reclaimed).
- Class 2 (published "grass") is a mixed street-edge band on every tile (7β32%) but only
0β8% green β **blanket-adding class 2 would wreck good tiles**; the green gate keeps
its reclaim to +0.2β1.8% of a tile.
**Fix (SHIPPED to branch `feat/green-reclaim`, flag `GREEN_RECLAIM`, off by default =
byte-identical):** `reclaim_green_turf` in `segmentation.py` β pixels the model assigned
to `GREEN_RECLAIM_CLASS_IDS` ([2, 6]) that also pass `run_vegetation_color_threshold`
are added back to the lawn mask before it gates the LiDAR points. Reuses the veg mask the
pipeline already computes (no extra pass). Zero-network unit tests in
`tests/test_segmentation_reclaim.py`.
**Limitation:** the green gate can't recover *dormant/dry* mislabeled turf (brown grass
fails the color threshold β Young St's unestablished front strips stay uncounted). That
remaining gap is Experiment 4 territory.
## Status / order (2026-07)
1. **Experiment 1 (sidewalk)** β **RETIRED (FAILED, owner verdict 2026-07).** Built behind
`SIDEWALK_CARVE` on `feat/sidewalk-carve` (β4%, byte-identical off) but the fixed-geometry ROW
sub-band carve did not hold up and will NOT ship. Superseded by the SAM/labeler sidewalk-road
line-projection direction and the leaf-on/leaf-off fusion carve (empirical, not a fixed band).
2. **Experiment 3 (season)** β **CLOSED**: premise not reproducible (analysis imagery is leaf-on).
3. **Experiment 2 (SAM)** β **EXPLORED**: no free accuracy win (LiDAR-dominated). Safe use is the
hybrid / hardscape-restriction mask (~β2%); real value is interpretability + auto-labeling for #4.
4. **Experiment 4 (train our own model)** β IN PROGRESS: Phase 1 bake-off + the
shipped cascade (`LAWN_CASCADE`) live in `docs/model-upgrade-plan.md` (the
plan of record for Phases 2β3, incl. per-step licensing); research facts in
`docs/phase2-findings.md`. Exp 5 (green reclaim, below) shipped as v0.2.
|