File size: 15,337 Bytes
cebc4ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7ee8be4
 
 
 
 
 
 
cebc4ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22e96c3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cebc4ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d33a76
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cebc4ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d33a76
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
466c0c5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d33a76
 
84e2326
 
 
 
7d33a76
 
 
31ebc97
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
# Accuracy Experiments

The estimation is accurate enough to ship (and now drives pricing), but the partner's
visual review surfaced a recurring **over-count** that inflates prices, plus a
seasonal-imagery weakness. This brief scopes three experiments to close that gap. Read
`CLAUDE.md` (architecture) and `ROADMAP.md` first.

## Why this matters

Our lawn estimate (`lawn_sqft`) is the pricing basis (Phase 6, live). The partner
reviewed production (Google-imagery) measurements against his own job sheet and confirmed
they're solid overall but **over-count** in three recurring ways, each of which pushes the
price up:

- **Sidewalk in the right-of-way** counted as lawn.
- **Backyard concrete** (patios/pads) counted as lawn.
- **Roof edges** bleeding into the lawn mask.

He also flagged a season issue: current Google imagery is sometimes **leaf-off (bare
trees) with dormant, dry grass**, which the green-keyed color baseline and the
segmentation model read poorly.

Reconciliation context (see the `lawn-estimator-pricing` memory): our measured lawn runs a
median ~1.27Γ— his recorded sqft, so pricing off it is ~+22% vs his historical charges. He
accepts that (he under-measured), but tightening the over-count directly improves price
accuracy β€” so accuracy is the current highest-value work.

## Hard constraints β€” read before touching anything

> **HISTORICAL NOTE (2026-07):** the "SAM is LOCAL ONLY" constraint below was written while
> SAM was an experiment. It is **superseded** β€” Experiment 2 shipped: SAM (`facebook/sam-vit-base`)
> is now a lazy-loaded prod dependency, pre-baked into the Docker image, behind the `SAM_RESTRICT`
> flag (ON in prod). It is imported by `restrict.py` and reached from `strategies.py`. Prod
> segmentation still runs on `mfaytin/mask2former-satellite`; SAM is a not-lawn *restriction*, not
> the segmenter. The rest of this brief is preserved as-written for provenance.

- **Do not perturb production or the byte-identical Douglas gate.** The estimation path is
  guarded by `data/regression/douglas_qa.csv`: run
  `python -m lawn_estimator.cli --csv data/regression/douglas_qa.csv --imagery naip` and
  diff the numeric columns against `data/regression/phase35_douglas.csv`. A change that
  *intentionally* shifts numbers (Experiment 1 will) is a deliberate re-baseline β€” gate it
  behind a config flag and review, don't silently move prod output.
- **SAM work is LOCAL and EXPERIMENTAL only.** Do NOT import SAM into `pipeline.py` /
  `segmentation.py` or the prod dependency set / Docker image. Use `notebooks/` or a scratch
  script and a separate optional dependency. Prod stays on `mfaytin/mask2former-satellite`.
- **Branch/PR flow:** work on a feature branch β†’ PR to `dev` (auto-deploys the staging
  Space) β†’ PR `dev` β†’ `main` (auto-deploys prod). CI (ruff + pytest) must be green. Never
  push to the `space` remote by hand. Details in the `lawn-estimator-cicd` memory.
- **Env:** always `"c:/Users/sergi/miniconda3/Scripts/conda.exe" run -p ./conda-env
  --no-capture-output python …` (invoking `conda-env/python.exe` directly crashes numpy's
  LAPACK with no traceback). `conda run` cannot execute `python -c` with newlines β€” write a
  temp `.py` file.
- **NAIP β‰  prod.** The QA baseline uses NAIP; production uses **Google** (`imagery=auto` +
  billed key), which reads ~10–15% lower. Evaluate accuracy on Google imagery, never the
  NAIP QA CSV.

## Evaluation β€” the owner judges, visually

**The partner's square footage is a loose reference, NOT ground truth.** He under-measures
(reconciliation showed our lawn ~1.27Γ— his recorded sqft), so do not tune to match
`data/lawn_care_schedule.csv` β€” treat it only as a rough sanity check.

**The real check is visual, and the OWNER makes the call.** Generate the 4-panel
visualization for each test address (panel 4 = the measured lawn, cyan outline) and surface
the images to the owner so he can verify accuracy himself β€” the same "View" experience as
the live app.

- Produce the PNGs by running the pipeline: `python -m lawn_estimator.cli --csv <addresses>
  --imagery google` (or `--address "…"`) β†’ PNGs land in `data/outputs/google/`; panel 4 is
  the measured lawn.
- **Show them where the owner can actually see them.** He is often away / on mobile, so the
  editor side panel isn't enough β€” publish a **shareable page that embeds the PNGs** (data-URI
  `<img>`s, openable on a phone) or point him to the staging/live app to run the addresses and
  hit **View**. For a geometry change (Experiment 1) show **before/after** of the same
  properties so the effect is obvious.
- `data/regression/phase35_douglas.csv` is the NAIP byte-identical gate (regression only, not
  an accuracy reference).

## Experiment 1 β€” Carve the sidewalk out of the right-of-way

**Idea:** public sidewalks sit at a fairly standard position/width inside the ROW (commonly
a ~4–5 ft walk set back a foot or two from the property line). The ROW extension currently
adds the full `ROW_BUFFER_FT` (12 ft) strip as candidate lawn on street-facing edges; a
standard sidewalk band within that strip is hardscape, not turf.

**Where:** `geometry.py` β€” `parallel_offset_zone` builds the outward quads on street-facing
edges (within `STREET_MAX_DIST_M`); `build_estimation_geometry` assembles the estimation
geometry + the per-frontage `extensions`. The sidewalk band is a parallel sub-strip to
subtract from the ROW quad (or to exclude from the lawn mask within the ROW zone).

**Notes / risks:** offsets/widths vary; a fixed assumption is an approximation β€” validate
against the viz. The RGB mask already excludes *some* pavement, so measure the
**incremental** effect and don't double-subtract. This touches estimation geometry, so it
WILL move Douglas QA numbers β†’ treat as a deliberate, reviewed re-baseline, ideally behind
a config flag (off by default) until proven. Highest-value, most tractable, pure geometry
β€” **start here.**

## Experiment 2 β€” SAM vs mask2former

**Idea:** Segment Anything (SAM/SAM2) is class-agnostic and may give tighter turf/hardscape
edges than `mfaytin/mask2former-satellite`. The roadmap frames it as boundary *refinement*
(prompt SAM with our existing lawn mask), not a wholesale replacement.

**Where (reference only β€” do NOT modify for prod):** `segmentation.py`
(`run_lawn_area_model`, `_run_segmentation`, `MODEL_LAWN_CLASS_IDS`).

**How:** local notebook/script; install SAM in a separate/optional env (likely needs GPU).
Run SAM on the same Google tiles for the partner addresses; compare resulting lawn
area/edges to mask2former, against the viz and his sqft. Decide whether SAM as a refinement
pass earns its weight. If it graduates, it returns as a proposal β€” never a silent import
into prod.

**Findings (2026-07 β€” tested locally on 10 Douglas lots through the real LiDAR+RGB math):**
- SAM runs on CPU with **no new dependency** (`transformers` ships `SamModel`). Naive
  point-prompt refinement *expands* the mask (wrong direction). "Segment-and-classify" (keep SAM
  pieces overlapping the mask2former lawn) drops real turf β€” **82% of what pure-SAM removes is
  grass** across the sample.
- The only safe design is the **hybrid**: mask2former for coverage, SAM subtracts only its
  *confirmed hardscape* pieces. Effect is small (**~βˆ’2%**) because the LiDAR path already ignores
  roof bleed (roofs return no ground points) and the sidewalk carve (Exp 1) handles the ROW.
- A SAM-first **per-segment classifier** (transparent feature vote per piece) *agrees* with
  mask2former within **Β±6%** on confident pieces; the only disagreement is shadowed/ambiguous
  strips, which no signal we have resolves β€” **LiDAR intensity is a coin flip (AUC 0.51)**.
- **Net:** no free accuracy win β€” the billed number is LiDAR-dominated and every segmentation
  swap agrees or worsens. See the `lawn-estimator-accuracy-exp2` memory + the review pages.

**The promising direction β€” SAM as a hardscape *restriction* mask (not a lawn finder):** SAM
segments house and concrete *beautifully* (crisp boundaries) even without a single "building"
class, and both LiDAR and mask2former sometimes eat into the roof. Use SAM's precise
house/driveway/patio segments to **restrict the whole pipeline** β€” exclude any LiDAR ground point
(and mask pixel) that falls inside them β€” to stop that roof/edge bleed. This is the near-term
lever, and it also makes SAM the natural **auto-labeler** for Experiment 4.

## Experiment 3 β€” Season / imagery vintage

**Idea:** current Google imagery can be leaf-off with dormant/dry grass. The color baseline
`run_vegetation_color_threshold` (`segmentation.py`) keys on green (hue + excess-green), so
dry turf fails it, and the model may weaken on dormant grass.

**Where:** imagery adapters `sources/imagery.py` (NAIP, county orthos) + `sources/google.py`;
`_fetch_analysis_image` in `pipeline.py` selects the source; the model + color baseline in
`segmentation.py`. Roadmap "Later" already flags "Winter/leaf-off imagery."

**How (measurement-only unless a change proves out):** compare sources/seasons for the same
addresses β€” Google (current, maybe dormant) vs NAIP (leaf-on ~2022) β€” and measure how
`lawn_sqft` and the masks shift. Quantify what dry-grass season costs and whether a source
preference or threshold/model tweak helps.

**Finding (2026-07): premise NOT reproducible β€” CLOSED.** Current Google *analysis* imagery is
leaf-ON; the dormant-looking tiles were the county-ortho **DISPLAY** layer (never analyzed). An
initial "signal" was an artifact of that conflation plus a hardscape-confounded greenness proxy.
NAIP is out (resolution too low). No code. Revisit only if Google actually serves leaf-off
imagery. See `lawn-estimator-accuracy-exp3`.

## Experiment 4 β€” Train our own segmentation model (LATER β€” the real ceiling-raise)

**Why:** every off-the-shelf swap (mask2former variants, SAM, per-segment) either agrees with the
current number or worsens it β€” the billed number is LiDAR-dominated and the disputed area is small
+ genuinely ambiguous (shadowed strips, flat hardscape). The only way to move the ceiling is a
model trained on *our own* metro imagery with *our* classes (house / driveway / sidewalk /
deck-patio / lawn / canopy). This is the roadmap's "Fine-tuned segmentation" item, promoted to a
first-class experiment.

**The bottleneck is labeling, not architecture β€” so this is about AUTO-LABELING / data-engine
techniques, not hand-labeling:**
- **SAM as the auto-annotator** (Meta's Segment Anything **data engine** β€” how SA-1B was built:
  model-assisted masks, human-in-the-loop on the hard cases, progressing toward fully automatic).
  We've shown SAM segments house/concrete crisply β†’ it proposes object masks we auto-classify
  (mask2former vote + color + LiDAR ground/roof) into training labels.
- **Karpathy-style "data engine" flywheel** (Tesla Autopilot): model proposes labels β†’ focus human
  review on the *uncertain / failure* cases (our shadowed strips) β†’ retrain β†’ repeat. Active
  learning, not bulk labeling.
- **Free supervision we already have:** LiDAR classification (building=6 vs ground=2) auto-labels
  roof vs at-grade; county building footprints (Microsoft/Google Open Buildings β€” a roadmap item)
  auto-label the house; parcel βˆ’ footprint βˆ’ SAM-hardscape β‰ˆ lawn candidate. Combine these weak/auto
  labels to bootstrap a training set at near-zero labeling cost.

**Model + validation:** fine-tune mask2former or train a DeepLabV3+/UNet on the auto-labeled metro
set; validate against LiDAR + owner review before it ever replaces `mfaytin/mask2former-satellite`.

**Constraints:** training needs a GPU (this box is CPU-only) β†’ cloud GPU, out of band. Strictly
experimental; prod stays on the current model until a validated replacement earns it. Method
references to evaluate: **Karpathy's auto-labeling / data-engine ("auto-research")** and **Meta's
AutoData / SAM data-engine** work.

## Experiment 5 β€” Green reclaim (turf the model mislabels as water/grass)

**Trigger (2026-07-16, owner report):** 8571 Young St, Omaha 68122 β€” a mostly-lawn
new-construction lot measured **1,215 sqft (16.4% of the estimation area)** while the
color baseline alone saw 4,399 sqft (59.5%). Obvious under-count, the opposite of the
over-counts above.

**Root cause (debug_hardscape.py + pinned-revision check β€” identical either way):** the
model splits **large uniform flat turf** into classes we don't count. On the Young St tile,
class 6 (published "water") covered **44% of the tile β€” the entire contiguous backyard
lawn β€” and 98.7% of those pixels pass the green color threshold**. The empirical `[1, 4]`
lawn mapping was derived on established, tree-heavy lots where class 1 dominates; on
treeless uniform turf (new construction, big open backyards) the model reads the texture
as water/grass, the mask excludes it, the LiDAR ground points there are dropped, and the
estimate comes out severalΓ— low.

**Tile survey (23 addresses: 5 QA + 18 schedule, Google imagery):**
- Class 6 is **absent (0%) on 20/23 tiles**. Where present it is either turf
  (Young St 44%, 657 J E George Blvd 3.3% @ 98.8% green, 17531 Madison St 2.4% @ 97%
  green) or **actual water that fails the green gate** (708 Kountze Memorial Dr,
  Bellevue lakefront: class 6 only 3.6% green β†’ correctly not reclaimed).
- Class 2 (published "grass") is a mixed street-edge band on every tile (7–32%) but only
  0–8% green β€” **blanket-adding class 2 would wreck good tiles**; the green gate keeps
  its reclaim to +0.2–1.8% of a tile.

**Fix (SHIPPED to branch `feat/green-reclaim`, flag `GREEN_RECLAIM`, off by default =
byte-identical):** `reclaim_green_turf` in `segmentation.py` β€” pixels the model assigned
to `GREEN_RECLAIM_CLASS_IDS` ([2, 6]) that also pass `run_vegetation_color_threshold`
are added back to the lawn mask before it gates the LiDAR points. Reuses the veg mask the
pipeline already computes (no extra pass). Zero-network unit tests in
`tests/test_segmentation_reclaim.py`.

**Limitation:** the green gate can't recover *dormant/dry* mislabeled turf (brown grass
fails the color threshold β€” Young St's unestablished front strips stay uncounted). That
remaining gap is Experiment 4 territory.

## Status / order (2026-07)

1. **Experiment 1 (sidewalk)** β€” **RETIRED (FAILED, owner verdict 2026-07).** Built behind
   `SIDEWALK_CARVE` on `feat/sidewalk-carve` (βˆ’4%, byte-identical off) but the fixed-geometry ROW
   sub-band carve did not hold up and will NOT ship. Superseded by the SAM/labeler sidewalk-road
   line-projection direction and the leaf-on/leaf-off fusion carve (empirical, not a fixed band).
2. **Experiment 3 (season)** β€” **CLOSED**: premise not reproducible (analysis imagery is leaf-on).
3. **Experiment 2 (SAM)** β€” **EXPLORED**: no free accuracy win (LiDAR-dominated). Safe use is the
   hybrid / hardscape-restriction mask (~βˆ’2%); real value is interpretability + auto-labeling for #4.
4. **Experiment 4 (train our own model)** β€” IN PROGRESS: Phase 1 bake-off + the
   shipped cascade (`LAWN_CASCADE`) live in `docs/model-upgrade-plan.md` (the
   plan of record for Phases 2–3, incl. per-step licensing); research facts in
   `docs/phase2-findings.md`. Exp 5 (green reclaim, below) shipped as v0.2.