| # DECISIONS.md |
|
|
| Running log of implementation decisions where the spec (`PROJECT_SPEC_dog_reunification_V1.md`) |
| left room, plus anything discovered that contradicts the spec. Newest entries at the bottom of |
| each section. |
|
|
| ## Implementer choices (from spec §19) |
|
|
| - **Frontend styling: Tailwind CSS.** Faster to build responsive/mobile-first UI; utility classes |
| keep styling co-located with components. (spec §5, §19) |
| - **Auth: JWT access tokens** (Bearer), password hashing via `passlib[bcrypt]`. Stateless, simple to |
| consume from the SPA, and matches the OpenAPI/codegen story. Refresh tokens deferred. (spec §11, §19) |
| - **Dog-detector crop: planned, never built.** The intent was an optional crop-to-dog pass before |
| embedding. It was not needed once the embedder was fine-tuned on full-body photos as well as face |
| crops, so no detector module exists and `Picture.is_probably_not_dog` is never set. (spec §8, §19) |
|
|
| ## Decisions where the spec left room |
|
|
| - **Python 3.13** is used (local interpreter); spec asked for 3.11+. No incompatibilities expected. |
| - **Finders without accounts do not create a `users` row.** Their contact lives on the `cases` row |
| (`finder_name/finder_email/finder_phone`), per spec §7.1 note. De-dupe into lightweight finder users |
| deferred. (spec §7.1) |
| - **Embeddings stored in their own `embeddings` table** (per spec §7.5/§7.6 refinement), vector as |
| float32 BLOB, L2-normalized at write time so cosine == dot product. |
| - **Single config object** via `pydantic-settings`, loaded from `.env`. All thresholds/radius levels are |
| named config values with documented defaults (spec §0, §15). |
|
|
| ## Resolved spec questions |
|
|
| - `REVIEW_THRESHOLD` / `STRONG_THRESHOLD` govern the **case matcher** only. The public photo search |
| returns the top N ranked results with no cutoff, deliberately: a page of low scores tells a |
| searcher their dog is probably not in the database, where an empty page is ambiguous. Retune with |
| `scripts/eval_matching.py` if the case workflow is put into real use (spec §9.3, §9.6). |
| - `RADIUS_LEVELS` default `[0, 10, 25, 50, 100, -1]` where `-1` == nationwide (no filter). |
| - Nearby-shelter lookup is a bundled static list for MVP (spec §12). |
|
|
| ## Estimated-breed candidate gate (extends spec §9.3 metadata gate) |
|
|
| - **New swap point: `BreedClassifier`** (`backend/app/ml/breed.py`), structured exactly like |
| `Embedder`: `Protocol` + `MockBreedClassifier` default + lazily-imported `HFBreedClassifier` + |
| `get_breed_classifier()`/`reset_breed_classifier_cache()` singletons. It produces **breed labels |
| only**, used purely as a candidate pre-filter — never for similarity (that matcher is selected |
| separately). |
| - **Predictions stored in their own versioned `breed_predictions` table** (mirrors `embeddings`): |
| per-picture, keyed by breed-model `name`/`version`, one row per top-K rank. Match time uses only |
| the *active* breed model's rows; other versions count as "no predictions" → fail-open. |
| - **Fail-open, symmetric top-K gate** (open choice #1: chose **symmetric** for recall). A candidate |
| is excluded **only on a definite conflict**: both query and candidate have predictions under the |
| active model and their top-K label sets do not intersect. Missing predictions on either side |
| never exclude — same recall-first discipline as `_metadata_compatible` and `geo.within_radius`. |
| The gate runs after the geo + color/size gates and **before** the cosine work, so it prunes the |
| expensive comparison. `BREED_FILTER_ENABLED=false` disables it entirely (predictions still |
| stored). |
| - **`BREED_TOP_K` default 7** (open choice #3; anywhere in 5–10), exposed as config to tune vs recall. |
| - **Human-entered breed is deliberately not folded in** (open choice #2: **no**, settled). |
| `KnownDog.breed` / `UnknownDog.est_breed` are free text, and owners are often wrong about mixed |
| breeds. A confident-but-wrong human label would exclude the correct dog outright, so the gate stays |
| purely model-driven to keep it recall-first. |
| - **Mock-first / HF opt-in.** Default `BREED_CLASSIFIER=mock` is deterministic and weight-free |
| (reuses the mock embedder's pixel hashing so identical images yield identical breed sets — this |
| keeps the seeded Rex match working). The real HF model |
| (`jhoppanne/Dogs-Breed-Image-Classification-V1`, ~43 MB) is enabled only via config and lazily |
| imports `torch`/`transformers`, exactly like the real embedders. Labels are read from |
| `model.config.id2label` at load time (no hardcoded breed list). |
| - **Best-effort in the upload pipeline.** Breed prediction is wrapped in try/except so an upload |
| never fails if the breed model errors or isn't downloaded; the embedding stays required. |
| - **Migrations:** `versions/` previously relied on `create_all`; added a baseline autogenerated |
| migration covering the full schema including `breed_predictions`. Existing dev DBs predate the |
| table — delete `backend/data/app.db` and re-seed, **or** run `alembic upgrade head` then |
| `python -m scripts.process_dataset --all` to populate predictions for already-uploaded photos. |
|
|
| ## Batch loading & test-data management (datasets) |
|
|
| - **New `datasets` table + nullable `dataset_id` FK** on `known_dogs`, `unknown_dogs`, `users` |
| (Alembic `31362988755b`, chained on the `a18bdb0cac22` baseline; batch-mode ALTERs with **named** |
| FK constraints since SQLite batch mode requires names). Cases/matches/pictures link transitively |
| through the dogs/users, not via their own `dataset_id`. |
| - **Purge is application-level and transactional**, not DB-cascade. Pictures link to dogs by |
| convention (`subject_type`/`subject_id`, no FK) and `embeddings`/`breed_predictions` FK pictures |
| with no `ON DELETE`, so a clean purge deletes children before parents in one transaction |
| (notifications → breed_predictions/embeddings → matches → pictures → cases → unknown_dogs → |
| known_dogs → users → dataset) and returns per-type counts. Everything-or-nothing via |
| commit/rollback. (`services/datasets.py`) |
| - **Purge also deletes media files.** The picture storage keys are collected *before* the rows are |
| deleted, then the files are removed *after* the DB commit succeeds (best-effort: a failed unlink |
| only leaves an orphan, never inconsistent data; a rolled-back purge deletes nothing). Count |
| returned as `media_files`. Earlier builds left media on disk (the storage step was a no-op), which |
| accumulated orphans across purges/re-seeds; `scripts/prune_media.py` removes any orphans not |
| referenced by a `pictures` row (dry-run by default) and stays as a safety net for manual DB resets. |
| - **One dog per folder, multiple pictures (default).** The CSV `folder` column is the identity key |
| (each DogFaceNet subfolder == one real dog). The loader **groups a folder's rows into a single |
| `KnownDog`/`UnknownDog`** and attaches every image as a `Picture` (first = primary); a known dog |
| gets one optional lost case, a found dog one found `Case`. This supersedes the original "one row == |
| one dog" behavior, which created a separate same-named dog per photo — wrong for multi-photo |
| identities and bad for matching (1-picture dogs). Pass `group_by_folder=False` / |
| `--one-dog-per-image` for the legacy per-row behavior. Owners/finders are still de-duped by email |
| and linked to the dataset; `found_zip` equals the dog's known ZIP so default-radius (level 0 = same |
| ZIP) matching works. **Note:** datasets loaded before this fix have one dog per image — purge and |
| re-load them. |
| - **Holdout default:** 1 image, or 2 when a dog has 5+ images and `--holdout` is not given; always |
| leaves ≥1 registration image (a single-image dog yields 0 holdouts, logged). |
| - **`--mark-lost` also opens a lost `Case`** per marked known dog (status `lost`, `last_known_zip` |
| set) so the known dataset is actually matchable; `--mark-lost-pct` controls the share (default 100). |
| - **Embedding reuse:** extracted `embed_picture()` in `services/images.py`, used by both the upload |
| pipeline and `embed-all` — embedding generation is never reimplemented. |
| - **Many-to-many match:** `run_matching_for_case` gained an optional `candidate_dataset_id` to scope |
| the candidate pool; the `/admin/datasets/{id}/match?candidate_dataset_id=` endpoint passes it through. |
| - **Progress feedback = polling.** No realtime channel existed, so dataset loads run in a worker |
| thread tracked by a process-local in-memory job registry (`services/jobs.py`); the dashboard polls |
| `GET /admin/jobs/{id}`. Ephemeral by design (single-process local target, A8); a durable queue is a |
| documented future swap. |
| - **Schema change ⇒ re-seed.** Existing dev DBs lack `dataset_id`; `create_all` does not ALTER, so |
| delete `backend/data/app.db` and re-seed, or `alembic upgrade head`. |
|
|
| ## Storage-only load (defer embeddings) |
|
|
| - `process_and_store_picture` gained `generate_embedding` / `generate_breed` flags (default `True`, |
| so normal uploads/seed are unchanged). The batch loader's new `skip_embeddings` option sets both |
| to `False` — images are validated, normalized (EXIF stripped), stored, and thumbnailed, but **no |
| embeddings or breed predictions are generated**. Exposed as `--skip-embeddings` (CLI), |
| `skip_embeddings` (load-job form), and a "Load images into storage only" checkbox in the |
| dashboard. Matching needs embeddings, so "Run matching after load" is force-disabled when skipping. |
| - **Batch embedding generator:** the existing `POST /admin/datasets/{id}/embed-all` (idempotent — |
| only embeds pictures lacking an embedding under the active model) is now surfaced as an **"Embed" |
| button per dataset row on the Data Loading page**, next to "Purge" (in addition to the dataset |
| detail page's "Run Embeddings"). This is the intended two-step flow: bulk-load into storage now, |
| generate embeddings later. `embed-all` runs synchronously; fine for the mock embedder. (If a real |
| CNN embedder on very large datasets makes the request too long, move `embed_all` onto the |
| background job system used by load — a documented future step.) |
|
|
| ## Real re-ID embedder: HuggingFace model (EMBEDDER=hf) |
|
|
| - **New `HFEmbedder`** (`backend/app/ml/embedder.py`) behind the existing `Embedder` swap point — |
| no API/schema changes. It loads an HF image-classification model |
| (`EMBEDDER_HF_MODEL`, default `jhoppanne/Dogs-Breed-Image-Classification-V1`) and uses its |
| **penultimate, pre-classifier pooled features** as the re-ID vector: last hidden state → |
| global-average-pool (CNN feature maps) or mean-over-tokens (transformer sequences) → L2-normalize. |
| Breed *labels* are NOT used for similarity — that stays the separate `BreedClassifier`. Embedding |
| `dim` is inferred at load with one dummy forward (architecture-agnostic). |
| - **torch/transformers lazily imported**, exactly like `HFBreedClassifier`; default |
| `EMBEDDER=mock` still needs none of it. torch is installed CPU-only. |
| - **Versioning makes the swap safe:** embeddings are tagged `model_name="hf-embed"`, |
| `model_version=<repo name>`. Matching only compares embeddings from the *active* model, so after |
| switching `EMBEDDER=mock→hf` you must regenerate vectors (admin **embed-all**); old mock vectors |
| are simply ignored, not mixed. |
| - **Same model, two roles:** this repo can use the model both as the breed gate |
| (`BREED_CLASSIFIER=hf`, labels) and the re-ID embedder (`EMBEDDER=hf`, features). They load |
| independent copies for interface cleanliness; sharing one load is a possible future optimization. |
| - **Perf note:** CPU embedding of the full DogFaceNet set (~8k images) via `embed-all` is slow and |
| runs synchronously today — fine to run once, but a background-job version of `embed-all` is the |
| documented next step for large real-model runs. |
|
|
| ## Breed filtering (HF softmax labels) |
|
|
| - **Same HF model, the classifier head this time.** `BreedClassifier=hf` runs |
| `jhoppanne/Dogs-Breed-Image-Classification-V1`'s softmax and stores the top-K labels in |
| `breed_predictions` (separate from the re-ID *embeddings*). `BREED_TOP_K=10` so the UI can filter |
| by top-1 through top-10. |
| - **Filter semantics (per-photo, recall-friendly):** a dog matches breed B at depth k if **any of |
| its photos** predicts B within rank `< k`. `breed_k=1` = "top match"; 5/10 = top-K. Displayed |
| `predicted_breeds` is the dog-level aggregate (distinct labels ordered by best rank then score |
| across the dog's photos), so a dog matched on one photo's top-1 may show a different label first — |
| expected for multi-photo dogs. |
| - **Active breed model is auto-detected from the DB** (`active_breed_model_in_db`, prefers non-mock, |
| most-used) — reading predictions never loads the HF model. `estimated_breeds_for` was refactored |
| to use this too, so match-candidate hydration no longer risks loading the breed model. |
| - **Endpoints:** `GET /admin/breeds` (distinct labels + counts) and |
| `GET /admin/dogs?breed=&breed_k=` (filter). Profiles now include `predicted_breeds`. |
| - **Tests pin `BREED_CLASSIFIER=mock`** (conftest) so the suite never loads the HF model; the `.env` |
| default is `hf` for real use. Generate predictions with `python -m scripts.process_dataset --all` |
| (`--limit N` for a partial run). |
| |
| ## Public "search by photo" (home page) |
| |
| - **What:** anonymous visitors can upload a dog photo (+ optional ZIP) on the home page and get the |
| found/unknown dogs already in the system ranked by photo similarity. `POST /search/by-photo` |
| (multipart `file`, optional `zip`, `top_k`). No auth, no DB writes. |
| - **Read-only by design:** it is a *lookup*, not a report. It never creates a `KnownDog`/`Case`, so |
| it doesn't pollute the matching pipeline or notify anyone. The copy steers users who find a likely |
| match to the real "report my dog as lost" flow (where confirmed, mediated matching happens). |
| - **Transient embedding:** `images.embed_bytes()` runs the *same* validate→normalize→encode pipeline |
| as stored uploads, writes to a `tempfile`, embeds via the active model, deletes the temp file. The |
| query photo is never persisted — same vector a stored picture would get, zero footprint. |
| - **Scoring reuses the matcher:** `matching.search_unknowns_by_vectors()` scores the active |
| found/unknown pool (`pending|lost|at_shelter`) by the same dog-level **max cosine** aggregate the |
| case matcher uses (`_dog_level_score`). No `REVIEW_THRESHOLD` cut here — this is a "show me the |
| closest" ranked list, so we return the top-K regardless of absolute score (recall-first; the model |
| eval showed scores run uniformly high, so a threshold would mostly mislead). |
| - **ZIP is optional + fail-open:** with a ZIP we scope to `DEFAULT_SEARCH_RADIUS=100` mi (fail-open |
| on unknown ZIPs, like the case matcher) and surface per-result `distance_miles`; empty ZIP = |
| nationwide. Results are always ranked by score, not distance. |
| - **Results carry their own photos** so the UI viewer (`PhotoTools` modal) shows every angle |
| without a separate auth'd detail endpoint (the admin `/admin/dog/...` route stays admin-only). |
| - **Tests:** `backend/tests/test_search.py` (ranking, no-auth, persists-nothing, ZIP distance, bad |
| image) and `frontend/src/components/__tests__/PhotoTools.test.tsx` (upload→render, photo viewer). |
|
|
| ## Public breed-estimation toy (home page) |
|
|
| - **What:** a for-fun home-page widget — upload a photo, get the top N (3/5/7/10) breed labels. |
| `POST /search/breed` (multipart `file`, `top_n`). No auth, no DB writes. |
| - **Reuses the breed swap point, not a new model:** `images.predict_breeds_bytes()` mirrors |
| `embed_bytes` (same validate→normalize→encode→tempfile→cleanup) and calls the active |
| `BreedClassifier` (`get_breed_classifier`). With `BREED_CLASSIFIER=hf` it's the real softmax; with |
| `mock` (tests/offline) it's deterministic pseudo-labels. The endpoint slices to `top_n`, so the |
| practical max is bounded by `BREED_TOP_K` (10 in `.env`). |
| - **Distinct from the match-time breed gate:** this is display/entertainment only and never touches |
| matching; it just exposes the classifier the gate already uses. |
| - **Note:** the home page presents this together with photo-search in a single `PhotoTools` widget — |
| one upload, a large centered preview, and buttons to run matching, breed estimation, or both. |
| - **Tests:** `test_search.py` (top_n length, deterministic labels, clamp to 10, bad image) + |
| `frontend/src/components/__tests__/PhotoTools.test.tsx` (upload→ranked list, run-both). |
| |
| ## ZIP filtering fix (geo centroids were missing) |
| |
| - **Bug:** photo-search ZIP filtering (and case-match radius) silently did nothing. Two causes: |
| (1) the app default `zip_centroid_file` was CWD-relative (`./data/...`), so launching uvicorn from |
| `backend/` looked in `backend/data/` and loaded **0 centroids**; (2) the shipped centroid CSV was a |
| 31-row national demo set, while the DogFaceNet demo data uses ~480 real Texas ZIPs — **0** of which |
| had centroids. `within_radius` fails open on unknown ZIPs, so every candidate passed → ZIP no-op, |
| distance always null. |
| - **Fix 1 (path):** `config.zip_centroid_file` now defaults to an **absolute** path resolved from the |
| repo root (`parents[2]/data/zip_centroids.csv`), so it's found regardless of CWD. `.env` / |
| `ZIP_CENTROID_FILE` still overrides. |
| - **Fix 2 (coverage):** a one-off generator produced `data/zip_centroids.csv` (539 |
| rows) covering the national demo ZIPs (kept verbatim so matching tests' distances are unchanged) |
| plus every Texas ZIP in `prepare_dogfacenet.ZIP_CODES`, mapped to a **metro-level** centroid by |
| ZIP3 prefix. Inter-metro distances are realistic (Houston↔Dallas ≈ 225 mi) so radius filtering |
| works; intra-metro distance ≈ 0. A real per-ZIP gazetteer is a drop-in CSV replacement later. |
| - **Verified:** 1395/1395 demo dogs now have centroids; a Houston search at 100 mi keeps 382/1395 |
| (drops other metros). Test: `test_search_by_photo_zip_excludes_out_of_radius`. |
|
|
| ## One model pass per stored image (embedding + breed together) |
|
|
| - **Problem:** every stored picture forwarded through the model TWICE — once in `HFEmbedder.embed` |
| (penultimate features) and once in `HFBreedClassifier.predict` (logits) — even though both use the |
| *same* model (`jhoppanne/Dogs-Breed-Image-Classification-V1`). Batch workflows were worse: running |
| `embed_dataset` then `backfill_breeds` = two full passes over the dataset. |
| - **Fix:** `HFEmbedder.embed_and_breed(paths, top_k)` does ONE forward with `output_hidden_states=True` |
| and returns both the pooled re-ID vector (from `hidden_states[-1]`) and the softmax breed labels |
| (from `logits`). Verified the vector is **identical** to the old standalone `embed()` — existing |
| embeddings stay valid, no regeneration needed. |
| - **Orchestration:** `images.embed_and_breed_picture(db, pic)` uses the single pass when |
| `_same_hf_model()` (embedder=hf, breed=hf, same repo id); otherwise falls back to the two separate |
| models (mock defaults / mismatched ids), embedding required + breed best-effort. Breed rows are |
| tagged via `hf_breed_name_version()` so they match `HFBreedClassifier`'s tags exactly. |
| - **Wired in:** `process_and_store_picture` (live uploads) calls it when both are requested; |
| `datasets.embed_all` (admin "Run Embeddings + Breeds" button) now reports `embedded` + `breeds`; |
| and `scripts/process_dataset.py` is the batch path that **replaces** `embed_dataset` + |
| `backfill_breeds` (one pass instead of two). `embed_dataset` remains for mock / split configs; |
| `backfill_breeds` has since been removed. |
| - **Scope:** applies to pictures stored *in the system*. The public front-page toys (`embed_bytes`, |
| `predict_breeds_bytes`) are intentionally separate one-off queries and out of scope. |
| - **Tests:** `test_embed_breed_single_pass.py` — a fake same-HF embedder proves the model runs |
| exactly once and both rows are written with the right tags; plus the mock-fallback idempotency. |
|
|
| ## Purge fix: datasets that share users (reused test emails) |
|
|
| - **Bug:** purging "DogFaceNet known" failed with `FOREIGN KEY constraint failed` on `DELETE FROM |
| users` and rolled the whole transaction back (dataset left fully intact). Root cause: the prepare |
| scripts mint owner emails `owner1@test.com…`, and `batch_loader._get_or_create_user` dedupes users |
| by email — so a second dataset (MPDD) whose CSV reused those emails attached its 96 dogs to the |
| FIRST dataset's users. Purging the first dataset then tried to delete users still owning the |
| second dataset's dogs. |
| - **Fix (`purge_dataset`):** a user may be legitimately shared across datasets, so instead of force- |
| deleting all of the dataset's users, only delete users with **no surviving references** |
| (`KnownDog.owner_id` / `Case.person_id` after the dataset's own rows are gone) and **detach the |
| rest** (`dataset_id = NULL`) so the dataset row can be removed without a dangling FK. New |
| `users_detached` count is returned. Also extended match cleanup to remove matches whose |
| **candidate** is one of the purged dataset's dogs (not just by case), so a cross-dataset match |
| can't be left pointing at a deleted dog. |
| - **Verified live:** purged DogFaceNet known (1393 dogs, 5969 pictures + media, 11938 embeddings, |
| 59690 breed preds; 1297 users deleted, 96 detached). MPDD (#3/#4) fully intact, all owners valid, |
| 0 media orphans. "DogFaceNet unknown" purge was checked and had cleaned up correctly (no DB or |
| media orphans). |
| - **Latent data-hygiene note:** reused test emails still ENTANGLE datasets (one user owning dogs in |
| several). The purge now tolerates it, but future prep runs (YT-BB-Dog, Sibetan) will keep sharing |
| `owner1@test.com…` unless the prepare scripts namespace emails per dataset. Recommended follow-up. |
| - **Test:** `test_purge_retains_user_shared_with_another_dataset` (shared owner across two datasets). |
| |
| ## Removed color/size as entered information (subjective/unreliable) |
| |
| - **Why:** color and size are subjective, inconsistently reported, and were pruning candidates in the |
| matcher — a wrong color/size guess could hide a correct match. |
| - **Matching:** deleted the `_metadata_compatible` color/size gate and the `apply_metadata_filter` |
| path in `run_matching_for_case`. Matching now gates only on ZIP radius + estimated breed (model |
| softmax), then ranks by embedding similarity. (Breed filtering was explicitly wanted; kept.) |
| - **Inputs removed:** `KnownDogCreate`/`KnownDogUpdate`, `LostCaseCreate.new_dog_color/size`, |
| `FoundSightedCreate`, and the `/cases/found` + `/cases/sighted` form params no longer accept |
| color/size. Frontend forms (MyDogs add-dog, ReportLost new-dog, ReportFound/Sighted) dropped the |
| fields, and the now-always-empty color/size display was removed from the owner views (MyDogs card, |
| DogDetail). |
| - **Deliberately kept:** the `color`/`size` DB columns (nullable) and the `*Out` response schemas — |
| no destructive migration, and any already-loaded/legacy data still round-trips. Batch-loader CSVs |
| and admin dataset views still carry their (often "unknown") values; these are inert now that |
| matching ignores them. FastAPI/Pydantic ignore the extra fields, so old clients/CSVs don't break. |
| - **Verified:** 71 backend + 14 frontend tests pass, clean typecheck/build. |
|
|
| ## Per-dataset + combined re-ID metrics (scripts/eval_datasets.py) |
| |
| - New read-only eval over the loaded (known+unknown) dataset pairs, auto-detected by name |
| ("<base> known" + "<base> unknown"). Reuses DB embeddings (active model). |
| - **Intra-subject similarity:** per dog, cosine over all pairs of its OWN photos (min/max/avg) → |
| `eval_out/intra_subject_similarity.csv`, plus a per-dataset summary and a sampled different-dog |
| baseline for context. |
| - **Retrieval quality:** found dog → ranked known gallery with the production dog-level max-cosine |
| score; recall@1/5/10 (+counts), mean/median rank, MRR, true-vs-best-wrong separation. Reported per |
| dataset AND combined (shared gallery; the other dataset's dogs become distractors). Identity is |
| namespaced per source (`base:folder`) so colliding folder numbers across datasets never cross-match |
| → this is how "filter matching by dataset" is realized for the eval. Summary → |
| `eval_out/retrieval_summary.csv`. |
| - Measures the MODEL (pure embedding retrieval, full gallery). Deliberately does NOT apply the |
| production ZIP-radius / breed gates: on this synthetic data each dog's found+known share a ZIP, so |
| a ZIP gate would inflate recall as an artifact rather than reflect the model. |
|
|
| ## Photo-level retrieval / CMC (scripts/eval_photo_retrieval.py) |
|
|
| - Different protocol from eval_datasets.py (which is dog-level: whole found dog vs whole known dogs). |
| Here every INDIVIDUAL photo (on the recombined pre-split whole dogs) is a query matched against ALL |
| other photos, its own excluded; a hit @k means a same-dog photo is in the top-k. Reported per |
| dataset then combined (shared photo gallery; other dataset = distractors). Recall@1/5/10/20 + mean/ |
| median rank + MRR. Identity namespaced per source so folder-number collisions don't cross-match. |
| - `_cmc` computes each query's nearest-same-dog rank as `1 + #{sim > s_true}` with s_true and the |
| count taken from the SAME similarity row (one matmul) — a first attempt used two matmuls and a |
| float mismatch nudged true matches off rank 1 (recall@1 low by ~5pts). Verified elementwise |
| identical to a brute-force argsort on MPDD. |
| - Results (hf-embed): MPDD r@1/5/10/20 = 65.1/81.0/87.3/91.8%; DogFaceNet 62.9/79.8/85.6/90.1%; |
| Combined 62.7/79.5/85.3/89.8%. Lower than the dog-level eval (single-photo query is weaker than |
| aggregating a whole dog + larger per-photo gallery). Summary -> eval_out/photo_retrieval_cmc.csv. |
| |
| ## Outlier photo removal + re-eval (scripts/remove_outliers.py) |
|
|
| - Quarantined the 22 flagged `removable_outlier` photos (10 MPDD + 12 DogFaceNet) — each the single |
| isolated photo whose removal lifts its dog's MIN pairwise similarity back above 0.5. Files moved to |
| `data/outlier_photos/` (reversible); Picture+Embedding+BreedPrediction rows deleted. Every affected |
| dog kept >=2 photos. Verified 0 DB-keys-missing-on-disk afterward. |
| - Re-ran both eval scripts on the cleaned data (counts −22 exactly): |
| - Consistency: dogs with min<0.5 fell MPDD 15->5, DFN 13->1 (remaining low-min dogs have >1 bad |
| photo, so no single removable outlier left). |
| - Photo-level CMC improved: MPDD r@1/5/10/20 65.1/81.0/87.3/91.8 -> 66.7/82.6/88.9/93.2; DFN |
| 62.9/79.8/85.6/90.1 -> 63.0/79.9/85.8/90.2; Combined 62.7/79.5/85.3/89.8 -> 62.9/79.6/85.5/90.0. |
| MPDD moved most (small set); DFN barely (12/8363). |
| - NOTE (pre-existing, unrelated): ~5240 orphan media files on disk with no DB row (from the earlier |
| DogFaceNet reload). Invisible to the DB-based evals; cleanup via scripts/prune_media.py when wanted. |
| |
| ## Fix: load progress bar freezes on large batch loads |
|
|
| Two independent causes, both hit only on LONG loads (many polls over many minutes): |
| - **Frontend (root cause):** the poll's `catch { clearInterval(t) }` permanently stopped polling on a |
| SINGLE failed request, so the bar froze at its last value while the load kept running server-side. |
| Now tolerates a run of transient failures (~10s / 20 polls) before giving up, and surfaces a "Lost |
| connection" error only if the job is truly unreachable (`failsRef` counter). |
| - **Backend (why polls failed):** SQLite had no `busy_timeout` and used the default rollback journal, |
| so the status poll's admin-auth SELECT could transiently hit "database is locked" during the load's |
| long write transaction. Added `PRAGMA journal_mode=WAL` (readers don't block on the writer) + |
| `PRAGMA busy_timeout=10000`. Bonus: the test suite got ~35% faster under WAL. |
| - Tests: new "keeps polling through a transient poll failure" AdminData test; all backend/frontend |
| green. |
| - Not changed (documented follow-up): the whole load is still ONE transaction (commit at the end), |
| so the SQLAlchemy session grows for very large loads. Periodic commit + expunge would bound memory |
| and smooth progress further, at the cost of all-or-nothing semantics. |
|
|
| ## Batch encoding progress bar (background embed job) |
|
|
| - The admin "Run Embeddings + Breeds" button was a synchronous blocking POST — no progress, and on a |
| large dataset it just spun / could time out. Converted encoding to a background JOB reusing the |
| same registry + polling the loader uses. |
| - Backend: `datasets.embed_all(db, dataset, *, progress=None, commit_every=25)` now commits + |
| reports `{processed,total,dataset_id}` every 25 pictures (and `expunge_all` to bound the session) |
| instead of one commit at the end — safe because encoding is additive/idempotent. New endpoint |
| `POST /admin/datasets/{id}/embed-all-job` runs it in a worker thread and returns a job id; the old |
| synchronous `embed-all` endpoint stays for small sets / tests. |
| - Frontend: shared `useJobPolling` hook (resilient polling, same transient-failure tolerance as the |
| load bar) + `JobProgressBar` component. DatasetDetail's encode button now starts the job, shows a |
| live bar, and prints the summary + refreshes stats on completion. |
| - Tests: `test_embed_all_reports_progress` (monotonic progress events), `test_embed_all_job_runs_in_ |
| background` (start->poll->done), and a DatasetDetail UI test (bar -> summary). All suites green. |
|
|
| ## Palette: warm tan/brown theme (replaced generic blue/white) |
|
|
| - The app consistently uses `brand-*` and `gray-*` Tailwind tokens, so retheming is a two-scale |
| change in tailwind.config.js — no component rewrites: |
| - `brand` → warm tan→coffee-brown scale (was blue). Recolors buttons, links, active nav, focus |
| rings, progress bars, and breed chips. |
| - `gray` → warmed taupe/sand neutrals (overrides Tailwind's cool default gray). Warms every |
| surface, border, and text token the app already uses (page bg gray-50 = warm cream, cards stay |
| white for a paper-on-desk feel). |
| - New `accent` = a muted dusty-denim blue (deliberately not the default bright blue); the 4 |
| generic `blue-100/blue-700` "found/unknown/open" badges now use it as a cool counterpoint. |
| - Contrast checked on the key pairs (brand-600/white ~6.6:1, brand-700 on brand-50, gray-900 on |
| gray-50, accent-700 on accent-100) — all comfortably readable. tsc/tests/build green. |
| |
| ## Date plausibility gate in matching |
|
|
| - A dog can't be FOUND before it was LOST, so a lost dog is only compared to found dogs whose found |
| date is on/after `lost_date - grace`. Added `_date_compatible(lost_date, found_date)` in |
| matching.py, applied (pre-cosine, alongside the ZIP + breed gates) in BOTH directions of |
| `run_matching_for_case`: lost-query→found-pool and found-query→lost-pool. Uses each case's |
| `event_date` (the user-entered lost/found date, not the report timestamp). |
| - Recall-first / fail-open: missing dates never exclude; only a definite conflict (found earlier than |
| lost minus grace) prunes. Config: `DATE_FILTER_ENABLED=true`, `MATCH_DATE_GRACE_DAYS=2` (grace |
| absorbs imprecise dates). The admin test-match (pure re-ID check) intentionally skips this gate. |
| - Batch-loaded datasets set both cases' event_date to the load date, so their found/lost dates are |
| equal → always compatible (dataset matching/eval unaffected). |
| - Tests: found-before-lost excluded despite identical photo; found within grace still matches. |
| |
| ## Narrow-first case matching + widen the constraints |
| |
| Directive: searches start narrow (location + top-10 breeds + top-10 matches) then widen. Scope |
| chosen: lost/found case matching, MANUAL widen button (keep it). What was already true vs added: |
| - Location: already narrow-first (starts at radius_levels[0]=0) with the manual `/widen` ladder. ✓ |
| - Top-10 breed gate: already applied (BREED_TOP_K=10, breed_filter_enabled). ✓ |
| - Top-10 matches: `TOP_N` was 5 → now **10** (the narrow-pass result cap). |
| - "Then widen the constraints": the widen ladder previously loosened only location. Now the WIDEST |
| rung (nationwide, radius -1) also **drops the breed gate** (`nationwide = radius == -1`; the breed |
| gate is skipped there). So the manual widen progresses location → … → breed, with no schema change |
| (state derived from the existing `search_radius_miles`). Date/ZIP/threshold gates unchanged. |
| - Tests: narrow pass caps at 10 matches; breed gate excludes a disjoint-breed dog at radius 0 but |
| matches it once widened to nationwide. |
|
|
| ## Privacy: shelter/vet location is public (only homes stay private) |
|
|
| - Reversed the earlier rule that hid `current_location_detail` (the shelter/vet holding a found |
| dog). A shelter/vet is a public place, and telling a matched owner where to reclaim their dog is |
| the point — so it's now included in the hydrated match candidate (hydrate.py) and `UnknownDogOut`. |
| Home locations remain protected: never collected below ZIP granularity, contact stays mediated. |
| - Frontend: MatchCard shows "📍 Being held at <shelter/vet>"; the found-report field is relabeled |
| (public, "use the shelter/vet name — not a home address"). |
| - Test flipped: `test_unknown_dog_shelter_shown_but_location_stays_zip_level` asserts the shelter |
| name IS returned while location stays ZIP-level. Updated model/schema/spec-§14 comments. |
|
|
| ## Removed the "sighted" workflow (found = you have the dog) |
|
|
| - A found report now always means the reporter HAS the dog: in their own custody (`pending`) or |
| dropped at a shelter/vet (`at_shelter`). Removed the seen-but-not-caught "sighted" flow entirely. |
| - Backend: deleted `POST /cases/sighted` and the shared `_create_found_sighted` helper (folded into a |
| single `create_found_case`); removed `CaseType.sighted` and the now-unused `UnknownDogStatus.lost` |
| ("sighted, not in custody"); dropped `lost` from the found/unknown matching pool filter (both the |
| case matcher and the public photo search); renamed `FoundSightedCreate`→`FoundCreate`; updated |
| seed.py's far-away demo dog from sighted→found. |
| - Data migration (raw SQL, enums already removed): 1 `sighted` case → `found`, 1 `lost` unknown → |
| `pending`. |
| - Frontend: removed the `/report/sighted` route + home button; `ReportFound` is found-only (no `kind` |
| prop); `api.createReport(kind)`→`api.createFound`; `CaseType` and MatchCard copy updated. |
| - All backend/frontend tests green; typecheck + build clean. |
|
|
| ## Removed seeded demo data (generated-dot "dogs") |
|
|
| - The synthetic `make_sample_images` dogs (Rex, Maple, 2 demo found dogs) and their info were noise |
| now that real datasets are loaded. New `scripts/purge_seed_data.py` deletes every non-dataset dog |
| (`dataset_id IS NULL`) + its pictures/embeddings/breeds/cases/matches/notifications/media, plus the |
| demo owner users — but ONLY non-admin users owning no dataset dog, so the admin and the 96 |
| detached (dataset_id-NULL) MPDD owners are preserved. |
| - Ran it: removed 2 known + 2 unknown dogs, 2 owners (alice/bob), 3 cases, 1 match, 4 pictures + 4 |
| media files. Verified 0 non-dataset dogs remain, admin intact, datasets intact (1489 known + 1489 |
| unknown), media 8862 = 8862 on disk (0 orphans). |
| - `scripts/seed.py` slimmed to create ONLY the admin account (no synthetic dogs), so re-seeding can't |
| reintroduce them; README updated. `make_sample_images.py` stays (the test suite uses it). |
| - All backend tests green. |
| |
| ## Generated realistic cases for every loaded dog (one-off script, since removed) |
| |
| - One LOST case per known dog + one FOUND case per unknown dog, rebuilt cleanly (old cases/matches |
| purged first). Pairing is by dataset + folder identity ("DogN" ↔ "Test found dog N"). |
| - Constraints honored: lost dates within the last 2 days; each found date is on/after its paired |
| lost date and never in the future; each found ZIP equals the paired lost ZIP or a neighboring ZIP |
| (another ZIP sharing the metro centroid — 40% of found reports); known-dog ZIPs (which were null) |
| taken from the paired found dog's ZIP. |
| - Each known dog got a fresh owner `owner{i}` (owner0…owner1488) as its owner + the lost case's |
| reporter. The ~2825 old shared/entangled test owners (now orphaned) were removed; admin kept. |
| - Verified all 1489 pairs: dates in range, found≥lost, all found ZIPs same-metro as their lost ZIP, |
| every known dog lost with an owner. Cases are NOT auto-matched (run matching separately if wanted). |
| |
| ## Admin dashboard: all cases, owners, and run-matching-on-demand |
| |
| - **All Cases** (`/admin/cases`, page + `GET /admin/cases`): every case, newest first, hydrated with |
| its dog (thumb/name/kind), reporter (owner for lost, finder for found), ZIP, date, status, and |
| match count. Filter by type (all/lost/found), paginated. Replaced the old thin `list[CaseOut]`. |
| - **Owners** (`/admin/owners`, page + `GET /admin/owners`): each owner with the dogs they own |
| (thumb/name/status), paginated. |
| - **Run matching on demand** (`POST /admin/cases/{id}/run-match`): runs the PRODUCTION matcher |
| (`run_matching_for_case` — narrow-first ZIP/breed/date gates, persists ranked Match rows) for a |
| chosen case and returns the hydrated matches. Surfaced as a "Run matching" button per row on the |
| All Cases page, opening a modal of the ranked candidates (MatchCard). Distinct from Test Matching |
| (`/admin/test-match`), which is the non-persisting pure-re-ID diagnostic. |
| - Admin dashboard links updated (All Cases / All Dogs / Owners / Test Matching / Data Loading). |
| - Tests: `test_admin_all_cases_and_owners`, `test_admin_run_case_match`. All suites green |
| (79 backend, 16 frontend); verified live on the 2978-case dataset. |
|
|
| ## Fix: run-matching 500 (matching must not load the model) + case detail view |
|
|
| - **Bug:** `POST /admin/cases/{id}/run-match` 500'd. Root cause: `run_matching_for_case` called |
| `get_embedder()` / `get_breed_classifier()` just to read the active model's name/version — which |
| loads the HF model over the network. On a server without a HF token that gets rate-limited/offline, |
| the load fails → 500. Matching only needs the PRECOMPUTED embeddings, never the model. |
| - **Fix:** added `active_embed_model_in_db(db)` (mirrors `active_breed_model_in_db`) and switched |
| `run_matching_for_case` + `rank_dog_against_opposite` to pick the model from stored rows — no model |
| load, works fully offline. Removed the now-dead `_active_breed_model()`. The public photo search |
| keeps `get_embedder()` (it must embed the uploaded photo anyway). Verified run-match returns 200 |
| with the hf-embed vectors under `EMBEDDER=mock HF_HUB_OFFLINE=1` (zero network calls). |
| - **Case detail:** new `GET /admin/cases/{id}` returns the case + its subject dog (all photos) + its |
| persisted matches (each candidate hydrated with photos). All-Cases rows are now clickable → |
| CaseDetailModal showing the dog's photos and the ranked matches, with a Run/Re-run matching button |
| inside. Tests: `test_admin_case_detail`; all suites green (80 backend, 16 frontend). |
|
|
| ## All cases moved to Houston ZIPs |
|
|
| - The case generator drew ZIPs from `HOUSTON_ZIPS` (the 77002–77099 subset of the |
| prepare list) instead of the spread-across-Texas dataset ZIPs. Conventions kept: one deterministic |
| ZIP per dog (per-folder seeded), found ZIP = same or a neighboring (other Houston) ZIP (~40%), |
| owner ZIP = the dog's ZIP, lost dates within the last 2 days, found ≥ lost. |
| - Owner creation is now idempotent (upsert `owner{i}@example.com`) so re-running doesn't hit the |
| email UNIQUE constraint. |
| - Re-ran: 1489 lost + 1489 found cases, all ZIPs Houston (95 distinct), 0 constraint violations. |
| - Consequence: all Houston ZIPs share one centroid, so at radius 0 every dog is in range of every |
| other — ZIP no longer narrows the pool; matching leans on photos (+ breed/date gates). Matches were |
| cleared by the rebuild; re-run matching (per case in the admin All Cases view, or in bulk) if wanted. |
|
|
| ## Side-by-side photo comparison (case dog vs each match) |
|
|
| - `MatchCard` now takes optional `queryPhotos` + `queryLabel`. When present, clicking a match's |
| thumbnail opens a TWO-COLUMN modal: the case's own dog photos on the left, the candidate match's |
| photos on the right (with the % match). Without those props it falls back to the single-dog gallery. |
| - Owner view (CaseDetail): new owner/admin-authorized `GET /cases/{id}/dog` returns the case's dog |
| (kind + all photos); CaseDetail fetches it and passes it to each MatchCard. |
| - Admin view (AdminCases → CaseDetailModal): passes the case's subject-dog photos (already in the |
| detail payload) to each MatchCard. So comparison works in both regular and admin case viewing. |
| - Tests: `test_case_dog_endpoint_returns_photos_and_guards_owner` (photos + 403 for non-owner). All |
| suites green (81 backend, 16 frontend). |
|
|
| ## Removed human-entered breed; display the model's estimated breed (top 3) |
|
|
| - Owners/finders are unreliable on breed, so it's no longer collected. Removed the breed input from |
| `KnownDogCreate`/`KnownDogUpdate`, `LostCaseCreate.new_dog_breed`, `FoundCreate.est_breed`, the |
| `/cases/found` `est_breed` form param, and the frontend forms (MyDogs add-dog, ReportLost new-dog). |
| - Everywhere breed was shown, it now shows the **top-3 estimated breeds** (model's HF-softmax |
| predictions, already surfaced as `estimated_breeds` / `predicted_breeds`): MatchCard (card + modal |
| + comparison), DogPhotosModal, PhotoSearch results, AdminDogs cards, AdminTestMatch candidates, |
| DogDetail (removed the human "Breed" field), and the ReportLost dog picker. Shared |
| `estimatedBreeds()` helper (top-3, "Breed not estimated" fallback). |
| - DB columns (`KnownDog.breed`, `UnknownDog.est_breed`) kept nullable for backward compat (same as |
| color/size); they're just no longer collected or displayed. Matching never used human breed anyway |
| (the gate uses the model's `breed_predictions`). |
| - All suites green (81 backend, 16 frontend), clean build. |
|
|
| ## Found-report 500 fix + owner re-run + admin delete buttons |
|
|
| - **Found-report 500 (root cause):** each found report ran `rematch_open_lost_cases_against`, which |
| re-ran FULL matching for EVERY open lost case (1490 now) — O(cases×candidates), holding the SQLite |
| write lock for minutes → timeout/500 and locking out other requests (the Houston co-location made |
| every dog a radius-0 candidate, worsening it). The found case's OWN matching already surfaces + |
| notifies matching lost dogs, so the auto-rematch is redundant. Made it bounded + geo-filtered and |
| OFF by default (`REMATCH_MAX_CASES=0`); raise it to opt in for small sets. Also fixed a latent FK |
| bug: `run_matching_for_case` deleted prior pending matches without detaching notifications that |
| referenced them (nulls `Notification.match_id` first now) — this surfaced once re-matching a case |
| that had produced a strong-match notification. |
| - **Owner re-run matching:** `POST /cases/{id}/rematch` (owner/admin, re-runs at current radius) + |
| a "Re-run matching" button on the owner CaseDetail (next to Widen). Two tests updated to re-run |
| explicitly now that auto-rematch is off by default. |
| - **Admin delete (temporary, for re-demoing):** `DELETE /admin/dogs/{kind}/{id}` (dog + pictures + |
| media + embeddings + breeds + its cases + matches referencing it) and `DELETE /admin/cases/{id}` |
| (case + matches, dog kept). Delete buttons on AdminDogs cards and AdminCases rows. |
| - Tests: owner rematch (+403), admin delete dog/case (+404/400). All green (84 backend, 16 frontend). |
|
|
| ## Admin dashboard as a button hub + admin nav + delete people |
|
|
| - Admin home (`Admin.tsx`) is now just a grid of bigger link cards (Pending Matches, All Cases, All |
| Dogs, Owners, Test Matching, Data Loading). The pending-matches review view moved to its own page |
| `AdminPending.tsx` at `/admin/pending`. |
| - Nav: when an admin is logged in, "My Dogs"/"My Cases" are hidden (only the Admin link + Log out). |
| - Delete people: `DELETE /admin/owners/{id}` (+ `delete_person` service) removes a person and their |
| dogs/pictures/media/embeddings/breeds/cases/matches; refuses to delete an admin. Delete button per |
| owner on the Owners page (dogs + cases already had delete buttons). |
| - Test: `test_admin_delete_person` (deletes owner+dogs, 404 on re-delete, 400 on admin). All green. |
|
|
| ## 1-decimal match % + tie-break by most-recent report |
|
|
| - Match percentages now show one decimal (e.g. 94.4%) so close scores stop collapsing onto the same |
| integer: ConfidenceBar, MatchCard header, and the PhotoSearch result badges. |
| - Matching ranking (`run_matching_for_case`): primary sort is the similarity ROUNDED to one decimal |
| (matches the display); candidates that tie at that precision are ordered most-recently-reported |
| first, using the candidate's case `created_at` (fetched alongside the date-gate query; ties with no |
| report time sort last). `CandidateScore` gained a `reported_at` field. |
| - Tests: `test_equal_score_matches_break_ties_by_most_recent_report`; two frontend tests updated for |
| the decimal. All green (87 backend, 16 frontend). |
|
|