SmartSight Coach LoRA fine-tunes

LoRA fine-tunes for the SmartSight AI coach — an on-device fitness and nutrition coach that runs entirely on the phone, no account and no server round-trip. Converted to .litertlm for Android/iOS inference. → smartsight.app/ai-coach

Base model changed at v42 (2026-08-12). v27–v38 were fine-tuned from google/gemma-4-E2B-it. Everything from v42 on is fine-tuned from Google's quantization-aware-trained checkpoint google/gemma-4-E2B-it-qat-q4_0-unquantized, whose weights are conditioned to survive 4-bit rounding. The frontmatter above reflects the CURRENT parent, so the model tree matches what actually shipped. This repo keeps the current best checkpoints plus earlier superseded milestones for history — not every candidate tested each round. Rounds that lose are documented internally with their real failure cases; they are never uploaded here.

More coming. The vision tower is being trained surface by surface — body fat, meal photos and nutrition-label reading — each measured against held-out data before anything ships. Progress and honest limitations for every round are written up below and at smartsight.app/ai-coach.

smart-coach-vision.litertlm — CURRENT CHAMPION (v46-seed123)

One multimodal model at 44% of the download, and the last surviving defect eliminated. Promoted 2026-08-12. Lineage: v27 → v31 → v34 → v38 → v42 → v44 → v46.

This is the biggest release since the model was first fine-tuned. Two things landed together: the coach stopped needing a second AI model to see, and the last defect that had survived every round since v38 was eliminated — on both training seeds, with nothing traded away.

1. One model instead of two

The app used to download two models: the coach, and a separate 3.66 GB Gemma-3n vision model for body-fat photos, meal photos and nutrition labels. The coach's own base checkpoint was multimodal all along — 211 vision tensors were being discarded at export.

Exporting them instead costs +0.17 GB and removes an entire model:

before v46
download for a user who wants photo features 2.63 + 3.66 = 6.29 GB 2.80 GB
AI engines resident 2, swapped per photo 1
vision backend GPU required runs on CPU

56% less to download, and a far larger share of phones can run it at all — the second model was what put photo features out of reach on mid-range devices.

2. The engine swap, and the crashes it caused, are gone

This matters more than the megabytes. Every body-fat or label scan used to evict the coach engine (LocalVisionModel.analyzeBatchcoachModel.close()), load the vision engine, run, and then the coach cold-loaded again on the next question. Two shipped crash classes trace directly to that juggling: a meal-plan loop reloading the coach onto the GPU during a pose session (killed with signal 9), and the coach plus speech model sitting resident while the camera and pose landmarker loaded. Both previously needed defensive workarounds.

One engine deletes the swap entirely. And because the consolidated tower runs vision on CPU while the standalone model hard-refuses CPU (Model requires one of [gpu]), vision also moves off the contended GPU that caused those crashes in the first place.

3. Text quality: unchanged, and measured

The vision export path forces different embedder packaging (externalize_embedder, single_token_embedder), so text had to be re-tested rather than inherited:

user-visible defects (89 held-out) throughput
v46 text-only 0 26.8 ch/s
v46 with vision tower 0 26.0 ch/s

Identical. The vision capability is free in quality terms.

4. The last surviving defect, eliminated

Meal generation used to re-serve foods the prompt had explicitly banned — the same item, portion string and all (Skyr, 120 g, Kebab meat, 120 g). It had survived every round since v38 and more training data had never moved it.

The cause was found by reading the shipping app code rather than the model's output. MealPlanStore sends the avoid-list semicolon-separated with portions, alongside a positive shortlist of allowed foods. The training corpus taught a comma-joined list of bare names with no shortlist — a prompt shape production had stopped sending. The app had switched to semicolons deliberately: "a comma-joined list of comma-containing names is one run-on string to a small model." The model had never trained on the prompt it actually receives.

Fix: swap 128 stale-format rows for 139 production-format ones. No new data collected.

v44 v46-seed123 v46-seed456
avoid-list violations (36-case held-out probe) 5/36 (14%) 0/36 0/36
user-visible defects (89 held-out) 2 0 0
throughput, same machine 23.3 ch/s 26.8 25.2
suites worse than the previous champion 0 of 9 0 of 9

Zero on both seeds. If the true rate were still 14%, scoring 0 across 36 cases has probability ~0.4% — and the mechanism was identified in the source before the round was run, not fitted afterwards. A per-suite diff confirms nothing was traded: eight suites identical, one better. It is also the first model to score zero user-visible defects across the whole battery.

5. The vision tower is trained, not merely present

On three real front-angle photos of a lean subject, rubric mode, same coach weights, only the tower differing:

photo untrained tower v46 tower
1 14–17% — "faint abs and some softness" 10–13% — "clear abs and visible muscle definition"
2 14–17% 12–15%
3 14–17% — "faint abs and some softness" 10–13% — "clear abs and visible muscle definition"

Three improvements at once: the reads came down, the description became factually correct (the untrained tower called clearly separated abs "faint"), and it stopped returning the same answer for every photo — which is what makes the app's self-consistency blending meaningful at all.

⚠️ It is still not accurate, and we know by how much. That subject is DEXA-verified at 7%, so the correct band is 6–8%. Training closed roughly half the error (+7–10 pp → +3–6 pp); the tower still reads a full band high. Diagnosis: the right class exists in training with 85 rows at true_bf 7.0, but those are augmentations of only five distinct images, all slim lean physiques — while the subject is muscular lean. Ten identities cannot teach calibration, and augmentation multiplies pixels rather than people. Photo body-fat estimates should continue to be treated as a guide, not a measurement.

What ships where

The app chooses automatically — there is nothing for a user to select. Builds already installed keep downloading the text-only file and behave exactly as before; the next release points at the consolidated model and stops fetching the separate vision model. Both files carry the same v46 coach weights, so no one is held on an older coach for stability's sake.

More at smartsight.app/ai-coach.

Earlier rounds, kept for context

What v44 had changed — non-English quality, fixed in the DATA (kept for context)

The v42 champion wrote English macro words inside Danish prose ("240g carbs og 70g fedt") and answered a Danish challenge request with an English title ("Volume Boost Challenge"). Measured cause: of 7,288 corpus rows, 13 contained any Danish/Norwegian, zero taught "kulhydrat", and ten demonstrated exactly the wrong behaviour — English macro words inside non-English prose.

v44 localises those ten rows (carbskulhydrat, fatfedt, word-boundary anchored so "fatigue" survives) and adds 26 Nordic follow-up rows: 7,314 rows total. The machine /command lines stay English, because the app parses them. Prompts are unchanged, so this needs no app update.

v42-seed123 v44-seed123
user-visible defects (89 held-out) 2 2
throughput, same machine 22.6 ch/s 23.6 ch/s
Danish macros "240g carbs og 70g fedt" "240g kulhydrater og 70g fedt"
Danish challenge title "Volume Boost Challenge" "Byg mere volumen"
/nutrition line emitted 5/7 7/7
Nordic hard-probe red flags 2 0

Throughput is quoted from a champion re-run on the SAME machine as the candidate; chars/sec is hardware-dependent and cross-machine numbers are not comparable.

Seed 123 was chosen over seed 456 even though 456 scored better on the automated count (1 defect vs 3 for the champion on its box). Reading the transcripts, 456 opened a muscle-building plan with "Velkommen til vægttab" and prescribed a calorie deficit, and emitted a bare /nutrition line carrying no values at all. Neither is visible to the scorer.

Regressions, stated plainly. Nordic follow-ups about an unlogged past activity got worse, not better, despite the 26 rows added for exactly that; and Danish morphology is now occasionally wrong ("vækning", "fat-mål"). Both are listed under weaknesses below and are the next round's targets.

What v42 had changed — two quantization findings and one data fix (kept for context)

  1. New base. Switched to Google's QAT checkpoint google/gemma-4-E2B-it-qat-q4_0-unquantized. Identical architecture, byte-identical tokenizer.json and chat_template.jinja, identical eos_token_id [1,106,50] — but the weights are conditioned to survive 4-bit rounding.
  2. New export recipe. The previous Hadamard-rotation export was measured to erase that conditioning completely: post-rotation the QAT and stock checkpoints quantize identically (relative RMSE 0.1288 vs 0.1287). Rotation smears the block-local structure QAT creates, and ai_edge_quantizer makes rotation and blockwise mutually exclusive. Switched to blockwise-64 int4 on attention/MLP layers (embedders stay on the rotation path for size): reconstruction error 0.129 → 0.060. Measured recipe sweep on one fixed adapter — HR 28 flagged, b32fc 20, b64fc 18, blockwise-everything 23.
  3. Data fix, not an app workaround. The coach emitted ChallengeGen JSON without its closing brace, so 3 of 4 challenges were dropped by the parser. Root cause was the corpus: it taught bare JSON whose only terminator is one character. The stock base masked this by fencing spontaneously; the QAT base reproduced what we taught. Fixed by fencing 29 ChallengeGen completions. Prompts unchanged, so it works on already-shipped app builds.

v42's own measurement against v38, on 89 held-out cases, scoring only defects that reach a user — the app already rescales meal calories (MealPlanScaler) and clamps challenge XP, so counting those measured nothing:

v38-seed456 v42-seed123
user-visible defects 3 2
throughput 17.6 ch/s 19.9 ch/s
ChallengeGen JSON 4/4 4/4 (native)
size 2.577 GB 2.629 GB

Both training seeds scored 2, so this is not a single lucky run. Seed 123 was chosen over seed 456 — same score — because 456 produced one confident wrong number (subtracting a 78 kg bodyweight goal from a 106.7 kg squat 1RM and presenting the difference as progress).

Full-transcript read. v42 fixes three real v38 errors: the long-debunked "don't let your knees go past your toes" cue, a garbled lat-pulldown setup, and inventing a streak figure from a claim it could not verify. It also uses the user's name and respects their listed equipment where v38 answers generically. One regression: a garbled deadlift cue ("shins close to your thighs").

Champion history — what each round actually fixed

Every entry here was a promoted champion, verified against the previous one on a held-out suite with a full read of every answer. Rejected rounds (v28-v30, v32, v33, v35-v37, v39, v40, v41) are not listed — they are documented internally with their real failure cases.

Round Beat What it fixed Base
v27 baseline first winning fine-tune — shorter, equally accurate general advice gemma-4-E2B-it
v31 v27 exercise-form answers reached the untouched baseline's zero-error record (v27 was wrong on 4/5 direct form questions); no more garbled or duplicated plan output. The unlock was the export recipe, not data — three consecutive data-only rounds failed first, then Hadamard-rotation int4 fixed it with zero retraining gemma-4-E2B-it
v31 export update same weights, AlgorithmName.HADAMARD_ROTATION custom op instead of the decomposed variant: 88% faster decode, ~85 MB smaller, quality statistically unchanged (5.8% vs 6.4% defect rate over 312 generations)
v34 v31 dietary-rule compliance (a "vegetarian, no nuts" request had served peanut butter); first correct full examples for face pull and plank; squat-vs-hip-thrust contrast cases gemma-4-E2B-it
v38 v34 AI challenges: unrealistic targets for low-frequency goals (150 progress photos in six weeks from zero) and non-English requests answering in English. Won over its own alternate seed on a Danish grammar defect — plural "jeres/jer" where the app addresses one person gemma-4-E2B-it
v42 v38 see above — QAT base, blockwise export, and the fenced-JSON data fix. Won over its own alternate seed, which produced a confident wrong number (subtracting a 78 kg bodyweight goal from a 106.7 kg squat 1RM) gemma-4-E2B-it-qat-q4_0-unquantized
v44 v42 Danish macro vocabulary (kulhydrat/fedt) and Danish challenge titles; /nutrition line now emitted on every applicable case gemma-4-E2B-it-qat-q4_0-unquantized
v46 v44 avoid-list violations eliminated (0/36 vs 5/36) by matching the corpus to the prompt format production actually sends gemma-4-E2B-it-qat-q4_0-unquantized

Recipe constant across every round since v22c: LoRA r=16 / alpha=32 / dropout=0.1, 2 epochs, lr 6e-5, NEFTune noise_alpha=5, completion-only loss masking (train_on_responses_only), per-task validation shards, two seeds per round with the winner chosen on behaviour rather than validation loss. Only the training data, the base, and the export recipe have moved.

Honest current weaknesses

Documented rather than hidden, because they are the targets for the next rounds:

  • Meal-plan portions are sized by habit, not arithmetic. The model anchors on ~100 g portions rather than solving the prompt's stated calorie budget. Mostly invisible in the app, which rescales a day to its target (MealPlanScaler, clamped 0.6x–2.0x) — but one case per model is still off after that clamp, and the underlying arithmetic is unsolved.
  • Avoid-lists are not always respected — foods explicitly banned for variety reappear (2/29 held-out cases on v44, v42 and v38 alike; this is the single defect class that has survived every round so far).
  • Nordic follow-ups about an UNLOGGED past activity regressed in v44. Asked in Norwegian about a hike "yesterday" that is not in the app, v42 correctly said the gap was only a logging omission; v44 answers as though advising about today. 26 Nordic follow-up rows were added to fix exactly this and did not move it — the cause is not simply coverage.
  • Occasional Danish/Norwegian word-formation slips in v44 — "vækning" (not a word) and the hybrid "fat-mål". The macro vocabulary itself is now correct, but morphology is not reliable.
  • The /nutrition command line is improvised. The corpus contains ZERO examples of it; the format lives only in the prompt. v44 emits the line on 7/7 applicable cases (v42 omitted it entirely on 2/7) but drops the trailing fat value on 3/7, so the app applies three of the four fields. Harmless by design — NutritionCommandParser accepts any subset — but the fat target silently stays stale.
  • Within-day duplicate or misplaced exercises still appear occasionally in generated routines (e.g. a chest isolation movement landing in a Legs day). Lower real-world impact than it sounds, since the app's own parser drops duplicate slots.
  • Confabulation-despite-coverage on a few specific lifts (squat/hip-thrust blending, farmer's carry) that does not reliably respond to more training data.
  • Superset pairing — avoiding two competing-muscle compounds back to back — is imperfect.
  • Tight macro-budget precision (±8–10%) is unsolved across every model tested, including the untouched baseline.

Evaluation discipline

Every promote/reject decision reads every row of the eval suite, never a sample — an early round reported a verdict from ~12 of 52 rows and missed hard timeout loops entirely. Prose quality is judged as its own axis alongside structural correctness, because a structurally clean answer can still be flat, templated, or subtly wrong in a second language. Candidates are always compared against both the previous champion and the untouched baseline.

Files

Which file the app fetches

Two files are published. The app decides which one to fetch based on its build version; there is nothing to select and nothing to configure.

file who gets it size
smart-coach-vision.litertlm current champion — text coaching AND vision (body-fat photos, meal photos, nutrition labels) in one model. Fetched by the next app release. 2.80 GB
smart-coach.litertlm text-only, fetched by app builds already installed. Unchanged on purpose. 2.63 GB

Both carry the same v46 coach weights, so an older build is not stuck on an older coach. They exist side by side because installed builds also fetch a separate 3.66 GB vision model — handing them the larger consolidated file would raise their total download on devices already tight on RAM. See the current-champion section above for what the consolidation actually changes.

  • smart-coach-vision.litertlm — ⭐ current champion, fetched by the next app release. The v46 coach WITH a working vision tower: one model for text coaching, body-fat photos, meal photos and nutrition labels (2.80 GB). The next app release points both the coach and vision URLs here and stops downloading the separate 3.66 GB vision model.
  • smart-coach.litertlmSTABLE, text-only (2.63 GB). The file app builds already in the wild download at runtime on Android and iOS, and the one they keep using. Deliberately left unchanged: those builds also fetch the separate vision model, so giving them the larger consolidated file would raise their total download on devices that are already RAM-constrained. Same v46 coach weights, so nobody is stuck on an older coach for stability's sake.
  • coach.litertlm — the original working int8 build from the v1 era (2.59 GB), kept as the historical starting point.
  • coach-finetuned-int4.litertlm — the first int4 conversion (2.56 GB), kept because it is the artifact that exposed the LiteRT-LM chat-template .get() incompatibility.
  • model.safetensors + config.json + tokenizer files — HF-format artifacts for reference.
Downloads last month
86
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hanemay/smartsightCoach

Finetuned
(17)
this model