| # Honest Submission Notes — TurboSkillSlug |
|
|
| This document is the unvarnished account of what TurboSkillSlug actually does, |
| what is genuinely strong, what is partial, and what is aspirational. If a claim |
| is not in here with its caveats, treat the marketing copy as marketing. |
|
|
| Demo video: https://youtu.be/qSP9olWRv7o |
| Space: https://huggingface.co/spaces/build-small-hackathon/TurboSkillSlug |
| Code: https://github.com/AnubhavBharadwaaj/turbo-skill-slug |
|
|
| ## What it does, honestly |
|
|
| You give the slug a build session two ways: narrate it aloud (audio), or drop a |
| Claude Code / Codex CLI session trace (`.jsonl`). It transcribes/parses, extracts |
| a structured record, and returns four artifacts: a transferable SKILL.md, a |
| grounded second-person recap in a fine-tuned "slug voice," a procedural SVG shell, |
| and a thermal receipt. The shell is born on screen as a scroll that unrolls along |
| its spiral arm, with a byōbu (folding-screen) battle drawn across it. |
|
|
| ## What is genuinely strong |
|
|
| - **Two custom 1.5B LoRA adapters, both published and both used.** Voice |
| (500 pairs, loss 0.89) and extraction (167 pairs, loss 0.81). The extraction |
| LoRA replaced a prompted Qwen-7B, bringing the primary pipeline to ~2.6B active |
| params (Whisper 809M + one 1.5B base serving both adapters on a single T4). |
| - **The groundedness eval is real and published with its raw data.** On 25 |
| held-out transcripts the fine-tuned 1.5B reaches semantic groundedness 0.76 vs |
| the 7B's 0.72, at a third the active size, by paraphrasing rather than copying. |
| - **The shell is fully procedural and traceable.** Every visual element derives |
| from a real session feature: dead ends are knots/fallen warriors, gotchas are |
| jewels/archers, the breakthrough is the aperture/dragon, sentiment is the color |
| arc. No image generation. |
| - **The SKILL.md is built for real uplift**, gotchas-first (symptom→cause→fix), |
| "what does NOT work and why," transferable principles. Built and de-leaked |
| deterministically from the structured extraction, not trusted to model prose. |
| - **Trace input works.** A real Claude Code or Codex session log runs through the |
| same pipeline as audio. Judges can feed their own logs. |
| - **Graceful degradation is real and now observable.** If the extract adapter |
| emits invalid JSON, the 7B fallback catches it. If the voice adapter is down, a |
| deterministic extraction-based voice net prevents placeholder recaps. Every |
| voice-path outcome is logged with a `[VOICE]` tag (EXTRACT_DOWN / VOICE_DOWN / |
| VOICE_EMPTY / VOICE_OK / FALLBACK_VOICE_OK / NET_FROM_EXTRACTION / …) so failures |
| are diagnosable, not mysterious. |
|
|
| ## What is partial or has real caveats |
|
|
| - **Extraction parse reliability: 21/25 vs the 7B's 24/25.** The smaller model |
| produces invalid JSON more often. In the app a brace-walking parser and field |
| validators recover most of this, and the 7B fallback catches the rest, but the |
| raw rate is a real cost of going small, and we report it unsoftened. (You will |
| see `[VOICE] EXTRACT_DOWN` in the logs when this fires; the fallback then runs.) |
| - **The groundedness metric is calibrated to 5/6, not 6/6.** A hand-labeled |
| calibration block (run before scoring, printed in the logs) missed one case: |
| a grounded paraphrase scored just under threshold. The miss understates the |
| LoRA rather than inflating it. 25 transcripts is a small sample; treat the |
| LoRA-vs-7B gap as "matches or slightly exceeds," not "beats." |
| - **SKILL.md quality depends on extraction quality.** A good session yields |
| transferable gotchas; a thin session yields thinner ones. An optional one-shot |
| enrichment pass expands terse gotchas, gated by env and best-effort. We caught |
| and fixed a real bug where few-shot examples in the enrichment prompt leaked |
| into output (a game-theory skill got tree-coloring gotchas); the prompts are |
| de-leaked and a guard now rejects any leaked phrasing. |
| - **The scroll-unroll animation ships ~1.7MB per shell** (14 stacked growth |
| stages, so the parchment lays down ALONG the spiral arm rather than as a radial |
| wipe). Browser-rendered via SMIL in a sandboxed iframe, with a "watch it unroll |
| again" replay. On a slow connection the first paint takes a beat. If SMIL fails, |
| the full shell still shows (nothing is hard-hidden); animation is enhancement, |
| not a dependency. |
| - **The 3D paper curl rides the spiral tip via animateMotion.** Verified |
| frame-by-frame in offline renders (cairosvg + headless Chromium). Exact |
| tip-tracking in every browser is the part most likely to need a small tweak. |
|
|
| ## What was freshly added near the deadline (use with that in mind) |
|
|
| - **Shared terrarium gallery.** Kept shells save to a Modal Volume; a grid shows |
| all of them newest-first; each has a complete clickable permalink that re-loads |
| and re-animates that shell. Freshly built and lightly tested. Honest caveats: |
| each grid card is its own iframe, so past ~30-40 shells the grid gets heavy |
| (lazy-loading is the fix, not yet done). Gallery saves depend on the Modal |
| endpoint URLs matching the client; if they drift, a save fails and the UI says |
| so rather than failing silently. |
| - **Battle Trace (experimental).** A temporal replay of the same session as a war |
| between you (the Agent) and the Environment, fed by the SAME extraction (no |
| second parser, no OTel). Framed as complementary to the shell: the shell is the |
| slug's *memory* of how the battle ended (the frozen folding screen); the trace |
| is the *replay* of it in time. It is Canvas-2D with simple figurative |
| combatants (a samurai general, a horned adversary, fallen warriors, a dragon at |
| the breakthrough). Honest about its level: these are clean vector figures, not |
| game-quality sprite art. Labeled "experimental" in the UI. |
|
|
| ## What is aspirational / not built |
|
|
| - **Session diff view** (compare two sessions' shells) is not built. |
| - **Closing the parse-rate gap** needs more training pairs and constrained |
| decoding; not done. |
| - **Higher-fidelity Battle Trace art** (detailed sprite-grade samurai) is not |
| built; the current figures are deliberately simple vector shapes (see the |
| Battle Trace note above). |
|
|
| ## Infrastructure honesty |
|
|
| - The Qwen-7B still exists in the codebase but ONLY as a labeled fallback when the |
| Modal extract path returns nothing usable (including the invalid-JSON case |
| above). The primary path is the 1.5B dual adapter. The README architecture table |
| reflects this. |
| - The dual-adapter server can hold one warm container for demo reliability |
| (~$0.60/hr). It should be stopped after judging (`modal app stop |
| slug-dual-serve`). |
|
|
| ## How to verify our claims |
|
|
| - Groundedness: `modal run semantic_eval.py` reproduces the table; raw |
| generations are published in the eval dataset so anyone can re-score. |
| - Trace input: drop the sample `.jsonl` (one click) or your own session log. |
| - Models: both LoRA adapters are public on the Hub. |
| - Shell traceability: change a session's sentiment or dead-end count and watch |
| the shell's colors and knots change accordingly. |
| - Failure handling: the Space logs tag every voice-path outcome with `[VOICE]`. |
|
|
| ## The one-sentence honest summary |
|
|
| A small, slow, genuinely ~2.6B pipeline that turns a coding session into a |
| transferable skill, a grounded recap, and a procedural shell, measured honestly |
| (including where the small model costs us), with graceful, observable degradation, |
| and every model and the eval data published for anyone to check. |
|
|