File size: 5,355 Bytes
71d239c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
# StoryCode β€” Handoff for the executing agent (Codex)

Full design rationale: `~/.claude/plans/so-uh-this-is-modular-curry.md`.
This file is the **do-this-next checklist**. Deadline: **2026-06-15**.

Read `AGENTS.md` first β€” it has the rules you must not break.

## What's already built & verified βœ…

The whole app is written, and the **deterministic core is tested green**
(`python tests/test_analyzer.py` β†’ 10/10, no GPU). The Gradio UI builds and every
panel renders. The only thing not yet exercised is the live model β€” because it
needs the Modal endpoint.

| File | State |
|---|---|
| `config.py`, `schema.py`, `ingest.py` (+ secret scan) | βœ… done & tested |
| `analyzer/` (`python_ast`, `js_treesitter`, `generic`, `graph`, `deps`, `__init__`) | βœ… done & tested β€” **the source of truth** |
| `diagram.py`, `story.py`, `narrate.py`, `db.py` | βœ… done; narration verified via the model-free fallback |
| `llm.py` | βœ… written; needs the live Modal endpoint to exercise the real path |
| `ui/theme.py`, `ui/styles.css`, `app.py` | βœ… written; UI builds. Needs a live `python app.py` run |
| `modal_app.py` (vLLM serving MiniCPM4.1-8B) | βœ… written; **you must deploy + version-pin it** |
| `zerogpu_backend.py` | βœ… break-glass fallback |
| `scripts/sample_project/`, `tests/test_analyzer.py` | βœ… done; tests pass |
| `README.md` (HF frontmatter+tags), `.env.example`, `requirements.txt` | βœ… done |

## Do these in order

1. **Deploy the GPU backend.**
   - In `modal_app.py`: confirm `VLLM_VERSION` against the **MiniCPM4.1-8B model
     card** (it uses custom code β†’ `--trust-remote-code`). Set a real shared secret:
     `modal secret create storycode-api MODAL_API_KEY=<secret>`, then
     `modal token new` and `modal deploy modal_app.py`.
   - Smoke-test the endpoint:
     `curl -H "Authorization: Bearer <secret>" <url>/v1/chat/completions -d '{"model":"openbmb/MiniCPM4.1-8B","messages":[{"role":"user","content":"Reply with JSON {\"ok\":true}"}]}'`
   - **Risk:** MiniCPM4.1 is a *hybrid-reasoning* model. If it emits long `<think>`
     chains or ignores `guided_json`, (a) check the model card for the flag that
     disables thinking and add it to the vLLM args, and (b) confirm the installed
     vLLM version supports `xgrammar` guided decoding. The contract `llm.py`
     expects is a plain OpenAI `/v1/chat/completions` that honours
     `extra_body={"guided_json": ...}`.

2. **Wire secrets.** Copy `.env.example` β†’ `.env` locally; on the Space set
   `MODAL_ENDPOINT_URL` (= `<url>/v1`) and `MODAL_API_KEY` (same shared secret).

3. **Run locally.** `pip install -r requirements.txt && python app.py`.
   Click **"Try the sample project"** β†’ you should get a Story + Plain-English
   panel + a rendered Architecture Map + the Safe-to-Edit list + Dependencies.
   Switch styles/difficulty β†’ the Story re-narrates (no re-analysis).

4. **Tune the narration prompts** in `story.py` (and only `story.py`) until the
   output is accurate AND in-voice for each of the 5 styles Γ— 3 difficulties.
   This is the main creative loop and your clearest Codex-attributed work. Keep
   the grounding rule intact (the model narrates facts; it never invents files).
   Re-run `python tests/test_analyzer.py` after β€” it must stay 10/10.

5. **Real-user proof (Backyard AI requirement).** Run the friend's actual
   Claude-generated app through StoryCode. Capture: a quote from them, before/after
   screenshots, and the story it produced. Put these in the README.

6. **Deploy + submission assets.**
   - Create a Gradio Space under `build-small-hackathon`; push this repo; set the
     two secrets. Verify it reaches Modal (check Modal logs). **Test on mobile.**
   - Record a 60–90s demo video (upload β†’ Story β†’ Map β†’ Safe-to-Edit); link in README.
   - Push to a public GitHub repo with **Codex-attributed commits**; link it in
     the README (OpenAI Codex prize requirement).
   - Post once on social; link it. Confirm README frontmatter tags are present.

## Post-MVP queue β€” only after steps 1–6 are deployed. One at a time.

In priority order (each is a clean, bounded task; details in the plan file):
1. **Grounded chat Q&A** β€” "Ask anything about your code", answered from the
   ProjectModel + summaries already in `gr.State` (no new analysis).
2. **"How do I change X?"** assistant β†’ "file X, line Y, change Z to W."
3. **Error detective** β€” paste an error + file β†’ detective-story fix.
4. **Dependency panel polish** β€” surface `analyzer/deps.py` risk flags more loudly.
5. **Save / revisit + Share** (wire `db.py` into the UI).
6. **Export story as PDF / Markdown.**
7. **GitHub URL ingestion** (add a `from_github` path to `ingest.py`).
8. **Before/after edit preview.**

## Gotchas
- **Don't** add `torch`/`vllm`/`transformers` to `requirements.txt` β€” the GPU is on
  Modal; the Space is a CPU container. (Those belong only in `modal_app.py`, or in
  the Space *only* if you switch to the ZeroGPU fallback.)
- Keep the UI custom (`ui/styles.css`) β€” Off-Brand badge + explicit requirement.
- **Don't claim Tiny Titan** β€” MiniCPM4.1-8B is ~8B, not ≀4B.
- The static analysis is the source of truth. Never let the model override the
  Architecture Map or the Safe-to-Edit verdicts. If you change `analyzer/`, update
  `tests/test_analyzer.py` to match and keep it green.