File size: 29,178 Bytes
ef8744f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
# Memory-LoRA: A Hypernetwork that Writes Repo-Specific Adapters for Gemma-4-E2B

> **One line:** we train a small neural network (a *hypernetwork*) that reads an
> embedding of a codebase and **emits a LoRA adapter** for a frozen
> `google/gemma-4-E2B`. The adapter injects repo-specific knowledge into the
> model with **zero extra tokens at inference time** β€” no RAG, no context stuffing.
> Everything runs **locally on an Apple-Silicon Mac** (M4 Pro, 64 GB).

This document is the onboarding bible for the project. It covers *what* we built,
*why* every non-obvious choice was made, *what data* we used, *every experiment we
ran and its result*, and the **hard-won lessons** (the MPS memory leak alone cost
us hours). Read it top to bottom once; after that use it as a reference.

---

## Table of Contents

1. [The idea in 60 seconds](#1-the-idea-in-60-seconds)
2. [Origin: the Code2LoRA paper](#2-origin-the-code2lora-paper)
3. [Architecture](#3-architecture)
4. [The target model: Gemma-4-E2B specifics](#4-the-target-model-gemma-4-e2b-specifics)
5. [Data pipeline: the 6 views](#5-data-pipeline-the-6-views)
6. [Data pipeline: QA generation](#6-data-pipeline-qa-generation)
7. [What is learnable β€” Tier A / B / C](#7-what-is-learnable--tier-a--b--c)
8. [Datasets inventory](#8-datasets-inventory)
9. [Experiments & results](#9-experiments--results)
10. [Key decisions & lessons learned](#10-key-decisions--lessons-learned)
11. [Repository map](#11-repository-map)
12. [How to run it](#12-how-to-run-it)
13. [Evaluation methodology](#13-evaluation-methodology)
14. [Costs & budget discipline](#14-costs--budget-discipline)
15. [Current status & roadmap](#15-current-status--roadmap)
16. [Glossary](#16-glossary)

---

## 1. The idea in 60 seconds

A coding agent (Claude Code, Codex, etc.) is great at reasoning but knows nothing
about *your* repo until you paste files into its context β€” which is slow, expensive,
and capped by the context window. The usual fix is RAG (retrieve chunks at query
time). We do something different and complementary:

**We bake the repo's "personality" directly into the model's weights, once, as a
LoRA adapter β€” and we generate that adapter with a neural network instead of
training it.**

```
                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   repo on disk ──────► β”‚  6-view extractor + Qwen3    β”‚ ──► 12288-d
                        β”‚  frozen embedding encoder    β”‚     repo embedding
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜          β”‚
                                                                 β–Ό
                                                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                                   β”‚  HYPERNETWORK (our head) β”‚
                                                   β”‚  MLP β†’ per-module A,B     β”‚
                                                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                                 β”‚  LoRA weights
                                                                 β–Ό
   Q: "what layer owns auth in this repo?"  ──►  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                                 β”‚ FROZEN Gemma-4-E2B + injected β”‚ ──► "the middleware
                                                 β”‚ LoRA  (zero extra tokens)     β”‚      layer, via ..."
                                                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

The magic: **the hypernetwork is trained across hundreds of repos**, so it learns
the *mapping* `repo embedding β†’ good adapter`. At inference on a brand-new repo it
has never seen, it embeds the repo once and produces an adapter in a single forward
pass. This is the same reason the source paper needed 400+ repos, not 1: breadth is
what makes the mapping generalize.

---

## 2. Origin: the Code2LoRA paper

We reverse-engineered **Code2LoRA** (arXiv 2606.06492v1) and found its released
code (`anonymous.4open.science/r/code2lora-6857`, MIT). The paper's contribution:
a *static hypernetwork* that maps a **repository** embedding β†’ a LoRA adapter for a
frozen code LLM, evaluated on **RepoPeftBench** with IR (in-repo) / CR (cross-repo)
splits. On a full H100 setup they report **63.8 % cross-repo Exact Match**.

Our project is the **Doc2LoRA variant** the paper itself cites β€” hypernetwork maps a
*document/repo view* β†’ LoRA β€” reimplemented against **Gemma-4-E2B**, trained fully
**locally on MPS** (no CUDA/H100), and extended in two directions the paper does not
cover:

- **Memory / recall**: the adapter should let the model *recall facts* about the
  repo, not just complete code.
- **Tech-Lead judgment**: architecture, data-flow, conventions, contracts, ops β€”
  the things a 20-year senior engineer "just knows" about a codebase.

We keep the paper's proven autograd trick almost verbatim (see Β§3) and change only
what Gemma-4 and Apple Silicon force us to change.

---

## 3. Architecture

Three frozen/learned pieces. Only the middle one (the head) is trained.

### 3.1 Frozen encoder β€” `memory_lora/encoder.py`

- **Qwen3-Embedding-0.6B**, frozen, no gradient flows through it.
- Each repo view is chunked into token windows (2048 tokens, 128 overlap), each
  chunk mean-pooled, then chunks combined with **mean + max pooling** β†’ a **2048-d**
  vector *per view*.
- Embeddings are **precomputed once and cached** to parquet β€” the encoder never runs
  during training.

### 3.2 The hypernetwork head β€” `memory_lora/core.py :: MemoryLoRAHead`

The only trained component. Design (kept close to the paper):

- **2-layer GELU MLP trunk** (`input_dim β†’ hidden_dim β†’ hidden_dim`), followed by
  **L2-normalize + √hidden_dim rescale** (stabilizes the magnitude of generated
  weights).
- **Per-module-type output heads**: for each target module *type* it emits an
  `A ∈ [rank, in_features]` and `B ∈ [out_features, rank]`. **One (A,B) pair per
  type, shared across all layers of that type** β€” this is what keeps the head
  tractable (188.6 M params) instead of exploding per-layer.
- **Squashing**: `tanh(raw) * exp(log_scale)` with a learned per-type `log_scale`
  (init **-3.5**). This starts the generated adapter near-zero (so training begins
  close to the base model) and lets each type learn its own output scale.
- **Defaults**: `hidden_dim=128`, `rank=16`, `dropout=0.1`.
  - *Why hidden_dim=128 and not the paper's 512/1024?* A 745 M-param head
    (hidden_dim=512) barely moved eval loss (~1.9 β†’ ~2.7) but was far heavier to
    train on MPS. 128 cuts head size dramatically with negligible quality loss
    locally. Bump it later on real GPUs.

### 3.3 The LoRA injection β€” `memory_lora/core.py :: LoRA`

```
   base nn.Linear (FROZEN)              hypernetwork output
        β”‚                                    β”‚
   x ──►│ Wx  ────────────────┐              β”‚
        β”‚ (input detached      +──► y = Wx + scaling Β· B (A x)
        β”‚  into base)          β”‚              β–²   β–²
   x β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β”‚   β”‚
                                     A,B are NON-detached tensors so
                                     autograd flows LM-loss β†’ head
```

Critical detail (straight from the paper's code): **A and B are plain, non-buffer
tensor attributes, not `nn.Parameter` and not detached**, so the gradient of the LM
loss flows *through* the injected weights *into the hypernetwork*. The base
`nn.Linear` is frozen and its input is detached. Get this wrong and the head never
learns.

### 3.4 Shape-qualified module types β€” the heterogeneity fix

Gemma-4-E2B is **not** a uniform stack (see Β§4). Two `q_proj`s can have different
shapes. If you key the head by bare type name (`q_proj`) you get
*"type q_proj inconsistent dims"* crashes. Fix: key by **shape-qualified type**,
e.g. `q_proj_1536x2048` vs `q_proj_1536x4096`. The v2 run discovered **14 shape
types** across **205 target modules**:

```
down_proj_12288x1536  down_proj_6144x1536   gate_proj_1536x12288  gate_proj_1536x6144
k_proj_1536x256       k_proj_1536x512       o_proj_2048x1536      o_proj_4096x1536
q_proj_1536x2048      q_proj_1536x4096      up_proj_1536x12288    up_proj_1536x6144
v_proj_1536x256       v_proj_1536x512
```

`get_module_specs(root_prefix="model.language_model.")` restricts wrapping to the
text decoder β€” the **vision and audio towers are never touched** (not even
inspected), so the multimodal forward path stays intact and they cost only idle RAM.

---

## 4. The target model: Gemma-4-E2B specifics

Verified by reading the actual safetensors header, not guessed:

- **Real model.** Google shipped Gemma 4 in March 2026. Apache-2.0, ungated.
  Class `Gemma4ForConditionalGeneration`, `model_type: "gemma4"`. Loaded via
  `AutoModelForImageTextToText`.
- **Requires `transformers >= 5.5.0.dev0`** β€” install from the `main` branch, not a
  pinned PyPI release (this is the single biggest environment risk; smoke-test first).
- **Decoder is nested**: layers live at `model.language_model.layers.{i}.*`, *not*
  `model.layers.*`. The layer-index regex had to change accordingly.
- **35 text layers, heterogeneous:**
  - Aggressive KV sharing β€” **20 of 35 layers lack their own `k_proj`/`v_proj`**
    (`num_kv_shared_layers=20`).
  - **Every 5th layer is wider** (the `*_4096`, `*_12288` shape variants above).
- **Device `mps`, precision bf16** (fall back to fp16 if unstable). No
  `flash_attention_2` on MPS β€” use `sdpa`, fall back to `eager`.

---

## 5. Data pipeline: the 6 views

`scripts/build_repo_multiview.py` clones a repo and extracts **6 complementary
views**, embeds each with Qwen3 β†’ 2048-d, and **concatenates to a 12288-d** repo
vector. The views encode the different "lenses" a senior engineer uses:

| View            | What it captures                                   | Source signals |
|-----------------|----------------------------------------------------|----------------|
| `v_graph`       | call / import / dependency structure               | AST for Python (`memory_lora/codegraph.py`); `IMPORT_RE`/`DEF_RE` regex fallback for other languages |
| `v_arch`        | architecture & layout                              | README, folder tree |
| `v_history`     | how the code evolved                               | `git log`, recent diffs |
| `v_contracts`   | behavioral contracts                               | test files |
| `v_conventions` | idioms & style                                     | representative source files |
| `v_ops`         | build / deploy / runtime                           | CI config, Dockerfile, build files |

**Multi-language from the start.** `CODE_EXTS` + regex fallbacks mean the graph view
works for 9 languages, not just Python (this was a deliberate correction β€” see Β§10).
Repos with < 3 code files are skipped. The build is **resume-safe** (skips repos
already in `multiview_sources.jsonl`) and **flushes the embeddings parquet every 10
repos**, so a crash never loses more than 10 repos of work.

---

## 6. Data pipeline: QA generation

The repo embedding is the *input*; the *target* is repo-scoped Q&A. Two generators,
both parallelized (`ThreadPoolExecutor`, `--workers 10`) with a **per-prompt disk
cache** (idempotent reruns) and a `--model` flag:

- **`scripts/generate_repo_scoped_qa.py`** β€” reads the same 6 views and asks the LLM
  for **8–12 repo-level judgment questions** ("what layer owns X", "what convention
  does this repo use for Y", "how does data flow through Z", "why is it structured
  this way"). Target scope = input scope (repo-level embedding ↔ repo-level QA).
- **`scripts/generate_commitpack_qa.py`** β€” **breadth** generator: one commit per
  *distinct* repo across CommitPackFT (25k+ distinct repos), 3–4 commit-scoped
  judgment questions (why / conventions / contracts / impact). For a hypernetwork,
  **distinct-repo count is the currency of generalization**, so we favor 1 commit Γ—
  many repos over many commits Γ— one repo.

**Discipline (both):** answers are **short judgment**, never file-path/line-number
lists. This is deliberate β€” see Tier A/B/C next.

**Models used (OpenRouter, OpenAI-compatible API):**

| Model | Role | Notes | Cost |
|-------|------|-------|------|
| `google/gemini-3.6-flash` | high-quality QA | **reasoning is mandatory** β†’ needs generous `max_tokens` (3000–4000) or it returns empty | ~$0.0021 / QA |
| `google/gemma-4-31b-it` | bulk / cheap QA | non-reasoning, clean JSON | ~$0.001 / repo (~$1 per 1000 repos) |

The OpenRouter key lives in a **git-ignored `.env` (mode 600)** and is never pasted
into a command line.

---

## 7. What is learnable β€” Tier A / B / C

A LoRA adapter has finite capacity. We classify repo knowledge by whether a LoRA can
hold it β€” this drives the entire QA design:

- **Tier A β€” Judgment & conventions** (LEARNABLE). "This repo puts business logic in
  services, not views." Compressible, generalizes. β†’ **This is what we train on.**
- **Tier B β€” Structural gist** (LEARNABLE). "Auth flows through middleware." The kind
  of thing, not the exact file.
- **Tier C β€” Exact recall & multi-hop** (NOT reliably learnable). "Line 412 of
  `foo.py` calls `bar()`." This needs **retrieval (RAG)**, not weights.

So Memory-LoRA and RAG are **complementary**: the adapter carries Tier A/B judgment
for free (zero tokens); RAG handles Tier C exact lookups. The QA prompts forbid exact
file/line answers precisely so we never ask the LoRA to do a job it structurally
can't.

---

## 8. Datasets inventory

Everything lives under `data/` (git-ignored blobs). Sizes are approximate.

| Path | What | Scale |
|------|------|-------|
| `data/real_code2lora/` | **RepoPeftBench** from the `code2lora` HF org β€” 500 Python repos, repo-commit embeddings + diffs | 73,849 repo-commit rows; ~1.2 GB |
| `data/commitpack/multilang_commits.jsonl` | **CommitPackFT** shards, 9 languages | 25k+ distinct repos |
| `data/docs/multiview_sources.jsonl` | 6-view `view_text` per repo (input to QA gen) | growing (1000s of repos) |
| `data/embeddings/multiview_embeddings.parquet` | 12288-d multi-view repo embeddings | 1000+ repos |
| `data/embeddings/aligned6_embeddings.parquet` | **assembled training inputs** (repos with β‰₯1 QA) | 1058 repos (current) |
| `data/qna/repo_scoped_qa.jsonl` | repo-level judgment QA | 11,232 QA |
| `data/qna/techlead_qa_commitpack.jsonl` | commit-scoped breadth QA | 9,245 QA |
| `data/qna/techlead_qa.jsonl` | SWE-bench tech-lead QA | 2,786 QA |
| `data/qna/aligned6_qna.jsonl` | **assembled training targets** | 8,540 QA (current) |
| `data/openrouter_cache/` | per-prompt response cache | ~19 MB |

**Language balancing.** SWE-bench is ~79 % Django. Left alone, the dataset was 46 %
Django. `scripts/consolidate_qa.py` applies a **per-repo cap** (default 12–15 QA/repo)
which collapses Django to **~2.0 %** while preserving the 2400+ distinct repos'
diversity. `assemble_6view_dataset.py` applies the same cap when building the final
aligned set.

---

## 9. Experiments & results

Chronological, with the actual numbers we measured. Two families of runs.

### 9.1 Reproducing the paper (single-view, real RepoPeftBench)

| Run | What | Result |
|-----|------|--------|
| `full1` (early) | first end-to-end hypernetwork on converted real data | CR **EM 0.056–0.083**, EditSim ~0.27 β€” pipeline works, undertrained |
| `sixview`/converted-real (best single-view ckpt) | after more training | **CR EM 0.524, EditSim 0.635** |
| Paper (reference, H100) | their full run | CR EM **0.638** |

**Headline:** on real code, after only ~2.4 % of one epoch of local MPS training, we
reached **52.4 % cross-repo Exact Match** vs the paper's 63.8 % on a full H100 setup.
The mechanism demonstrably works β€” the generated adapter recovers repo-specific
identifiers the base model does not know.

### 9.2 The 6-view Tech-Lead model (the current line of work)

Loss is causal-LM cross-entropy on QA targets; lower is better. Three eval suites:
`cr_val` / `cr_test` (held-out *repos*) and `ir_test` (held-out *QA* of train repos).

| Run | Dataset | Best held-out `cr_test` loss | Notes |
|-----|---------|------------------------------|-------|
| `sixview_v1` | 515 repos / 3,988 QA (415 train repos) | **2.848** (step ~1060) | Overfit afterward: train loss fell to 1.75 while `cr_test` drifted to 3.35. Classic small-dataset ceiling. |
| `sixview_v2` | **1,058 repos / 8,540 QA (858 train repos)** | *in progress* | Resumed from `sixview_v1/head.best.pt`; 2Γ— the data specifically to break v1's ceiling. |

`sixview_v1` metrics trajectory (from `runs/sixview_v1/metrics.jsonl`):

```
step 1245  cr_test 2.962   ir_test 2.535   (end of epoch 2 β€” near best)
step 1400  cr_test 3.245   ir_test 2.593   (overfitting begins)
step 1600  cr_test 3.352   ir_test 2.655   (train loss still falling β†’ ceiling hit)
```

The v1β†’v2 story is the core empirical lesson: **the small aligned set was the
bottleneck, not the architecture** β€” hence the push to build 1000+ more repos.

---

## 10. Key decisions & lessons learned

The expensive knowledge. Read this section twice.

### 10.1 ⚠️ The MPS gradient-checkpointing memory leak (the big one)

**Symptom:** training with `gradient_checkpointing_enable()` (`use_reentrant=False`)
**leaked ~12 GB per step** and OOM'd the whole machine within a few steps.

**Diagnosis** (`scripts/diag_mps_leak.py`): forward-only was stable; train + checkpoint
leaked 39 GB β†’ 18 GB free in 2 steps. Isolated the checkpointing path as the cause.

**Fix:** **`--no-gradient-checkpointing`.** We have enough unified memory to hold
activations without it once the multimodal towers sit idle. This is documented as a
standing memory (`mps-gradient-checkpointing-leak.md`).

### 10.2 ⚠️ `psutil` RSS is blind to MPS memory

Our first memory safety-net used `psutil` RSS / `ps -o rss` β€” it reported **< 1 GB**
while `top` showed **55–83 GB** actually in use. MPS (GPU) allocations don't show up
in process RSS.

**Fix:** the safety check uses **`psutil.virtual_memory().available`** (system-wide)
with a `--min-available-gb` floor (default 5). To *observe* MPS memory, use
`top -l 1 -pid <PID> -stats mem`, not `ps`.

### 10.3 Memory competition between concurrent jobs

Three concurrent jobs once pushed available memory under the 10 GB floor and training
self-stopped. **Lesson:** during MPS training, run data builds/embedding on **CPU**
(`--device cpu`) so they don't contend for the GPU/unified memory. We now routinely
run training (MPS) + QA gen (network) + multiview build (CPU) together without
contention.

### 10.4 Don't lose hours of training

Every long run writes **checkpoints every 50 steps** (overwriting `head.latest.pt`),
**every 30 minutes** (timestamped `head.tNNNNm.pt`), **per-epoch** (`head.epN.pt`),
and a **`head.best.pt`** on eval improvement. Runs are launched with `nohup … &
disown` so they survive terminal/session death. `sixview_v1` in fact survived a full
session interruption and kept training. Resume with `--resume-from <ckpt>` (loads head
weights; optimizer restarts fresh).

### 10.5 Data-quality corrections (user-driven)

- **"I still see lots of Django."** SWE-bench is Django-dominated. β†’ per-repo cap +
  multi-language sourcing dropped Django 46 % β†’ 2.0 %.
- **"It must be good for any programming language."** β†’ 9-language diversity via
  CommitPackFT and language-agnostic view extraction.
- **"Where's the code context in the QA?"** β†’ clarified the two-channel design: the
  **repo embedding is the context channel**, the QA is only the target. They are
  joined by `doc_id` at assembly time.

### 10.6 OpenRouter gotchas

- `gemini-3.6-flash` **returned empty** until we raised `max_tokens` β€” it's a
  mandatory-reasoning model that spends tokens on hidden reasoning before content.
  Reasoning **cannot be disabled** (400 error).
- CommitPackFT's HF loader is deprecated β†’ fetch raw `data.jsonl` directly.
- `global MODEL` after use is a `SyntaxError` β†’ set via `globals()["MODEL"] = ...`.

### 10.7 Performance fix worth knowing

Loading embeddings was 5+ min because `_list_to_f32_array` used a Python loop.
Vectorized via `col.combine_chunks().flatten().to_numpy()` β†’ **~220Γ— faster**.

---

## 11. Repository map

```
memory_lora/                  # the library (importable package)
  core.py                     # LoRA wrapper, MemoryLoRAHead hypernetwork,
                              #   get_module_specs / replace_with_lora / inject_lora_weights,
                              #   load_doc_rows / load_qna_rows
  encoder.py                  # Qwen3 chunk + embed + mean/max pool (frozen)
  codegraph.py                # Python AST extractor (imports, sigs, call graph)
  data_paths.py               # local parquet/jsonl path resolver

scripts/
  build_repo_multiview.py     # clone β†’ 6 views β†’ 12288-d embeddings (multi-language, resume-safe)
  generate_repo_scoped_qa.py  # repo-level judgment QA (aligned to the 6 views)
  generate_commitpack_qa.py   # commit-scoped breadth QA across 1000s of distinct repos
  generate_techlead_qa.py     # SWE-bench tech-lead QA
  generate_synthetic_dataset.py  # original synthetic doc + QA generator
  consolidate_qa.py           # per-repo cap β†’ language/domain balancing
  assemble_6view_dataset.py   # join embeddings ↔ all QA by repo β†’ aligned6_{embeddings,qna}
  augment_paraphrases.py      # QA paraphrase augmentation
  convert_real_code2lora.py   # RepoPeftBench β†’ our schema
  build_doc_embeddings.py     # encoder pass over documents
  merge_corpora.py            # combine multiple corpora
  train_memory_lora.py        # THE trainer (MPS, one-repo-per-step, checkpoints, TensorBoard)
  train_direct_lora.py        # baseline: plain per-repo LoRA (no hypernetwork)
  eval_memory_lora.py         # EM / EditSim recall eval on cr/ir splits
  show_eval_examples.py       # dump concrete base-vs-adapted examples
  test_embed_this_repo.py     # embed the current repo (pipeline demo)
  test_recall_this_repo.py    # query the adapted model about this repo
  diag_mps_leak.py            # the memory-leak isolation harness

data/                         # git-ignored: embeddings, qna, sources, caches
runs/                         # git-ignored: checkpoints, logs, metrics.jsonl, tb/
requirements.txt              # torch 2.13 (MPS), transformers@main, openai, pyarrow, tensorboard…
.env                          # git-ignored, mode 600: OPENROUTER_API_KEY
```

---

## 12. How to run it

### Setup

```bash
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt          # installs transformers from git main
echo "OPENROUTER_API_KEY=sk-or-..." > .env && chmod 600 .env
```

Smoke-test the environment first (gates everything): confirm `transformers` main
loads `google/gemma-4-E2B` on `mps` and runs a text-only forward pass.

### Build data

```bash
# 1) multi-view embeddings for a repo list (CPU to stay off the GPU during training)
python scripts/build_repo_multiview.py \
  --repos-file data/multilang_repo_list.txt --max-repos 1000 --device cpu

# 2) repo-scoped QA (cheap model) β€” appends, resume-safe, cached
./venv/bin/python scripts/generate_repo_scoped_qa.py \
  --model google/gemma-4-31b-it --workers 10

# 3) balance + assemble the aligned training set
python scripts/consolidate_qa.py --per-repo-cap 12
python scripts/assemble_6view_dataset.py
# -> data/embeddings/aligned6_embeddings.parquet + data/qna/aligned6_qna.jsonl
```

### Train (the exact `sixview_v2` command)

```bash
nohup ./venv/bin/python scripts/train_memory_lora.py --output-dir sixview_v2 \
  --resume-from runs/sixview_v1/head.best.pt \
  --embeddings-path data/embeddings/aligned6_embeddings.parquet \
  --qna-path data/qna/aligned6_qna.jsonl --epochs 100 --max-hours 8 \
  --checkpoint-every-steps 50 --checkpoint-every-minutes 30 --epoch-ckpt-every 5 \
  --eval-every-steps 300 --limit-eval-docs 40 --max-seq-len 512 --fixed-seq-len \
  --max-qna-per-doc 12 --lm-micro-batch 2 --device mps --no-gradient-checkpointing \
  --rank 16 --head-hidden-dim 128 --head-dropout 0.1 --weight-decay 0.05 \
  --early-stop-patience 25 --lr 8e-5 --lr-total-steps 9000 --min-available-gb 5 \
  > runs/sixview_v2_train.log 2>&1 &
disown
```

**Flags you must not forget:** `--no-gradient-checkpointing` (the leak),
`--device cpu` for builds during training (contention), `--min-available-gb`
(system-wide memory floor).

### Watch it

```bash
tensorboard --logdir runs/sixview_v2/tb    # train/loss, train/lr, eval/{suite}_loss
tail -f runs/sixview_v2_train.log
```

### Evaluate & inspect

```bash
python scripts/eval_memory_lora.py --ckpt runs/sixview_v2/head.best.pt   # EM / EditSim
python scripts/show_eval_examples.py                                     # base vs adapted
```

---

## 13. Evaluation methodology

- **Splits (deterministic, by `md5(repo) % 100`):** 80 % train / 10 % `cr_val` /
  10 % `cr_test` **by repo**, so cross-repo suites are **entirely held-out
  repositories** the hypernetwork never trained on. Within train repos, ~15 % of QA
  is held out β†’ `ir_test` (in-repo generalization to unseen questions of seen repos).
- **Metrics:** causal-LM **eval loss** during training (fast, every N steps on
  `--limit-eval-docs` docs to stay cheap on CPU), plus generation-time **Exact Match
  (EM)** and **EditSim** for the recall eval.
- **The proof spot-check:** query the *adapted* model with repo-specific questions and
  confirm the *base* (un-adapted) model gets them wrong/vague β€” proving the
  **adapter**, not the base model's pretraining, does the work.

CPU eval of a float32 5B model is slow (~20 min for a full pass) β†’ we cap eval docs
(e.g. 10–40) for in-loop evals and run full EM eval separately.

---

## 14. Costs & budget discipline

- **Spend baseline:** $31.00 (`runs/spend_baseline.txt`); ~**$33.90 total** to date;
  ~**$11 remaining**. The project is run under explicit budget caps ("spend at most
  $4 more") with spend-guards.
- **Unit economics:** `gemini-3.6-flash` β‰ˆ **$0.0021/QA**; `gemma-4-31b-it` β‰ˆ
  **$0.001/repo (~$1 per 1000 repos)** β€” which is exactly why the 1000-repo expansion
  uses the gemma model.
- **Free levers:** the per-prompt cache makes reruns free; embedding and training are
  local (electricity only).

---

## 15. Current status & roadmap

**Live right now (three jobs in parallel, no contention):**

- **`sixview_v2` training** β€” resumed from `head.best.pt` on the doubled
  **1,058-repo / 8,540-QA** dataset (858 steps/epoch), MPS. First held-out eval at
  step 300 tells us whether doubling the data broke v1's 2.848 ceiling.
- **QA generation** β€” `gemma-4-31b-it` filling in all ~1,032 new repos for the
  **complete** dataset (next training run).
- **Multiview build** β€” cloning/embedding toward the full +1,000-new target (CPU).

**Roadmap:**

1. Finish the complete 1000-new-repo dataset (embeddings + QA).
2. Assemble the full aligned set (~1,650 repos) and train **`sixview_v3`** on it.
3. Run generation-time **EM/EditSim** on the 6-view model (base vs adapted).
4. Push `head_hidden_dim` back up once on real GPUs; the 128 default was an
   MPS-locality compromise.
5. Broaden Tier-A/B QA toward agent-harness use cases (Jira/ticket tracking,
   diff/impact reasoning) already scaffolded in `generate_techlead_qa.py`.

**Open questions:**

- Does the 12288-d 6-view embedding actually beat the single 2048-d view on
  generation EM, or only on loss? (loss says yes; EM eval pending)
- What's the real Tier-B ceiling β€” how much structural gist fits in rank-16?
- Optimal per-repo QA cap for the breadth/depth trade-off.

---

## 16. Glossary

- **Hypernetwork** β€” a network that outputs the weights of another network. Here:
  repo embedding β†’ LoRA (A,B) matrices.
- **LoRA** β€” Low-Rank Adaptation: `y = Wx + scaling Β· B(Ax)`, with `A,B` low-rank
  (rank 16). We *generate* A,B instead of training them per-repo.
- **6 views** β€” graph / arch / history / contracts / conventions / ops; each 2048-d,
  concatenated to 12288-d.
- **CR / IR** β€” cross-repo (held-out repos) / in-repo (held-out QA of seen repos).
- **EM / EditSim** β€” Exact Match / edit-distance similarity of generated vs gold.
- **Tier A/B/C** β€” judgment (learnable) / structural gist (learnable) / exact recall
  (needs RAG).
- **MPS** β€” Apple's Metal Performance Shaders GPU backend for PyTorch.
- **RepoPeftBench** β€” the paper's benchmark; 500 Python repos, repo-commit
  embeddings + diffs, IR/CR splits.

---

*Maintained as living documentation. If you change a default, a path, or a flag,
update the matching section here β€” onboarding depends on it.*