memory-lora-gemma4 / docs /DEEP_DIVE.md
moncefem's picture
Rewrite the model card
ef8744f
|
Raw
History Blame Contribute Delete
29.2 kB
# Memory-LoRA: A Hypernetwork that Writes Repo-Specific Adapters for Gemma-4-E2B
> **One line:** we train a small neural network (a *hypernetwork*) that reads an
> embedding of a codebase and **emits a LoRA adapter** for a frozen
> `google/gemma-4-E2B`. The adapter injects repo-specific knowledge into the
> model with **zero extra tokens at inference time** β€” no RAG, no context stuffing.
> Everything runs **locally on an Apple-Silicon Mac** (M4 Pro, 64 GB).
This document is the onboarding bible for the project. It covers *what* we built,
*why* every non-obvious choice was made, *what data* we used, *every experiment we
ran and its result*, and the **hard-won lessons** (the MPS memory leak alone cost
us hours). Read it top to bottom once; after that use it as a reference.
---
## Table of Contents
1. [The idea in 60 seconds](#1-the-idea-in-60-seconds)
2. [Origin: the Code2LoRA paper](#2-origin-the-code2lora-paper)
3. [Architecture](#3-architecture)
4. [The target model: Gemma-4-E2B specifics](#4-the-target-model-gemma-4-e2b-specifics)
5. [Data pipeline: the 6 views](#5-data-pipeline-the-6-views)
6. [Data pipeline: QA generation](#6-data-pipeline-qa-generation)
7. [What is learnable β€” Tier A / B / C](#7-what-is-learnable--tier-a--b--c)
8. [Datasets inventory](#8-datasets-inventory)
9. [Experiments & results](#9-experiments--results)
10. [Key decisions & lessons learned](#10-key-decisions--lessons-learned)
11. [Repository map](#11-repository-map)
12. [How to run it](#12-how-to-run-it)
13. [Evaluation methodology](#13-evaluation-methodology)
14. [Costs & budget discipline](#14-costs--budget-discipline)
15. [Current status & roadmap](#15-current-status--roadmap)
16. [Glossary](#16-glossary)
---
## 1. The idea in 60 seconds
A coding agent (Claude Code, Codex, etc.) is great at reasoning but knows nothing
about *your* repo until you paste files into its context β€” which is slow, expensive,
and capped by the context window. The usual fix is RAG (retrieve chunks at query
time). We do something different and complementary:
**We bake the repo's "personality" directly into the model's weights, once, as a
LoRA adapter β€” and we generate that adapter with a neural network instead of
training it.**
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
repo on disk ──────► β”‚ 6-view extractor + Qwen3 β”‚ ──► 12288-d
β”‚ frozen embedding encoder β”‚ repo embedding
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ HYPERNETWORK (our head) β”‚
β”‚ MLP β†’ per-module A,B β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ LoRA weights
β–Ό
Q: "what layer owns auth in this repo?" ──► β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ FROZEN Gemma-4-E2B + injected β”‚ ──► "the middleware
β”‚ LoRA (zero extra tokens) β”‚ layer, via ..."
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
The magic: **the hypernetwork is trained across hundreds of repos**, so it learns
the *mapping* `repo embedding β†’ good adapter`. At inference on a brand-new repo it
has never seen, it embeds the repo once and produces an adapter in a single forward
pass. This is the same reason the source paper needed 400+ repos, not 1: breadth is
what makes the mapping generalize.
---
## 2. Origin: the Code2LoRA paper
We reverse-engineered **Code2LoRA** (arXiv 2606.06492v1) and found its released
code (`anonymous.4open.science/r/code2lora-6857`, MIT). The paper's contribution:
a *static hypernetwork* that maps a **repository** embedding β†’ a LoRA adapter for a
frozen code LLM, evaluated on **RepoPeftBench** with IR (in-repo) / CR (cross-repo)
splits. On a full H100 setup they report **63.8 % cross-repo Exact Match**.
Our project is the **Doc2LoRA variant** the paper itself cites β€” hypernetwork maps a
*document/repo view* β†’ LoRA β€” reimplemented against **Gemma-4-E2B**, trained fully
**locally on MPS** (no CUDA/H100), and extended in two directions the paper does not
cover:
- **Memory / recall**: the adapter should let the model *recall facts* about the
repo, not just complete code.
- **Tech-Lead judgment**: architecture, data-flow, conventions, contracts, ops β€”
the things a 20-year senior engineer "just knows" about a codebase.
We keep the paper's proven autograd trick almost verbatim (see Β§3) and change only
what Gemma-4 and Apple Silicon force us to change.
---
## 3. Architecture
Three frozen/learned pieces. Only the middle one (the head) is trained.
### 3.1 Frozen encoder β€” `memory_lora/encoder.py`
- **Qwen3-Embedding-0.6B**, frozen, no gradient flows through it.
- Each repo view is chunked into token windows (2048 tokens, 128 overlap), each
chunk mean-pooled, then chunks combined with **mean + max pooling** β†’ a **2048-d**
vector *per view*.
- Embeddings are **precomputed once and cached** to parquet β€” the encoder never runs
during training.
### 3.2 The hypernetwork head β€” `memory_lora/core.py :: MemoryLoRAHead`
The only trained component. Design (kept close to the paper):
- **2-layer GELU MLP trunk** (`input_dim β†’ hidden_dim β†’ hidden_dim`), followed by
**L2-normalize + √hidden_dim rescale** (stabilizes the magnitude of generated
weights).
- **Per-module-type output heads**: for each target module *type* it emits an
`A ∈ [rank, in_features]` and `B ∈ [out_features, rank]`. **One (A,B) pair per
type, shared across all layers of that type** β€” this is what keeps the head
tractable (188.6 M params) instead of exploding per-layer.
- **Squashing**: `tanh(raw) * exp(log_scale)` with a learned per-type `log_scale`
(init **-3.5**). This starts the generated adapter near-zero (so training begins
close to the base model) and lets each type learn its own output scale.
- **Defaults**: `hidden_dim=128`, `rank=16`, `dropout=0.1`.
- *Why hidden_dim=128 and not the paper's 512/1024?* A 745 M-param head
(hidden_dim=512) barely moved eval loss (~1.9 β†’ ~2.7) but was far heavier to
train on MPS. 128 cuts head size dramatically with negligible quality loss
locally. Bump it later on real GPUs.
### 3.3 The LoRA injection β€” `memory_lora/core.py :: LoRA`
```
base nn.Linear (FROZEN) hypernetwork output
β”‚ β”‚
x ──►│ Wx ────────────────┐ β”‚
β”‚ (input detached +──► y = Wx + scaling Β· B (A x)
β”‚ into base) β”‚ β–² β–²
x β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚
A,B are NON-detached tensors so
autograd flows LM-loss β†’ head
```
Critical detail (straight from the paper's code): **A and B are plain, non-buffer
tensor attributes, not `nn.Parameter` and not detached**, so the gradient of the LM
loss flows *through* the injected weights *into the hypernetwork*. The base
`nn.Linear` is frozen and its input is detached. Get this wrong and the head never
learns.
### 3.4 Shape-qualified module types β€” the heterogeneity fix
Gemma-4-E2B is **not** a uniform stack (see Β§4). Two `q_proj`s can have different
shapes. If you key the head by bare type name (`q_proj`) you get
*"type q_proj inconsistent dims"* crashes. Fix: key by **shape-qualified type**,
e.g. `q_proj_1536x2048` vs `q_proj_1536x4096`. The v2 run discovered **14 shape
types** across **205 target modules**:
```
down_proj_12288x1536 down_proj_6144x1536 gate_proj_1536x12288 gate_proj_1536x6144
k_proj_1536x256 k_proj_1536x512 o_proj_2048x1536 o_proj_4096x1536
q_proj_1536x2048 q_proj_1536x4096 up_proj_1536x12288 up_proj_1536x6144
v_proj_1536x256 v_proj_1536x512
```
`get_module_specs(root_prefix="model.language_model.")` restricts wrapping to the
text decoder β€” the **vision and audio towers are never touched** (not even
inspected), so the multimodal forward path stays intact and they cost only idle RAM.
---
## 4. The target model: Gemma-4-E2B specifics
Verified by reading the actual safetensors header, not guessed:
- **Real model.** Google shipped Gemma 4 in March 2026. Apache-2.0, ungated.
Class `Gemma4ForConditionalGeneration`, `model_type: "gemma4"`. Loaded via
`AutoModelForImageTextToText`.
- **Requires `transformers >= 5.5.0.dev0`** β€” install from the `main` branch, not a
pinned PyPI release (this is the single biggest environment risk; smoke-test first).
- **Decoder is nested**: layers live at `model.language_model.layers.{i}.*`, *not*
`model.layers.*`. The layer-index regex had to change accordingly.
- **35 text layers, heterogeneous:**
- Aggressive KV sharing β€” **20 of 35 layers lack their own `k_proj`/`v_proj`**
(`num_kv_shared_layers=20`).
- **Every 5th layer is wider** (the `*_4096`, `*_12288` shape variants above).
- **Device `mps`, precision bf16** (fall back to fp16 if unstable). No
`flash_attention_2` on MPS β€” use `sdpa`, fall back to `eager`.
---
## 5. Data pipeline: the 6 views
`scripts/build_repo_multiview.py` clones a repo and extracts **6 complementary
views**, embeds each with Qwen3 β†’ 2048-d, and **concatenates to a 12288-d** repo
vector. The views encode the different "lenses" a senior engineer uses:
| View | What it captures | Source signals |
|-----------------|----------------------------------------------------|----------------|
| `v_graph` | call / import / dependency structure | AST for Python (`memory_lora/codegraph.py`); `IMPORT_RE`/`DEF_RE` regex fallback for other languages |
| `v_arch` | architecture & layout | README, folder tree |
| `v_history` | how the code evolved | `git log`, recent diffs |
| `v_contracts` | behavioral contracts | test files |
| `v_conventions` | idioms & style | representative source files |
| `v_ops` | build / deploy / runtime | CI config, Dockerfile, build files |
**Multi-language from the start.** `CODE_EXTS` + regex fallbacks mean the graph view
works for 9 languages, not just Python (this was a deliberate correction β€” see Β§10).
Repos with < 3 code files are skipped. The build is **resume-safe** (skips repos
already in `multiview_sources.jsonl`) and **flushes the embeddings parquet every 10
repos**, so a crash never loses more than 10 repos of work.
---
## 6. Data pipeline: QA generation
The repo embedding is the *input*; the *target* is repo-scoped Q&A. Two generators,
both parallelized (`ThreadPoolExecutor`, `--workers 10`) with a **per-prompt disk
cache** (idempotent reruns) and a `--model` flag:
- **`scripts/generate_repo_scoped_qa.py`** β€” reads the same 6 views and asks the LLM
for **8–12 repo-level judgment questions** ("what layer owns X", "what convention
does this repo use for Y", "how does data flow through Z", "why is it structured
this way"). Target scope = input scope (repo-level embedding ↔ repo-level QA).
- **`scripts/generate_commitpack_qa.py`** β€” **breadth** generator: one commit per
*distinct* repo across CommitPackFT (25k+ distinct repos), 3–4 commit-scoped
judgment questions (why / conventions / contracts / impact). For a hypernetwork,
**distinct-repo count is the currency of generalization**, so we favor 1 commit Γ—
many repos over many commits Γ— one repo.
**Discipline (both):** answers are **short judgment**, never file-path/line-number
lists. This is deliberate β€” see Tier A/B/C next.
**Models used (OpenRouter, OpenAI-compatible API):**
| Model | Role | Notes | Cost |
|-------|------|-------|------|
| `google/gemini-3.6-flash` | high-quality QA | **reasoning is mandatory** β†’ needs generous `max_tokens` (3000–4000) or it returns empty | ~$0.0021 / QA |
| `google/gemma-4-31b-it` | bulk / cheap QA | non-reasoning, clean JSON | ~$0.001 / repo (~$1 per 1000 repos) |
The OpenRouter key lives in a **git-ignored `.env` (mode 600)** and is never pasted
into a command line.
---
## 7. What is learnable β€” Tier A / B / C
A LoRA adapter has finite capacity. We classify repo knowledge by whether a LoRA can
hold it β€” this drives the entire QA design:
- **Tier A β€” Judgment & conventions** (LEARNABLE). "This repo puts business logic in
services, not views." Compressible, generalizes. β†’ **This is what we train on.**
- **Tier B β€” Structural gist** (LEARNABLE). "Auth flows through middleware." The kind
of thing, not the exact file.
- **Tier C β€” Exact recall & multi-hop** (NOT reliably learnable). "Line 412 of
`foo.py` calls `bar()`." This needs **retrieval (RAG)**, not weights.
So Memory-LoRA and RAG are **complementary**: the adapter carries Tier A/B judgment
for free (zero tokens); RAG handles Tier C exact lookups. The QA prompts forbid exact
file/line answers precisely so we never ask the LoRA to do a job it structurally
can't.
---
## 8. Datasets inventory
Everything lives under `data/` (git-ignored blobs). Sizes are approximate.
| Path | What | Scale |
|------|------|-------|
| `data/real_code2lora/` | **RepoPeftBench** from the `code2lora` HF org β€” 500 Python repos, repo-commit embeddings + diffs | 73,849 repo-commit rows; ~1.2 GB |
| `data/commitpack/multilang_commits.jsonl` | **CommitPackFT** shards, 9 languages | 25k+ distinct repos |
| `data/docs/multiview_sources.jsonl` | 6-view `view_text` per repo (input to QA gen) | growing (1000s of repos) |
| `data/embeddings/multiview_embeddings.parquet` | 12288-d multi-view repo embeddings | 1000+ repos |
| `data/embeddings/aligned6_embeddings.parquet` | **assembled training inputs** (repos with β‰₯1 QA) | 1058 repos (current) |
| `data/qna/repo_scoped_qa.jsonl` | repo-level judgment QA | 11,232 QA |
| `data/qna/techlead_qa_commitpack.jsonl` | commit-scoped breadth QA | 9,245 QA |
| `data/qna/techlead_qa.jsonl` | SWE-bench tech-lead QA | 2,786 QA |
| `data/qna/aligned6_qna.jsonl` | **assembled training targets** | 8,540 QA (current) |
| `data/openrouter_cache/` | per-prompt response cache | ~19 MB |
**Language balancing.** SWE-bench is ~79 % Django. Left alone, the dataset was 46 %
Django. `scripts/consolidate_qa.py` applies a **per-repo cap** (default 12–15 QA/repo)
which collapses Django to **~2.0 %** while preserving the 2400+ distinct repos'
diversity. `assemble_6view_dataset.py` applies the same cap when building the final
aligned set.
---
## 9. Experiments & results
Chronological, with the actual numbers we measured. Two families of runs.
### 9.1 Reproducing the paper (single-view, real RepoPeftBench)
| Run | What | Result |
|-----|------|--------|
| `full1` (early) | first end-to-end hypernetwork on converted real data | CR **EM 0.056–0.083**, EditSim ~0.27 β€” pipeline works, undertrained |
| `sixview`/converted-real (best single-view ckpt) | after more training | **CR EM 0.524, EditSim 0.635** |
| Paper (reference, H100) | their full run | CR EM **0.638** |
**Headline:** on real code, after only ~2.4 % of one epoch of local MPS training, we
reached **52.4 % cross-repo Exact Match** vs the paper's 63.8 % on a full H100 setup.
The mechanism demonstrably works β€” the generated adapter recovers repo-specific
identifiers the base model does not know.
### 9.2 The 6-view Tech-Lead model (the current line of work)
Loss is causal-LM cross-entropy on QA targets; lower is better. Three eval suites:
`cr_val` / `cr_test` (held-out *repos*) and `ir_test` (held-out *QA* of train repos).
| Run | Dataset | Best held-out `cr_test` loss | Notes |
|-----|---------|------------------------------|-------|
| `sixview_v1` | 515 repos / 3,988 QA (415 train repos) | **2.848** (step ~1060) | Overfit afterward: train loss fell to 1.75 while `cr_test` drifted to 3.35. Classic small-dataset ceiling. |
| `sixview_v2` | **1,058 repos / 8,540 QA (858 train repos)** | *in progress* | Resumed from `sixview_v1/head.best.pt`; 2Γ— the data specifically to break v1's ceiling. |
`sixview_v1` metrics trajectory (from `runs/sixview_v1/metrics.jsonl`):
```
step 1245 cr_test 2.962 ir_test 2.535 (end of epoch 2 β€” near best)
step 1400 cr_test 3.245 ir_test 2.593 (overfitting begins)
step 1600 cr_test 3.352 ir_test 2.655 (train loss still falling β†’ ceiling hit)
```
The v1β†’v2 story is the core empirical lesson: **the small aligned set was the
bottleneck, not the architecture** β€” hence the push to build 1000+ more repos.
---
## 10. Key decisions & lessons learned
The expensive knowledge. Read this section twice.
### 10.1 ⚠️ The MPS gradient-checkpointing memory leak (the big one)
**Symptom:** training with `gradient_checkpointing_enable()` (`use_reentrant=False`)
**leaked ~12 GB per step** and OOM'd the whole machine within a few steps.
**Diagnosis** (`scripts/diag_mps_leak.py`): forward-only was stable; train + checkpoint
leaked 39 GB β†’ 18 GB free in 2 steps. Isolated the checkpointing path as the cause.
**Fix:** **`--no-gradient-checkpointing`.** We have enough unified memory to hold
activations without it once the multimodal towers sit idle. This is documented as a
standing memory (`mps-gradient-checkpointing-leak.md`).
### 10.2 ⚠️ `psutil` RSS is blind to MPS memory
Our first memory safety-net used `psutil` RSS / `ps -o rss` β€” it reported **< 1 GB**
while `top` showed **55–83 GB** actually in use. MPS (GPU) allocations don't show up
in process RSS.
**Fix:** the safety check uses **`psutil.virtual_memory().available`** (system-wide)
with a `--min-available-gb` floor (default 5). To *observe* MPS memory, use
`top -l 1 -pid <PID> -stats mem`, not `ps`.
### 10.3 Memory competition between concurrent jobs
Three concurrent jobs once pushed available memory under the 10 GB floor and training
self-stopped. **Lesson:** during MPS training, run data builds/embedding on **CPU**
(`--device cpu`) so they don't contend for the GPU/unified memory. We now routinely
run training (MPS) + QA gen (network) + multiview build (CPU) together without
contention.
### 10.4 Don't lose hours of training
Every long run writes **checkpoints every 50 steps** (overwriting `head.latest.pt`),
**every 30 minutes** (timestamped `head.tNNNNm.pt`), **per-epoch** (`head.epN.pt`),
and a **`head.best.pt`** on eval improvement. Runs are launched with `nohup … &
disown` so they survive terminal/session death. `sixview_v1` in fact survived a full
session interruption and kept training. Resume with `--resume-from <ckpt>` (loads head
weights; optimizer restarts fresh).
### 10.5 Data-quality corrections (user-driven)
- **"I still see lots of Django."** SWE-bench is Django-dominated. β†’ per-repo cap +
multi-language sourcing dropped Django 46 % β†’ 2.0 %.
- **"It must be good for any programming language."** β†’ 9-language diversity via
CommitPackFT and language-agnostic view extraction.
- **"Where's the code context in the QA?"** β†’ clarified the two-channel design: the
**repo embedding is the context channel**, the QA is only the target. They are
joined by `doc_id` at assembly time.
### 10.6 OpenRouter gotchas
- `gemini-3.6-flash` **returned empty** until we raised `max_tokens` β€” it's a
mandatory-reasoning model that spends tokens on hidden reasoning before content.
Reasoning **cannot be disabled** (400 error).
- CommitPackFT's HF loader is deprecated β†’ fetch raw `data.jsonl` directly.
- `global MODEL` after use is a `SyntaxError` β†’ set via `globals()["MODEL"] = ...`.
### 10.7 Performance fix worth knowing
Loading embeddings was 5+ min because `_list_to_f32_array` used a Python loop.
Vectorized via `col.combine_chunks().flatten().to_numpy()` β†’ **~220Γ— faster**.
---
## 11. Repository map
```
memory_lora/ # the library (importable package)
core.py # LoRA wrapper, MemoryLoRAHead hypernetwork,
# get_module_specs / replace_with_lora / inject_lora_weights,
# load_doc_rows / load_qna_rows
encoder.py # Qwen3 chunk + embed + mean/max pool (frozen)
codegraph.py # Python AST extractor (imports, sigs, call graph)
data_paths.py # local parquet/jsonl path resolver
scripts/
build_repo_multiview.py # clone β†’ 6 views β†’ 12288-d embeddings (multi-language, resume-safe)
generate_repo_scoped_qa.py # repo-level judgment QA (aligned to the 6 views)
generate_commitpack_qa.py # commit-scoped breadth QA across 1000s of distinct repos
generate_techlead_qa.py # SWE-bench tech-lead QA
generate_synthetic_dataset.py # original synthetic doc + QA generator
consolidate_qa.py # per-repo cap β†’ language/domain balancing
assemble_6view_dataset.py # join embeddings ↔ all QA by repo β†’ aligned6_{embeddings,qna}
augment_paraphrases.py # QA paraphrase augmentation
convert_real_code2lora.py # RepoPeftBench β†’ our schema
build_doc_embeddings.py # encoder pass over documents
merge_corpora.py # combine multiple corpora
train_memory_lora.py # THE trainer (MPS, one-repo-per-step, checkpoints, TensorBoard)
train_direct_lora.py # baseline: plain per-repo LoRA (no hypernetwork)
eval_memory_lora.py # EM / EditSim recall eval on cr/ir splits
show_eval_examples.py # dump concrete base-vs-adapted examples
test_embed_this_repo.py # embed the current repo (pipeline demo)
test_recall_this_repo.py # query the adapted model about this repo
diag_mps_leak.py # the memory-leak isolation harness
data/ # git-ignored: embeddings, qna, sources, caches
runs/ # git-ignored: checkpoints, logs, metrics.jsonl, tb/
requirements.txt # torch 2.13 (MPS), transformers@main, openai, pyarrow, tensorboard…
.env # git-ignored, mode 600: OPENROUTER_API_KEY
```
---
## 12. How to run it
### Setup
```bash
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt # installs transformers from git main
echo "OPENROUTER_API_KEY=sk-or-..." > .env && chmod 600 .env
```
Smoke-test the environment first (gates everything): confirm `transformers` main
loads `google/gemma-4-E2B` on `mps` and runs a text-only forward pass.
### Build data
```bash
# 1) multi-view embeddings for a repo list (CPU to stay off the GPU during training)
python scripts/build_repo_multiview.py \
--repos-file data/multilang_repo_list.txt --max-repos 1000 --device cpu
# 2) repo-scoped QA (cheap model) β€” appends, resume-safe, cached
./venv/bin/python scripts/generate_repo_scoped_qa.py \
--model google/gemma-4-31b-it --workers 10
# 3) balance + assemble the aligned training set
python scripts/consolidate_qa.py --per-repo-cap 12
python scripts/assemble_6view_dataset.py
# -> data/embeddings/aligned6_embeddings.parquet + data/qna/aligned6_qna.jsonl
```
### Train (the exact `sixview_v2` command)
```bash
nohup ./venv/bin/python scripts/train_memory_lora.py --output-dir sixview_v2 \
--resume-from runs/sixview_v1/head.best.pt \
--embeddings-path data/embeddings/aligned6_embeddings.parquet \
--qna-path data/qna/aligned6_qna.jsonl --epochs 100 --max-hours 8 \
--checkpoint-every-steps 50 --checkpoint-every-minutes 30 --epoch-ckpt-every 5 \
--eval-every-steps 300 --limit-eval-docs 40 --max-seq-len 512 --fixed-seq-len \
--max-qna-per-doc 12 --lm-micro-batch 2 --device mps --no-gradient-checkpointing \
--rank 16 --head-hidden-dim 128 --head-dropout 0.1 --weight-decay 0.05 \
--early-stop-patience 25 --lr 8e-5 --lr-total-steps 9000 --min-available-gb 5 \
> runs/sixview_v2_train.log 2>&1 &
disown
```
**Flags you must not forget:** `--no-gradient-checkpointing` (the leak),
`--device cpu` for builds during training (contention), `--min-available-gb`
(system-wide memory floor).
### Watch it
```bash
tensorboard --logdir runs/sixview_v2/tb # train/loss, train/lr, eval/{suite}_loss
tail -f runs/sixview_v2_train.log
```
### Evaluate & inspect
```bash
python scripts/eval_memory_lora.py --ckpt runs/sixview_v2/head.best.pt # EM / EditSim
python scripts/show_eval_examples.py # base vs adapted
```
---
## 13. Evaluation methodology
- **Splits (deterministic, by `md5(repo) % 100`):** 80 % train / 10 % `cr_val` /
10 % `cr_test` **by repo**, so cross-repo suites are **entirely held-out
repositories** the hypernetwork never trained on. Within train repos, ~15 % of QA
is held out β†’ `ir_test` (in-repo generalization to unseen questions of seen repos).
- **Metrics:** causal-LM **eval loss** during training (fast, every N steps on
`--limit-eval-docs` docs to stay cheap on CPU), plus generation-time **Exact Match
(EM)** and **EditSim** for the recall eval.
- **The proof spot-check:** query the *adapted* model with repo-specific questions and
confirm the *base* (un-adapted) model gets them wrong/vague β€” proving the
**adapter**, not the base model's pretraining, does the work.
CPU eval of a float32 5B model is slow (~20 min for a full pass) β†’ we cap eval docs
(e.g. 10–40) for in-loop evals and run full EM eval separately.
---
## 14. Costs & budget discipline
- **Spend baseline:** $31.00 (`runs/spend_baseline.txt`); ~**$33.90 total** to date;
~**$11 remaining**. The project is run under explicit budget caps ("spend at most
$4 more") with spend-guards.
- **Unit economics:** `gemini-3.6-flash` β‰ˆ **$0.0021/QA**; `gemma-4-31b-it` β‰ˆ
**$0.001/repo (~$1 per 1000 repos)** β€” which is exactly why the 1000-repo expansion
uses the gemma model.
- **Free levers:** the per-prompt cache makes reruns free; embedding and training are
local (electricity only).
---
## 15. Current status & roadmap
**Live right now (three jobs in parallel, no contention):**
- **`sixview_v2` training** β€” resumed from `head.best.pt` on the doubled
**1,058-repo / 8,540-QA** dataset (858 steps/epoch), MPS. First held-out eval at
step 300 tells us whether doubling the data broke v1's 2.848 ceiling.
- **QA generation** β€” `gemma-4-31b-it` filling in all ~1,032 new repos for the
**complete** dataset (next training run).
- **Multiview build** β€” cloning/embedding toward the full +1,000-new target (CPU).
**Roadmap:**
1. Finish the complete 1000-new-repo dataset (embeddings + QA).
2. Assemble the full aligned set (~1,650 repos) and train **`sixview_v3`** on it.
3. Run generation-time **EM/EditSim** on the 6-view model (base vs adapted).
4. Push `head_hidden_dim` back up once on real GPUs; the 128 default was an
MPS-locality compromise.
5. Broaden Tier-A/B QA toward agent-harness use cases (Jira/ticket tracking,
diff/impact reasoning) already scaffolded in `generate_techlead_qa.py`.
**Open questions:**
- Does the 12288-d 6-view embedding actually beat the single 2048-d view on
generation EM, or only on loss? (loss says yes; EM eval pending)
- What's the real Tier-B ceiling β€” how much structural gist fits in rank-16?
- Optimal per-repo QA cap for the breadth/depth trade-off.
---
## 16. Glossary
- **Hypernetwork** β€” a network that outputs the weights of another network. Here:
repo embedding β†’ LoRA (A,B) matrices.
- **LoRA** β€” Low-Rank Adaptation: `y = Wx + scaling Β· B(Ax)`, with `A,B` low-rank
(rank 16). We *generate* A,B instead of training them per-repo.
- **6 views** β€” graph / arch / history / contracts / conventions / ops; each 2048-d,
concatenated to 12288-d.
- **CR / IR** β€” cross-repo (held-out repos) / in-repo (held-out QA of seen repos).
- **EM / EditSim** β€” Exact Match / edit-distance similarity of generated vs gold.
- **Tier A/B/C** β€” judgment (learnable) / structural gist (learnable) / exact recall
(needs RAG).
- **MPS** β€” Apple's Metal Performance Shaders GPU backend for PyTorch.
- **RepoPeftBench** β€” the paper's benchmark; 500 Python repos, repo-commit
embeddings + diffs, IR/CR splits.
---
*Maintained as living documentation. If you change a default, a path, or a flag,
update the matching section here β€” onboarding depends on it.*