File size: 4,647 Bytes
31dc8dc | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 | # SpecForge Roadmap
The consolidated, phase-by-phase roadmap behind [`../../plan.md`](../../plan.md) (the reconciled
architecture). It **folds in** the former online-disaggregation roadmap PR (#618) so there is one
roadmap home. Each track doc gives, per phase: **Goal / Target state / Implementation
(files + symbols) / Tests / Done-when**.
## Standing decisions (apply across all tracks)
- **Substrate is canonical.** `SampleRef` (metadata, control plane; `assert_no_tensors`) +
`FeatureStore` (tensors: Local/SharedDir/Mooncake) + `FeatureDataLoader` β `TrainBatch`. There is
**no separate `HiddenStateStream`** source of truth β the loader *is* the stream; online/offline/
disaggregated vary only in (ref source + `FeatureStore`), shielded from training.
- **Frozen target, no weight sync.** "Train-with-decode" = a **frozen** target streams hidden
states over a **fixed** prompt set; the draft is not in the generation loop. Weight-sync /
hot draft-update / weight-version registry / staleness gate / on-policy are **out of scope**;
`draft_weight_version` is kept **only as provenance**.
- **Ray is OPEN.** A *candidate* for the O2 scale-out orchestration layer β likely needed for
multi-node N-producer/M-trainer scale-out β but **not committed and not a non-goal**. See the
decision gate in [online-disaggregation.md](./online-disaggregation.md) Β§O2.
- **Preserve the training seam** (`TrainerCore` / `DraftTrainStrategy` / `TrainingBackend` +
`StepContext`); a domain `Trainer` + managers *wrap* it, they do not replace it. It is relocated
**intact** (not rewritten) from `runtime/training` to top-level `training/` in the move-only
step `E0` β see [domain-refactor.md](./domain-refactor.md).
- **One implementation home per concern; `runtime/` is substrate-only.** `runtime/` holds only the
DataFlow spine (`control_plane` + `data_plane` + `contracts`). All training-execution code lives
in top-level `training/`, all rollout/capture-execution code in top-level `inference/`, and
`modeling/` holds model definitions only (no orchestration, no capture factory). **New code is
born in its final home** β the Phase-D managers land directly in `training/`, never deeper in
`runtime/`; the existing seam and target engine are relocated once, in `E0`. There is **no facade
package**.
## Tracks
| Track | Doc | Scope |
|---|---|---|
| Domain / architecture | [domain-refactor.md](./domain-refactor.md) | Strategy/registry, `TargetEngine`, domain `Trainer` + managers, drafts registry, config/CLI/export |
| Online disaggregation | [online-disaggregation.md](./online-disaggregation.md) | Live frozen-target generation, cross-process control plane, scale-out (Ray = open), hardening |
| Eval & breadth | [eval-and-breadth.md](./eval-and-breadth.md) | Acceptance-length eval harness; new algorithms = a `StrategySpec` + loss |
## Phase status at a glance
| Phase | Track | Size | Status |
|---|---|---|---|
| A β Composable launch | domain | L | **in review** (#627/#628/#629) |
| O1.1 β Shared cross-process control plane | online | M | in review (~#624) |
| O1.2 β Async streaming loop + one-process builder | online | M | in review (~#625) |
| B β Domain abstractions (`TargetEngine` + `Trainer`) | domain | L | next |
| O1.3 β Live SGLang-server hidden-state capture | online | L | next (π΄ spike gates it) |
| E1 β Acceptance-length eval harness | eval | M | next |
| C β Colocated lightweight path | domain | M | later |
| D β Training managers (no_sync / resume / ckpt / eval) | domain | L | later |
| E0 β Layout consolidation (move-only: seam + target engine β top-level homes) | domain | M | later (front of E) |
| E β Composition & run surface (drafts registry, config/CLI, export) | domain | L | later |
| O2 β Scale-out orchestration (Ray = open) | online | L | later |
| O3 β Hardening (RDMA pool, restart, observability) | online | L | later |
| E2 β Algorithm breadth (MTP/β¦) | eval | L | later |
## Dependencies (cross-track)
```
domain: A(rev.) ββ¬ββΆ B ββ¬ββΆ C
β βββΆ D ββΆ E0 ββΆ E
βββΆ B.TargetEngine ββββββββββββββ
online: O1.1(review) ββΆ O1.2(review) ββΆ O1.3 βββββ ββΆ O2(Ray=open) ββΆ O3
eval: E1 ββΆ E2 (parallel / orthogonal to both tracks)
```
- Domain **B** (`TargetEngine`) unblocks online **O1.3** (a real `SGLangServerEngine` capture
backend), so it is the highest-leverage next domain step.
- **E1**'s `Evaluator` is the same one domain **D** wires into the trainer loop β build once.
|