# ARCHON SFT v2 Phase 6.5 — Readiness Validation Report Date: 2026-06-02 Validator: Claude Sonnet 4.6 (subagent) Scope: 7 Phase 6.5 artifacts, cross-validation of 5 criteria --- ## VERDICT: NOT READY — 3 blocking conflicts, 1 data source defect Synthetic generation must NOT launch until conflicts C1 and C2 are resolved. Conflict C3 and defect D1 can be fixed in parallel with generation prep. --- ## VALIDATION RESULTS ### V1 — HF Search Results Status: PASS (with noted risk) Single usable peer-equal dataset identified: `Multi-Agent-LLMs/DEBATE` (apache-2.0). 7,200 usable rows across debate + relay + memory + report paradigm configs. Peer-equal check confirmed: no supervisor/manager/orchestrator pattern in sampled rows. Rejection of camel-ai datasets on NC licence + structural hierarchy grounds is correct. UltraInteract_sft classification as N1-only is correct (SFT split is flat, no multi-agent turns). Risk noted: DEBATE tasks are trivia/QA domains (manta ray biology, seating puzzles), not ARCHON-domain (network/cybersec/Etherlink). Format transfer is accepted but domain gap is real. Mitigation already documented in adversarial check section — adequate. ### V2 — Synthetic Generation Plan Feasibility Status: PARTIAL PASS — cost reasonable, hours reasonable, sample targets INCONSISTENT (see C1/C2) Cost: $220-250 USD (generation plan estimate with prompt caching is the reliable figure; the YAML $125 figure underestimates by using a lower $/call approximation). Hours: 28-35h sequential, ~18-20h with 2 parallel workers. Feasible. Pilot gates (V1-V4) are well-designed and proportionate. Parallel launch strategy (N1+N4 terminal 1, N2 terminal 2, N3 terminal 3) is sound. ### V3 — MTP Task Profiles Consistency Status: PASS 4 new task types added correctly: family_consult [1.0, 0.65, 0.50, 0.40, 0.30] (N2 structured — moderately predictable) family_negotiate [1.0, 0.55, 0.40, 0.30, 0.20] (N3 debate — mid-predictability) family_swarm [1.0, 0.40, 0.25, 0.18, 0.12] (N4 emergent — low predictability) family_role_rotate [1.0, 0.70, 0.55, 0.45, 0.35] (meta escalation — structured) Predictability gradient N2>N3>N4 is correct and rationale is documented. `general_conv_fr` was NOT duplicated (already present). Comment confirms this. Phase 6 submix sum assertion: 0.01+0.04+0.06+0.04 = 0.15 — PASS. Curriculum continuity: 7 phases, steps 0-50000, no gaps, no overlaps — PASS. DPO sub-window 42000-50000 aligned in both Python file and design doc — PASS. Legacy `multi_agent_coord`/`delegation_chain` kept at low weight (0.8) for rehearsal — correct. ### V4 — Mix Distribution Target Coherence Status: FAIL — sample target inconsistency across three documents (see C1, C2) ### V5 — METATRON / "Master" Connotation Audit Status: CONDITIONAL PASS — one spec doc has a METATRON leak (see C3) "master" usage in all Phase 6.5 artifacts is definitional negation only (describing what to avoid). No content encodes master-subordinate behavior. File name `_out_synthetic_prompts_master.md` is a title, not a connotation issue. Generator script: METATRON correctly replaced by `[trading-specialist]` in ROSTER, SYSTEM_BASE, and anti-pattern list. Forbidden-pattern validator checks "metatron" string. Generation plan gate V2 checks "metatron" absence. PASS. `_out_phase_6_5_design.md` line 75: schema example lists `METATRON` as a participant: "participants": ["ARCHON", "<1-3 peers from NEXUS/CYPHER/AETHER/METATRON>"] This is a design doc inconsistency — the actual generator script never uses METATRON. Risk: low (design doc is not read by the generator), but must be corrected before sharing. ### V6 — Coverage: HF + Synthetic vs Total Target Status: FAIL — N4 coverage is severely under-specified (see C2) --- ## CONFLICTS (blocking unless noted) ### C1 — BLOCKING: N4 sample count three-way inconsistency | Document | N4 target | |---|---| | _out_multi_agent_sources.yaml | 10,000 (to_generate) | | _out_synthetic_generation_plan.md | 3,000 | | _out_mix_distribution_target_phase65.json | 20,000 (target_samples in N4_family_swarm) | The mix JSON is the authoritative training target: 20,000 N4 samples. The YAML says 10,000 to generate (which could mean HF covers 10K — but HF has 0 N4 rows). The generation plan says only 3,000 — this is the most severe undercount. Resolution required: set N4 target uniformly to 10,000 across all three documents (20,000 mix target is too aggressive for entirely-synthetic N4; 10K is the YAML intent and matches the HF search gap analysis). OR keep mix at 20K and update synth plan to 20K (cost ~$200 for N4 alone via Opus templates + Sonnet expansion, per YAML estimate). Decision needed from Jescy before launching. ### C2 — BLOCKING: Aggregate N2-N4 coverage gap Computed from authoritative sources: | Niveau | HF available | Synth plan generates | Total | Mix JSON target | |---|---|---|---|---| | N2 | 2,400 | 5,000 | 7,400 | 20,000 | | N3 | 3,600 | 7,000 | 10,600 | 30,000 | | N4 | 0 | 3,000 | 3,000 | 20,000 | | **Total N2-N4** | **6,000** | **15,000** | **21,000** | **70,000** | 21,000 actual vs 70,000 mix target = 30% coverage. This is a 3.3× shortfall. Possible resolutions: (a) Reduce mix JSON targets to match synth capacity (N2: 7.5K, N3: 10.5K, N4: 3K or 10K). Pro: preserves current synth budget. Con: N4 remains thin. (b) Scale up synth generation to match mix targets (cost ~$800-1000 total, 3-4 days). Pro: full coverage. Con: significant cost and time increase. (c) Accept mismatch with oversampling: synth samples x10 weight to fill quota. Risk: memorisation at high oversampling factor. ### C3 — NON-BLOCKING: METATRON in design doc schema `_out_phase_6_5_design.md` line 75 lists METATRON as a possible participant in the generator schema example. The generator script never actually uses METATRON (correctly uses `[trading-specialist]`). Fix: replace METATRON with `[trading-specialist]` in the schema example. One-line edit, no functional impact. --- ## DEFECTS (non-blocking but must fix before mix build) ### D1 — UltraInteract_sft N2 filter is broken Mix JSON entry for N2_family_consult uses UltraInteract_sft with filter `trajectory_length >= 4`. The HF search report (section 2) explicitly confirms: "The parent_id tree is the preference learning split (_pair), NOT this SFT split." The SFT split is flat (instruction, response) pairs — no `trajectory_length` field exists. This filter will either fail at runtime or return 0 rows. Impact: 8,000 planned N2 samples from this source will be unavailable. Mitigation: remove UltraInteract N2-filtered source from N2 mix entry. Replace with additional synth generation (+3,000 N2) or oversample DEBATE relay/report configs. Net effect on C2 gap: worsens N2 shortfall by 8K additional samples. --- ## SUMMARY TABLE | Check | Status | Action | |---|---|---| | HF peer-equal datasets identified | PASS | None | | Synthetic cost feasible | PASS | Use $220-250 budget | | Synthetic hours feasible | PASS | 28-35h sequential | | MTP profiles consistent | PASS | None | | Phase 6 submix sum = 15% | PASS | None | | Curriculum continuity 0-50K | PASS | None | | DPO sub-window alignment | PASS | None | | N4 sample count aligned | FAIL — C1 | Resolve before launch | | N2-N4 coverage vs mix targets | FAIL — C2 | Resolve before launch | | METATRON absent in generated content | PASS | None | | METATRON in design doc schema | WARN — C3 | One-line fix | | UltraInteract N2 filter valid | FAIL — D1 | Remove/replace before mix build | | DPO egalitarian dependency order | PASS | None | --- ## TOTAL COST ESTIMATE (Phase 6.5 complete) Scenario A — current synth plan (17K samples): Synthetic generation: $220-250 GPU cost Phase 6 (8K steps, RTX PRO 6000 ~$1.2/hr, ~8-10h): $10-12 Total Phase 6.5: ~$230-262 Scenario B — scaled to match mix targets (~54K N2-N4 samples): Synthetic generation: ~$800-1000 GPU Phase 6: ~$10-12 Total Phase 6.5: ~$810-1012 Recommended: resolve C1/C2 first, then choose scenario. --- ## PRE-LAUNCH CHECKLIST Before running `python _out_synthetic_generator.py`: - [ ] C1 resolved: N4 target aligned across all three documents - [ ] C2 resolved: synth targets or mix JSON targets reconciled - [ ] C3 fixed: METATRON replaced by [trading-specialist] in _out_phase_6_5_design.md line 75 - [ ] D1 acknowledged: UltraInteract N2 source removed from mix JSON or filter corrected - [ ] Dry-run passes: `python _out_synthetic_generator.py --niveau all --dry-run` - [ ] Pilot 50 samples N2 inspected manually before full run - [ ] ANTHROPIC_API_KEY set in environment - [ ] synthetic_out/ and logs/ directories created