# Mindweave experiment — 2026-07-23 This report records the first external-data experiment for CFO-HRM. The result is a shadow-mode learning artifact, not evidence that the model is ready for a real accounting close. ## Data and split The experiment uses revision-pinned, CC BY-NC 4.0 Mindweave bank and accounting ledger datasets. Exact revisions, file hashes, and attribution are recorded in `THIRD_PARTY_DATA.md` and each run's `dataset_manifest.json`. The adapter verified: - 3,208 unique provenance-derived bank-to-cash-GL pairs; - 24 source-less monthly bank fees, conservatively labeled unmatched; - 3,232 bank rows in 145 independently shuffled fixed-size cases; and - no unresolved rows in the pinned snapshot. The chronological split prevents period leakage: | Split | Periods | Cases | Bank rows | Matches | Fees | | --- | --- | ---: | ---: | ---: | ---: | | Train | 2024-01 through 2025-06 | 92 | 2,021 | 2,003 | 18 | | Validation | 2025-07 through 2025-09 | 23 | 509 | 506 | 3 | | Test | 2025-10 through 2025-12 | 30 | 702 | 699 | 3 | Independent deterministic shuffling reduced same-position bank/GL targets from 98.6% to 4.1%, removing a positional shortcut. ## Training The robust run starts from the synthetic-pretrained checkpoint at step 900 and uses: - 150 outer steps and four HRM deep-supervision segments; - 75% Mindweave and 25% synthetic retention batches; - 0.25 reference, 0.10 description, and 0.05 source-module dropout on external training rows only; and - external validation plus a fixed, disjoint synthetic retention evaluation. Checkpoint selection prioritizes external match F1, external exception macro recall, zero false auto-matches, synthetic retention metrics, and then external calibration. It selected step 50. ## Held-out results All selective metrics use a 0.90 confidence threshold. ### Raw test split | System | Assignment accuracy | Match F1 | Exception accuracy | Bank-fee recall | ECE | Auto-match coverage | False auto-match rate | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | Synthetic-pretrained HRM | 99.43% | 99.43% | 7.41% | 100% | 14.76% | 52.99% | 0% | | Deterministic evidence baseline | 100% | 100% | 99.57% | 0% | 0.83% | 99.57% | 0% | | Robust fine-tuned HRM, step 50 | 100% | 100% | 100% | 100% | 7.95% | 83.90% | 0% | The baseline's perfect matching result is important: this synthetic dataset is largely solvable with exact amount, date, and provenance evidence. The HRM does not establish an advantage on matching. Its useful measured improvement here is exception classification, particularly recognizing bank fees. ### Stress tests | Test | Assignment accuracy | Match F1 | Bank-fee recall | Selective coverage | False auto-match rate | | --- | ---: | ---: | ---: | ---: | ---: | | Reference masked | 95.44% | 95.42% | 33.33% | 80.06% | 0% | | Source module, reference, and description masked | 96.30% | 96.28% | 0% | 48.29% | 0% | The model abstained safely at the chosen threshold in these tests, but the accuracy and fee-recall degradation show that it still depends heavily on clean synthetic identifiers and contextual fields. On the run's fixed synthetic retention split, the selected checkpoint achieved 95.71% match F1 and 34.22% exception macro recall. The corresponding starting checkpoint achieved 95.41% and 37.86%, respectively: matching was retained, while exception behavior still shifted across domains. ## End-to-end controlled execution Held-out test case 8 contains 19 bank rows, including one monthly fee. The promoted checkpoint proposed 18 matches and one unmatched bank fee. Go controls reported: ```text review_queue: exception_review controls_passed: 18 blocked_candidates: 0 unmatched_bank_rows: 1 requires_human_review: true posting_authorized: false audit verification: 1 hash-chained event verified ``` The review request, decision, and audit files were all written with mode `0600`. No component contains an ERP posting or approval client. ## Decision Promote `runs/mindweave-robust-20260723/best.pt` only as the current local shadow-mode checkpoint. Keep the deterministic baseline beside it and report both. The next meaningful evaluation needs harder grouped transactions, reference corruption, split/aggregate payments, GL-only items, and ultimately approved human-reconciled real exports.