| --- |
| license: cc-by-nc-4.0 |
| library_name: cfo-hrm |
| datasets: |
| - mindweave/bank-transactions-us |
| - mindweave/accounting-ledger-us |
| tags: |
| - pytorch |
| - finance |
| - accounting |
| - bank-reconciliation |
| - hierarchical-reasoning |
| - structured-data |
| - human-in-the-loop |
| --- |
| |
| # CFO-HRM Mindweave reconciliation checkpoint |
|
|
| This is the promoted **shadow-mode learning checkpoint** for CFO-HRM, a |
| native-PyTorch hierarchical recurrent model that scores candidate links between |
| bank transactions and general-ledger cash lines and classifies bank-row |
| exceptions. |
|
|
| It is not a general-purpose language model. It accepts the fixed structured |
| tensors produced by the CFO-HRM data pipeline. It does not read arbitrary |
| natural-language prompts, approve accounting decisions, or post journal |
| entries. |
|
|
| ## Intended use |
|
|
| This checkpoint is intended for: |
|
|
| - local research and education about hierarchical reasoning on structured |
| accounting tasks; |
| - reproducing the reported synthetic bank-to-GL reconciliation experiment; |
| - shadow-mode candidate scoring, followed by deterministic accounting controls |
| and human review; and |
| - studying confidence thresholds, abstention, exception routing, and robustness |
| to missing evidence. |
|
|
| It is not intended for autonomous reconciliation, journal-entry creation, |
| approval, ERP posting, financial reporting, audit evidence, commercial use, or |
| decisions involving real funds. Do not use a prediction as an accounting |
| conclusion. |
|
|
| ## Model details |
|
|
| - Framework: native PyTorch |
| - Trainable parameters: 1,665,352 |
| - Input capacity: up to 24 bank rows and 24 GL rows per case |
| - Hidden size: 128 |
| - Attention heads: 4 |
| - High-level cycles/layers: 2 / 1 |
| - Low-level cycles/layers: 2 / 1 |
| - Deep-supervision segments at evaluation: 4 |
| - Exception classes: 6 |
| - Continuous features per row: 3 |
| - Categorical fields per row: 6 |
| - Selected fine-tuning step: 50 |
| - Checkpoint file: `best.pt` |
| - Checkpoint SHA-256: |
| `65049fa087ff6e4fa2fc28262c237ab7b4987b25cc292104bdf7b78b047662e9` |
|
|
| The model embeds categorical and continuous bank/GL features into a shared |
| sequence. Recurrent high-level and low-level reasoning modules repeatedly |
| exchange state. Pair heads score candidate bank-to-GL links, an unmatched head |
| scores the unmatched option, and an exception head classifies each bank row. |
| The model itself does not enforce accounting policy. CFO-HRM applies a |
| one-to-one assignment decoder and then performs entity, currency, amount, date, |
| confidence, uniqueness, and source-lineage checks outside the network. |
|
|
| This implementation is inspired by the Hierarchical Reasoning Model research |
| architecture, but is a small purpose-built reconciliation model rather than an |
| official implementation or checkpoint from that work. |
|
|
| ## Training data |
|
|
| Fine-tuning used two synthetic Mindweave datasets: |
|
|
| | Dataset | Pinned revision | License | |
| | --- | --- | --- | |
| | [`mindweave/bank-transactions-us`](https://huggingface.co/datasets/mindweave/bank-transactions-us) | `290e657395a1ce25837642b8309ca9a741f85303` | CC BY-NC 4.0 | |
| | [`mindweave/accounting-ledger-us`](https://huggingface.co/datasets/mindweave/accounting-ledger-us) | `b2f2d0890353438b4473f2c3b2d2b5e15a5132a0` | CC BY-NC 4.0 | |
|
|
| The checkpoint was initialized from a locally synthetic-pretrained CFO-HRM |
| checkpoint at step 900. Robust fine-tuning ran for 150 outer steps with four |
| deep-supervision segments, a batch size of 16, a learning rate of `5e-5`, and |
| 25% synthetic retention batches. External training rows used independent |
| categorical dropout of 0.25 for reference, 0.10 for description, and 0.05 for |
| source module. Validation selected step 50. |
|
|
| Labels for positive pairs were derived only from unique |
| `(company_id, source_module, source_id)` journal-entry provenance and the bank |
| account's mapped cash GL line. Twenty-four narrowly identified source-less |
| monthly fees were labeled unmatched. Ambiguous provenance is retained for |
| review and excluded from match supervision rather than silently treated as |
| unmatched. |
|
|
| Bank and GL rows were independently and deterministically shuffled within each |
| case. This reduced same-position targets from 98.6% to 4.1% and removed the |
| near-diagonal positional shortcut present in the source ordering. |
|
|
| ### Chronological split |
|
|
| | Split | Period | Cases | Bank rows | Matched rows | Bank fees | |
| | --- | --- | ---: | ---: | ---: | ---: | |
| | Train | 2024-01 through 2025-06 | 92 | 2,021 | 2,003 | 18 | |
| | Validation | 2025-07 through 2025-09 | 23 | 509 | 506 | 3 | |
| | Test | 2025-10 through 2025-12 | 30 | 702 | 699 | 3 | |
|
|
| The raw third-party datasets are **not included in this model repository**. |
| They remain attribution-required, non-commercial CC BY-NC 4.0 data. Download |
| them from the upstream dataset repositories at the pinned revisions and review |
| their terms before use. Do not infer that this checkpoint or its training data |
| is cleared for commercial use. |
|
|
| ## Evaluation |
|
|
| The following values were measured on the held-out October–December 2025 test |
| split. Selective metrics use a confidence threshold of 0.90. |
|
|
| ### Raw test split |
|
|
| | System | Assignment accuracy | Match F1 | Exception accuracy | Bank-fee recall | ECE | Auto-match coverage | False auto-match rate | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| | Synthetic-pretrained HRM | 99.43% | 99.43% | 7.41% | 100% | 14.76% | 52.99% | 0% | |
| | Deterministic evidence baseline | 100% | 100% | 99.57% | 0% | 0.83% | 99.57% | 0% | |
| | This checkpoint | 100% | 100% | 100% | 100% | 7.95% | 83.90% | 0% | |
|
|
| Additional measured values for this checkpoint on the raw split: |
|
|
| - match precision and recall: 100% / 100%; |
| - exact-case accuracy: 100%; |
| - dollar-weighted accuracy: 100%; |
| - selective assignment accuracy: 100%; |
| - selective coverage: 84.33%; |
| - selective dollar coverage: 83.40%; |
| - abstention rate: 15.67%; and |
| - assignment Brier score: 3.83%. |
|
|
| The deterministic baseline's perfect matching result is a central caveat. This |
| synthetic test set is largely solvable through exact amount, date, and |
| provenance evidence. These results do not establish a matching advantage for |
| HRM. The measured improvement over that baseline is exception classification, |
| especially recognizing the three held-out bank fees. |
|
|
| ### Missing-evidence stress tests |
|
|
| | Test | Assignment accuracy | Match F1 | Exception accuracy | Bank-fee recall | ECE | Selective coverage | Auto-match coverage | False auto-match rate | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | |
| | Reference masked | 95.4416% | 95.4220% | 99.5726% | 33.3333% | 5.2496% | 80.0570% | 79.6296% | 0% | |
| | Source module, reference, and description masked | 96.2963% | 96.2804% | 87.7493% | 0% | 11.2771% | 48.2906% | 47.8632% | 0% | |
|
|
| For comparison, the deterministic baseline with references masked achieved |
| 99.1453% assignment accuracy and 99.1416% match F1, with 0% bank-fee recall and |
| 62.9630% selective coverage. With source module, reference, and description all |
| masked, it retained 99.1453% assignment accuracy and 99.1416% match F1 but |
| abstained on 99.5726% of rows: selective coverage was only 0.4274%, auto-match |
| coverage was 0%, and bank-fee recall was 0%. |
|
|
| On the fixed synthetic retention split, this checkpoint achieved 95.71% match |
| F1 and 34.22% exception macro recall. The starting checkpoint achieved 95.41% |
| and 37.86%, respectively. |
|
|
| ## Local use |
|
|
| The checkpoint requires the CFO-HRM source package and the two pinned upstream |
| datasets. It is not compatible with the Hugging Face `pipeline()` API. |
|
|
| ```bash |
| git clone https://github.com/Aznatkoiny/cfo-hrm.git |
| cd cfo-hrm |
| uv sync --extra dev |
| |
| # Download this model repository separately, then verify the checkpoint. |
| sha256sum /path/to/model-repository/best.pt |
| ``` |
|
|
| The expected digest is: |
|
|
| ```text |
| 65049fa087ff6e4fa2fc28262c237ab7b4987b25cc292104bdf7b78b047662e9 |
| ``` |
|
|
| Place the upstream CSV files under the paths shown below, preserving the |
| dataset layouts, and evaluate: |
|
|
| ```bash |
| uv run cfo-hrm-eval-mindweave \ |
| /path/to/model-repository/best.pt \ |
| --bank-dir data/raw/mindweave/bank-transactions-us \ |
| --ledger-dir data/raw/mindweave/accounting-ledger-us \ |
| --split test \ |
| --device cuda \ |
| --confidence-threshold 0.9 \ |
| --output evaluation.json |
| ``` |
|
|
| For CPU-only evaluation, use `--device cpu --no-amp`. |
|
|
| To produce a versioned review request for the separate control layer: |
|
|
| ```bash |
| uv run cfo-hrm-export-review \ |
| /path/to/model-repository/best.pt \ |
| --bank-dir data/raw/mindweave/bank-transactions-us \ |
| --ledger-dir data/raw/mindweave/accounting-ledger-us \ |
| --split test \ |
| --case-index 8 \ |
| --device cuda \ |
| --output review-request.json |
| ``` |
|
|
| The Go orchestrator must then validate that proposal. In the supplied |
| end-to-end example, it routed 18 control-passing matches and one unmatched fee |
| to `exception_review`, set `requires_human_review: true`, and set |
| `posting_authorized: false`. |
|
|
| ## Safety boundary |
|
|
| Use this model only as one component in a controlled review workflow: |
|
|
| 1. Treat model outputs as proposals, never approvals. |
| 2. Apply deterministic accounting controls independently of the model. |
| 3. Abstain and route low-confidence, unmatched, or policy-sensitive rows to a |
| human reviewer. |
| 4. Preserve source hashes, checkpoint hashes, record IDs, control results, and |
| review decisions in a tamper-evident audit trail. |
| 5. Keep ERP posting and approval credentials outside the model process. |
|
|
| The accompanying CFO-HRM Go control layer deliberately sets |
| `posting_authorized: false`. The model repository contains no ERP posting |
| client and must not be extended to bypass segregation of duties or human |
| approval. |
|
|
| ## Limitations |
|
|
| - All external training and test records are synthetic. |
| - The test split contains only three bank fees and no unresolved rows, so the |
| reported exception percentages have high variance. |
| - Exact identifiers and clean provenance make the raw matching task unusually |
| easy. |
| - The model degrades materially when contextual evidence is removed, especially |
| on fee recall. |
| - It was not evaluated on split payments, aggregate payments, many-to-many |
| reconciliation, substantial reference corruption, GL-only exceptions, real |
| ERP exports, other countries, other accounting standards, or adversarial |
| inputs. |
| - Maximum case size is fixed at 24 bank rows and 24 GL rows. |
| - Confidence scores are not guarantees; calibration can shift on any new |
| company, system, period, or data-generating process. |
| - This checkpoint cannot replace professional judgment, internal controls, |
| accounting policy, audit procedures, or regulatory compliance. |
|
|
| Before any real-data pilot, perform company-specific data mapping, privacy and |
| security review, out-of-period evaluation, calibration, control testing, |
| reviewer acceptance testing, and comparison with deterministic baselines. Keep |
| the system in shadow mode until those gates are independently satisfied. |
|
|
| ## License and attribution |
|
|
| The CFO-HRM source code is Apache-2.0. This checkpoint was fine-tuned on |
| CC BY-NC 4.0 third-party datasets and is therefore published under |
| CC BY-NC 4.0 as the conservative model-artifact license. This does not broaden |
| the upstream permissions or constitute legal advice. Attribute the two |
| Mindweave dataset repositories, retain their license notices, and do not use the |
| checkpoint or data commercially without establishing the necessary rights. |
|
|
|
|