tachyone-multi (System One decision engine)

Status: released (v0.9.0), revision 2026-10-08 β€” both checkpoints are the B-15 template-train retrain (English carries its P3 stack rebuilt on the new trunk); revision commits are listed in docs/huggingface.md. Trained on NVIDIA L4 23GB and published as LoRA adapters (munod/tachyone-en, munod/tachyone-multi); measured numbers below come from benchmarks/report.md.

This revision is B-17: the teacher expanded the tone banks, the multilingual arm was retrained, and B-1's per-language ECE acceptance closes. 352 phrases authored by a local qwen3.6:35b teacher went into every neutral/request bank (five domains Γ— six languages, inserted before the held-out phrase β€” the holdout evals are byte-identical, sha-pinned); the multilingual checkpoint retrained with the unchanged B-15 recipe and English was not touched (byte-identical numbers). On phrasing never seen in any training: multilingual 0.9264 (routed), English 0.9825 β€” the in-template β†’ holdout gap collapsed βˆ’0.1105 β†’ βˆ’0.0165 β€” and all 12 (primitive, language) cells now sit at ECE ≀ 0.05 (worst score/fr 0.0257; score/it 0.0521 β†’ 0.0014, accuracy 0.8699 β†’ 0.9928) β€” B-1 criterion 3 (NFR-C06) CLOSED on the frozen yardstick. Also improved: choice holdout 0.9248 β†’ 0.9808, the noul tripwire 0.7807 β†’ 0.9735, every language β‰₯ 0.9928 (es 1.0000), in-template multilingual ECE 0.0228 β†’ 0.0036 (zero gate exceptions everywhere). Disclosures, never smoothed: the expansion targeted a diagnosed mechanism (BACKLOG B-17 carries the per-domain table); off-domain MASSIVE 0.074 β†’ 0.058 with ECE raw 0.046 β†’ 0.197 β€” in-domain gains re-sharpened off-domain confidence (the L-007/L-015 trade), still 3.4Γ— its chance; and the English P3 choice ECE 0.5243 (pooled fit at T = 10.0) stands exactly as previously disclosed.

The v0.9.0/v0.10.0 revisions were B-15/B-16 (history). B-15 retrained both checkpoints on template_split: "train" after the data recipe was corrected end to end (sampler strides decoupled so every language trains every label, evals made leave-one-template-out, phrase content Γ—3, volume 1.7Γ—) and every training set regenerated so training and the holdout eval share 0% of templates. The multilingual lineage carries the two-epoch interleaved touch-up that fixed the B-14 noul cell inversion, and the trainer now ships a per-cell validation monitor. On phrasing never seen in any training (eval_*_domains_holdout): multilingual 0.8355, English 0.9825 β€” +0.127 / +0.131 over the previously released adapters, with the in-template β†’ holdout gap at βˆ’0.1105 / βˆ’0.0179. Disclosures, never smoothed: English choice confidence is refitted on the rebuilt never-trained holdout and its pooled fit pegged the grid at T = 10.0, so in-domain choice ECE reads 0.5243 honestly instead of in-sample (P3's in-domain ≀ 0.05 acceptance recorded NOT met for the rebuilt asset; accuracy is untouched); and B-1's per-language ECE ≀ 0.05 on holdout, which after the B-16 confidence map (2026-10-09) reads 11 of 12 cells ≀ 0.05 with score/it 0.0521 recorded NOT met (B-15 measured 7 failing cells, worst 0.1495; full numbers in .specs/project/BACKLOG.md B-1 and B-16). The multilingual choice/score answers now report a fitted confidence β€” a 2D peakedness Γ— strength table per (primitive, language) keyed by a detector that reads the state and the localized question (B-16); probabilities and decisions are untouched, and the in-template multilingual ECE trade (0.0069 β†’ 0.0228, still zero gate exceptions) is disclosed with it.

The v0.8.0 English artifact was B-13: the mixture retrain (history). Same recipe and seed as the B-5 run, one factor changed β€” the training data: 35,540 records (the 21,000 five-domain records plus 11,000 pinned public records β€” MultiNLI, BoolQ, Banking77, licences in training/data/jev_sources.lock.json β€” and 3,540 executable-rule-tree family records), then the identical frozen-trunk choice-bank fit. It sweeps both English splits 0.964 β†’ 1.000 (gate strict 1.000; unseen-text rows 0.9987) β€” and, measured against a fresh same-recipe control that moves +0.6, it lifts Intelligence on the 231 public JevBench items (evaluation-only, never trained on) from 8.4 to 15.0. The probes below are the honest external check: XNLI 0.341 β†’ 0.566, typed-decisions 0.269 β†’ 0.367.

The v0.8.0 multilingual artifact was B-5b: five-domain coverage (history). Its trunk is joint-trained on 30,000 multilingual five-domain records (support keeps its 18,000; four new domains 3,000 each) and its choice heads were re-fitted on the frozen trunk by training/fit_choice_bank.py β€” the same construction as the English artifact, recipe shipped next to the weights as choice_bank_fit.json. On the five-domain split it scores 0.9975 (previous adapter zero-shot: 0.561), gate strict 1.000, all six languages at ECE ≀ 0.004.

Labels (B-11 + B-12, ADR-0014/ADR-0015): every label in all three primitives is derived from the text it accompanies β€” noul from its phrase bank (request β†’ 1, neutral/empty β†’ 0), score from the tone's level (empty β†’ middle), choice from the option the state names (empty β†’ the catch-all other). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left every noul label in es/de/nl at 0, and gave 7.8% of score rows a "near-tie" the text never showed.

Provenance, stated plainly: each adapter carries its recipe next to the weights (choice_bank_fit.json) β€” the English adapter is the B-15 five-domain bank over the template-train trunk (en_tt, 6 epochs) with the P3 stack rebuilt on it; the multilingual adapter is the B-15 joint run (52,200 records, 8 epochs, then the 2-epoch interleaved touch-up β†’ multi_tt_il) plus the frozen-trunk head fit. Both heads-first constructions exist because training the corrected labels makes noul+score trivial and costs choice (an identical-recipe control landed 13 points lower, L-006) β€” the fix that worked was re-fitting the head on a trunk that already knew the data.

Model details

  • Developed by: The Tachyone Authors.
  • Model type: non-autoregressive encoder with three task distributions (noul, choice, score), answering typed questions about a state in one forward pass.
  • Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
  • Adapters: munod/tachyone-en, munod/tachyone-multi (LoRA; load base + adapter).
  • License: Apache-2.0.
  • Repository: https://github.com/munod/tachyone

Uses

Tachyone answers atomic choice / score / noul questions about a state and returns typed values with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as a drop-in and runs locally/offline with no API key. Compose several atomic answers in code rather than asking one broad question.

Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β€” decompose those into atomic questions and combine results in code.

Bias, risks, and limitations

  • Probabilities are only meaningful after calibration; the shipped temperature must be applied (see docs/training.md).
  • Synthetic training data can inherit generator biases; public probes are evaluation-only.
  • Confidence is a property of the distribution, not a guarantee of correctness.

Training

Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out split (training/fit_calibration.py). Configs and seed live under training/configs/.

Evaluation

Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95 latency per primitive and language).

Full-scale run (NVIDIA L4 23GB): 36,000 five-domain English / 52,200 five-domain multilingual records generated with template_split: "train" (the holdout evals share 0% of training templates) / 1,500–7,500 eval deterministic synthetic records (fully localized per language β€” all five domains ship the seven training languages, a learnable other option with rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual) plus a dedicated low-rank choice head (near-identity init); 6 epochs English / 8 epochs multilingual, batch 8 Γ— grad-accum 4, bf16 + gradient checkpointing, and a per-cell validation monitor (cell_monitor.jsonl) so one (primitive, domain, language) collapse cannot hide behind the aggregate. Both published artifacts are bank-over-frozen-trunk: the trunk is trained jointly, then training/fit_choice_bank.py re-fits the choice heads on cached encodings (rank 128, 8 epochs at lr 1e-4), shipping its recipe as choice_bank_fit.json; the multilingual trunk additionally carries the two-epoch interleaved touch-up (training/interleave_continue.py).

Checkpoint Split Accuracy ECE (calibrated) p50 (ms)
English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) support 1.0000 0.188 57.7
English, five-domain split 5 domains 0.9983 0.191 57.7
English, holdout phrasing (never in training) 5 domains 0.9825 0.188 58.0
Multilingual (mmBERT-base + LoRA r=64 + fitted bank) β€” routed runtime support 0.9300 0.038 52.8
Multilingual, five-domain split 5 domains 0.9221 0.026 53.5
Multilingual, holdout phrasing (never in training) 5 domains 0.9264 0.032 53.3

The English ECE column reads the pooled never-trained fit (below): choice is flattened to T=10.0 to calibrate held-out rows, which shows up honestly in-domain (0.5243 choice ECE β€” accuracy untouched). The multilingual rows are the routed system: ~13–15% of rows fall through to the English checkpoint by language routing (empty state or no Latin-language signal β€” .specs/project/BACKLOG.md L-013), which is why the five-domain multilingual row sits below the multilingual checkpoint's own 0.9885 when it answers everything.

Per primitive (English, five-domain): choice 0.9956, noul 0.9996, score 0.9996; (multilingual, five-domain): choice 0.9564, noul 0.9136, score 0.8768; (holdout phrasing): English choice 0.9600 / noul 0.9956 / score 0.9920, multilingual choice 0.8960 / noul 0.7836 / score 0.8268. Label audit: every noul row is judged against its own text β€” 0 contradictory in every eval set (positive rates 0.472–0.486), per language in benchmarks/report.md.

Confidence comes from evidence (P3 β€” stack rebuilt on the v0.9.0 trunk). The English adapter ships two extra assets: state_prototypes.json (K=32 spherical k-means centroids of its own 36,000 training states) and confidence_calibration.json (that bank plus a fitted noul map, with sha256 provenance of both inputs). A noul answer keeps its direction β€” the cosine still decides yes/no β€” and takes its magnitude from strength = max_k cos(state, centroid), clamped to [0.5, 1]. The temperatures come from the rebuilt never-trained holdout (a stride slice of the new training file plus never-trained public rows, re-asserted against the public items before a byte was written): choice T = 10.0 β€” the pooled fit pegged the top of the pre-registered DEFAULT_GRID, a direct consequence of training on template_split: "train" while the holdout keeps its hard rows β€” noul T = 0.75 with the map, score pinned at 0.1 because its answer is the expected value (the pin's own record: rounded-EV accuracy 0.53 at T=1 vs 0.9996 shipped). Accuracy is untouched by construction (argmax and direction never move). Recorded NOT met: P3's in-domain ECE ≀ 0.05 acceptance for the rebuilt asset β€” choice reads 0.5243 in-domain under T=10.0 (the honest price of calibrating held-out rows; the fix is a fit-basis redesign, not another fit β€” see .specs/features/jevbench/spec.md P3 results for why no legal fit set can see bench difficulty). The multilingual adapter has no evidence stack; its served temperature_calibration.json is fitted on 15,000 pooled in-template + holdout predictions (choice 6.0, noul 0.25, score 0.25 β€” fitted where the model is not saturated, per L-015). The public probes were re-measured on this v0.9.0 artifact (2026-10-08): XNLI 0.566 β†’ 0.333 (exactly chance β€” the B-13 mixture's MultiNLI layer was not carried into the recomposed recipe, follow-up recorded in .specs/project/BACKLOG.md B-15; published, never re-fixed), typed-decisions 0.367 β†’ 0.306, MASSIVE 0.051 β†’ 0.074 (chance 0.017, the best reading yet), with off-domain ECE raw down across the board (0.096 / 0.042 / 0.068). JevBench Calibration 0.0 β†’ 52.6 with its recorded gate of 60 not met remains a v0.8.0-artifact measurement β€” JevBench was set aside by decision after its maintainer changed the submission methodology.

What the English rows mean (in-sample, stated plainly). The support and five-domain splits read 1.0000 / 0.9983 because their phrasing is seen β€” training excluded the holdout template pool, but these splits still draw the seen templates β€” so saturation here is a ceiling, not a claim. The honest number is the holdout row: 0.9825 overall (choice 0.9600), phrasing no training ever emitted. The external checks quoted by earlier revisions β€” JevBench Intelligence 8.4 β†’ 15.0 against a fresh control's +0.6 β€” were measured on the v0.8.0 mixture artifact (JevBench set aside by decision after its maintainer changed the submission methodology). The probes were re-measured on this artifact (2026-10-08): XNLI 0.333 (chance 0.333), typed-decisions 0.306, MASSIVE 0.074 (chance 0.017). Full tables: docs/benchmarks.md.

The v0.8.0 multilingual gates (B-5b β€” history; v0.9.0's gates are B-15's). The previous adapter measured 0.561 zero-shot on the new five-domain split (worst new domain agent_tools 0.438, support 0.8413) β€” the gates were support β‰₯ 0.8413, worst new domain β‰₯ 0.70, per-domain ECE ≀ 0.05 (exceptions declared), gate strict published. That artifact posted support 1.0000, worst new domain 0.993, per-domain ECE 0.0005–0.0038 (zero exceptions) and gate strict 1.000 over 2,500 choice rows. v0.9.0 (B-15) posts on the unchanged current benchmark: multilingual support 0.9760, worst new domain 0.9747, per-domain ECE 0.0019–0.0202 (zero exceptions), strict 1.000, every language β‰₯ baseline; overall 0.9885 / ECE 0.0069 when the multilingual checkpoint answers everything β€” the routed rows above show the pair as served. On the support-only routed split (13–15% of rows fall through to the English checkpoint by language routing, .specs/project/BACKLOG.md L-013) the per-language story is unchanged in kind: three of six languages sit above the 0.05 ECE target on that split, and on the holdout the B-1 per-language ECE criterion was NOT met at B-15 (7 cells, worst score/it 0.1495) β€” the B-16 cycle (serve-key detector + fitted confidence, 2026-10-09) moved it to 11/12 with score/it 0.0521 recorded NOT met β€” and the B-17 cycle (teacher bank expansion + retrain, 2026-10-10) closed it: all 12/12 cells ≀ 0.05 on the same frozen holdout (worst score/fr 0.0257, score/it 0.0014 at accuracy 0.9928), choice 0.9248 β†’ 0.9808 and every other axis improved or byte-identical. NFR-C06 met; the trail is in BACKLOG B-1/B-15/B-16/B-17. The CUDA-graph fast path (TACHYONE_FAST=1) gives 2.53Γ— p50 on English (17.31 β†’ 6.83 ms) and 3.71Γ— on the multilingual five-domain path (14.23 β†’ 3.84 ms) β€” re-measured on the v0.9.0 artifacts (2026-10-08): 0 top-label flips multilingual, 1 of 16 sampled on English with max answer-probability difference 0.024 (disclosed as measured; NFR-P01 met by both paths). The gate and the keyed heads execute after the encode, which the graphed path never sees.

Robustness (B-4). On a noisy view (one surface edit β€” typo/accents/casing β€” applied to 15% of states) English moves 1.0000 β†’ 0.9953 and the multilingual routed pair 0.9107 β†’ 0.9087, so the released adapters are robust to this noise model.

Full tables and environment are in benchmarks/report.md.

Citation

@misc{tachyone2026,
  title        = {tachyone: a local-first System One decision engine},
  author       = {The tachyone Authors},
  year         = {2026},
  howpublished = {\url{https://github.com/munod/tachyone}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support