tachyone-multi (System One decision engine)
Status: released (
v0.9.0), revision 2026-10-08 β both checkpoints are the B-15 template-train retrain (English carries its P3 stack rebuilt on the new trunk); revision commits are listed indocs/huggingface.md. Trained on NVIDIA L4 23GB and published as LoRA adapters (munod/tachyone-en,munod/tachyone-multi); measured numbers below come frombenchmarks/report.md.This revision is B-17: the teacher expanded the tone banks, the multilingual arm was retrained, and B-1's per-language ECE acceptance closes. 352 phrases authored by a local
qwen3.6:35bteacher went into everyneutral/requestbank (five domains Γ six languages, inserted before the held-out phrase β the holdout evals are byte-identical, sha-pinned); the multilingual checkpoint retrained with the unchanged B-15 recipe and English was not touched (byte-identical numbers). On phrasing never seen in any training: multilingual 0.9264 (routed), English 0.9825 β the in-template β holdout gap collapsed β0.1105 β β0.0165 β and all 12(primitive, language)cells now sit at ECE β€ 0.05 (worstscore/fr0.0257;score/it0.0521 β 0.0014, accuracy 0.8699 β 0.9928) β B-1 criterion 3 (NFR-C06) CLOSED on the frozen yardstick. Also improved:choiceholdout 0.9248 β 0.9808, thenoultripwire 0.7807 β 0.9735, every language β₯ 0.9928 (es1.0000), in-template multilingual ECE 0.0228 β 0.0036 (zero gate exceptions everywhere). Disclosures, never smoothed: the expansion targeted a diagnosed mechanism (BACKLOG B-17 carries the per-domain table); off-domain MASSIVE 0.074 β 0.058 withECE raw0.046 β 0.197 β in-domain gains re-sharpened off-domain confidence (the L-007/L-015 trade), still 3.4Γ its chance; and the English P3choiceECE 0.5243 (pooled fit at T = 10.0) stands exactly as previously disclosed.The v0.9.0/v0.10.0 revisions were B-15/B-16 (history). B-15 retrained both checkpoints on
template_split: "train"after the data recipe was corrected end to end (sampler strides decoupled so every language trains every label, evals made leave-one-template-out, phrase content Γ3, volume 1.7Γ) and every training set regenerated so training and the holdout eval share 0% of templates. The multilingual lineage carries the two-epoch interleaved touch-up that fixed the B-14noulcell inversion, and the trainer now ships a per-cell validation monitor. On phrasing never seen in any training (eval_*_domains_holdout): multilingual 0.8355, English 0.9825 β +0.127 / +0.131 over the previously released adapters, with the in-template β holdout gap at β0.1105 / β0.0179. Disclosures, never smoothed: Englishchoiceconfidence is refitted on the rebuilt never-trained holdout and its pooled fit pegged the grid at T = 10.0, so in-domainchoiceECE reads 0.5243 honestly instead of in-sample (P3's in-domain β€ 0.05 acceptance recorded NOT met for the rebuilt asset; accuracy is untouched); and B-1's per-language ECE β€ 0.05 on holdout, which after the B-16 confidence map (2026-10-09) reads 11 of 12 cells β€ 0.05 withscore/it0.0521 recorded NOT met (B-15 measured 7 failing cells, worst 0.1495; full numbers in.specs/project/BACKLOG.mdB-1 and B-16). The multilingualchoice/scoreanswers now report a fitted confidence β a 2D peakedness Γ strength table per(primitive, language)keyed by a detector that reads the state and the localized question (B-16); probabilities and decisions are untouched, and the in-template multilingual ECE trade (0.0069 β 0.0228, still zero gate exceptions) is disclosed with it.The v0.8.0 English artifact was B-13: the mixture retrain (history). Same recipe and seed as the B-5 run, one factor changed β the training data: 35,540 records (the 21,000 five-domain records plus 11,000 pinned public records β MultiNLI, BoolQ, Banking77, licences in
training/data/jev_sources.lock.jsonβ and 3,540 executable-rule-tree family records), then the identical frozen-trunkchoice-bank fit. It sweeps both English splits 0.964 β 1.000 (gate strict 1.000; unseen-text rows 0.9987) β and, measured against a fresh same-recipe control that moves +0.6, it lifts Intelligence on the 231 public JevBench items (evaluation-only, never trained on) from 8.4 to 15.0. The probes below are the honest external check: XNLI 0.341 β 0.566, typed-decisions 0.269 β 0.367.The v0.8.0 multilingual artifact was B-5b: five-domain coverage (history). Its trunk is joint-trained on 30,000 multilingual five-domain records (support keeps its 18,000; four new domains 3,000 each) and its
choiceheads were re-fitted on the frozen trunk bytraining/fit_choice_bank.pyβ the same construction as the English artifact, recipe shipped next to the weights aschoice_bank_fit.json. On the five-domain split it scores 0.9975 (previous adapter zero-shot: 0.561), gate strict 1.000, all six languages at ECE β€ 0.004.Labels (B-11 + B-12, ADR-0014/ADR-0015): every label in all three primitives is derived from the text it accompanies β
noulfrom its phrase bank (requestβ 1,neutral/empty β 0),scorefrom the tone's level (empty β middle),choicefrom the option the state names (empty β the catch-allother). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left everynoullabel ines/de/nlat 0, and gave 7.8% ofscorerows a "near-tie" the text never showed.Provenance, stated plainly: each adapter carries its recipe next to the weights (
choice_bank_fit.json) β the English adapter is the B-15 five-domain bank over the template-train trunk (en_tt, 6 epochs) with the P3 stack rebuilt on it; the multilingual adapter is the B-15 joint run (52,200 records, 8 epochs, then the 2-epoch interleaved touch-up βmulti_tt_il) plus the frozen-trunk head fit. Both heads-first constructions exist because training the corrected labels makesnoul+scoretrivial and costschoice(an identical-recipe control landed 13 points lower, L-006) β the fix that worked was re-fitting the head on a trunk that already knew the data.
Model details
- Developed by: The Tachyone Authors.
- Model type: non-autoregressive encoder with three task distributions (
noul,choice,score), answering typed questions about a state in one forward pass. - Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
- Adapters:
munod/tachyone-en,munod/tachyone-multi(LoRA; load base + adapter). - License: Apache-2.0.
- Repository: https://github.com/munod/tachyone
Uses
Tachyone answers atomic choice / score / noul questions about a state and returns typed values
with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as
a drop-in and runs locally/offline with no API key. Compose several atomic answers in code
rather than asking one broad question.
Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β decompose those into atomic questions and combine results in code.
Bias, risks, and limitations
- Probabilities are only meaningful after calibration; the shipped temperature must be
applied (see
docs/training.md). - Synthetic training data can inherit generator biases; public probes are evaluation-only.
- Confidence is a property of the distribution, not a guarantee of correctness.
Training
Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD
proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out
split (training/fit_calibration.py). Configs and seed live under training/configs/.
Evaluation
Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95
latency per primitive and language).
Full-scale run (NVIDIA L4 23GB): 36,000 five-domain English / 52,200 five-domain
multilingual records generated with template_split: "train" (the holdout evals share 0% of
training templates) / 1,500β7,500 eval deterministic synthetic records (fully localized per
language β all five domains ship the seven training languages, a learnable other option with
rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64
multilingual) plus a dedicated low-rank choice head (near-identity init); 6 epochs English /
8 epochs multilingual, batch 8 Γ grad-accum 4, bf16 + gradient checkpointing, and a
per-cell validation monitor (cell_monitor.jsonl) so one (primitive, domain, language)
collapse cannot hide behind the aggregate. Both published artifacts are
bank-over-frozen-trunk: the trunk is trained jointly, then training/fit_choice_bank.py
re-fits the choice heads on cached encodings (rank 128, 8 epochs at lr 1e-4), shipping its
recipe as choice_bank_fit.json; the multilingual trunk additionally carries the two-epoch
interleaved touch-up (training/interleave_continue.py).
| Checkpoint | Split | Accuracy | ECE (calibrated) | p50 (ms) |
|---|---|---|---|---|
| English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) | support | 1.0000 | 0.188 | 57.7 |
| English, five-domain split | 5 domains | 0.9983 | 0.191 | 57.7 |
| English, holdout phrasing (never in training) | 5 domains | 0.9825 | 0.188 | 58.0 |
| Multilingual (mmBERT-base + LoRA r=64 + fitted bank) β routed runtime | support | 0.9300 | 0.038 | 52.8 |
| Multilingual, five-domain split | 5 domains | 0.9221 | 0.026 | 53.5 |
| Multilingual, holdout phrasing (never in training) | 5 domains | 0.9264 | 0.032 | 53.3 |
The English ECE column reads the pooled never-trained fit (below): choice is flattened to
T=10.0 to calibrate held-out rows, which shows up honestly in-domain (0.5243 choice ECE β
accuracy untouched). The multilingual rows are the routed system: ~13β15% of rows fall
through to the English checkpoint by language routing (empty state or no Latin-language signal
β .specs/project/BACKLOG.md L-013), which is why the five-domain multilingual row sits
below the multilingual checkpoint's own 0.9885 when it answers everything.
Per primitive (English, five-domain): choice 0.9956, noul 0.9996, score
0.9996; (multilingual, five-domain): choice 0.9564, noul 0.9136, score
0.8768; (holdout phrasing): English choice 0.9600 / noul 0.9956 / score
0.9920, multilingual choice 0.8960 / noul 0.7836 / score 0.8268.
Label audit: every noul row is judged against its own text β 0 contradictory in every
eval set (positive rates 0.472β0.486), per language in benchmarks/report.md.
Confidence comes from evidence (P3 β stack rebuilt on the v0.9.0 trunk). The English
adapter ships two extra assets: state_prototypes.json (K=32 spherical k-means centroids of
its own 36,000 training states) and confidence_calibration.json (that bank plus a fitted
noul map, with sha256 provenance of both inputs). A noul answer keeps its direction β the
cosine still decides yes/no β and takes its magnitude from
strength = max_k cos(state, centroid), clamped to [0.5, 1]. The temperatures come from the
rebuilt never-trained holdout (a stride slice of the new training file plus never-trained
public rows, re-asserted against the public items before a byte was written): choice
T = 10.0 β the pooled fit pegged the top of the pre-registered DEFAULT_GRID, a direct
consequence of training on template_split: "train" while the holdout keeps its hard rows β
noul T = 0.75 with the map, score pinned at 0.1 because its answer is the
expected value (the pin's own record: rounded-EV accuracy 0.53 at T=1 vs 0.9996 shipped).
Accuracy is untouched by construction (argmax and direction never move). Recorded NOT met:
P3's in-domain ECE β€ 0.05 acceptance for the rebuilt asset β choice reads 0.5243
in-domain under T=10.0 (the honest price of calibrating held-out rows; the fix is a fit-basis
redesign, not another fit β see .specs/features/jevbench/spec.md P3 results for why no
legal fit set can see bench difficulty). The multilingual adapter has no evidence stack; its
served temperature_calibration.json is fitted on 15,000 pooled in-template + holdout
predictions (choice 6.0, noul 0.25, score 0.25 β fitted where the model is not
saturated, per L-015). The public probes were re-measured on this v0.9.0 artifact
(2026-10-08): XNLI 0.566 β 0.333 (exactly chance β the B-13 mixture's MultiNLI layer was
not carried into the recomposed recipe, follow-up recorded in .specs/project/BACKLOG.md
B-15; published, never re-fixed), typed-decisions 0.367 β 0.306, MASSIVE 0.051 β
0.074 (chance 0.017, the best reading yet), with off-domain ECE raw down across the board
(0.096 / 0.042 / 0.068). JevBench Calibration 0.0 β 52.6 with its recorded gate of 60
not met remains a v0.8.0-artifact measurement β JevBench was set aside by decision
after its maintainer changed the submission methodology.
What the English rows mean (in-sample, stated plainly). The support and five-domain splits
read 1.0000 / 0.9983 because their phrasing is seen β training excluded the holdout template
pool, but these splits still draw the seen templates β so saturation here is a ceiling, not a
claim. The honest number is the holdout row: 0.9825 overall (choice 0.9600), phrasing no
training ever emitted. The external checks quoted by earlier revisions β JevBench Intelligence
8.4 β 15.0 against a fresh control's +0.6 β were measured on the v0.8.0 mixture
artifact (JevBench set aside by decision after its maintainer changed the submission
methodology). The probes were re-measured on this artifact (2026-10-08): XNLI 0.333
(chance 0.333), typed-decisions 0.306, MASSIVE 0.074 (chance 0.017). Full tables:
docs/benchmarks.md.
The v0.8.0 multilingual gates (B-5b β history; v0.9.0's gates are B-15's). The previous adapter measured
0.561 zero-shot on the new five-domain split (worst new domain agent_tools 0.438, support
0.8413) β the gates were support β₯ 0.8413, worst new domain β₯ 0.70, per-domain ECE β€ 0.05
(exceptions declared), gate strict published. That artifact posted support 1.0000,
worst new domain 0.993, per-domain ECE 0.0005β0.0038 (zero exceptions) and gate
strict 1.000 over 2,500 choice rows. v0.9.0 (B-15) posts on the unchanged current
benchmark: multilingual support 0.9760, worst new domain 0.9747, per-domain ECE 0.0019β0.0202
(zero exceptions), strict 1.000, every language β₯ baseline; overall 0.9885 / ECE 0.0069 when
the multilingual checkpoint answers everything β the routed rows above show the pair as
served. On the support-only routed split (13β15% of
rows fall through to the English checkpoint by language routing, .specs/project/BACKLOG.md
L-013) the per-language story is unchanged in kind: three of six languages sit above the
0.05 ECE target on that split, and on the holdout the B-1 per-language ECE criterion was
NOT met at B-15 (7 cells, worst score/it 0.1495) β the B-16 cycle (serve-key
detector + fitted confidence, 2026-10-09) moved it to 11/12 with score/it 0.0521
recorded NOT met β and the B-17 cycle (teacher bank expansion + retrain, 2026-10-10)
closed it: all 12/12 cells β€ 0.05 on the same frozen holdout (worst score/fr
0.0257, score/it 0.0014 at accuracy 0.9928), choice 0.9248 β 0.9808 and every other
axis improved or byte-identical. NFR-C06 met; the trail is in BACKLOG
B-1/B-15/B-16/B-17.
The CUDA-graph fast path
(TACHYONE_FAST=1) gives 2.53Γ p50 on English (17.31 β 6.83 ms) and 3.71Γ on the
multilingual five-domain path (14.23 β 3.84 ms) β re-measured on the v0.9.0 artifacts
(2026-10-08): 0 top-label flips multilingual, 1 of 16 sampled on English with max
answer-probability difference 0.024 (disclosed as measured; NFR-P01 met by both paths). The gate and
the keyed heads execute after the encode, which the graphed path never sees.
Robustness (B-4). On a noisy view (one surface edit β typo/accents/casing β applied to 15% of states) English moves 1.0000 β 0.9953 and the multilingual routed pair 0.9107 β 0.9087, so the released adapters are robust to this noise model.
Full tables and environment are in
benchmarks/report.md.
Citation
@misc{tachyone2026,
title = {tachyone: a local-first System One decision engine},
author = {The tachyone Authors},
year = {2026},
howpublished = {\url{https://github.com/munod/tachyone}}
}