File size: 4,430 Bytes
0dacd10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f1f9596
599df58
 
 
0dacd10
 
 
 
 
 
 
c550fe4
 
f1f9596
0dacd10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
426039c
 
f1f9596
 
 
0dacd10
 
 
4c84db2
599df58
 
4c84db2
599df58
4c84db2
599df58
f1f9596
 
 
426039c
599df58
 
4c84db2
599df58
 
ef1ed78
 
0dacd10
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
---
license: apache-2.0
library_name: jeba
language:
  - en
  - pt
  - es
  - fr
  - de
  - multilingual
tags:
  - decision-engine
  - system-one
  - calibration
  - multilingual
  - local-first
pipeline_tag: text-classification
---

# jeba-multi (System One decision engine)

> **Status: released (`v0.3.0`).** Trained on a single RTX 3060 12GB and published as LoRA
> adapters ([`munod/jeba-en`](https://huggingface.co/munod/jeba-en),
> [`munod/jeba-multi`](https://huggingface.co/munod/jeba-multi)); measured numbers below come
> from `benchmarks/report.md`.

## Model details

- **Developed by:** The jeba Authors.
- **Model type:** non-autoregressive encoder with three task distributions (`noul`, `choice`,
  `score`), answering typed questions about a state in one forward pass.
- **Trunk:** ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
- **Adapters:** [`munod/jeba-en`](https://huggingface.co/munod/jeba-en),
  [`munod/jeba-multi`](https://huggingface.co/munod/jeba-multi) (LoRA; load base + adapter).
- **License:** Apache-2.0.
- **Repository:** <https://github.com/munod/jeba>

## Uses

jeba answers atomic `choice` / `score` / `noul` questions about a state and returns typed values
with probabilities and `confidence`. It speaks the TypeSafe Jev `/v1/systemone` wire protocol as
a drop-in and runs **locally/offline** with no API key. Compose several atomic answers in code
rather than asking one broad question.

**Out of scope:** free-form text generation, multi-step reasoning, and any decision requiring
extended deliberation — decompose those into atomic questions and combine results in code.

## Bias, risks, and limitations

- Probabilities are only meaningful after **calibration**; the shipped temperature must be
  applied (see `docs/training.md`).
- Synthetic training data can inherit generator biases; public probes are evaluation-only.
- Confidence is a property of the distribution, not a guarantee of correctness.

## Training

Deterministic synthetic JSONL (`training/generate_data.py`) supervised with an RLCD
proper-scoring objective (`training/finetune_rlcd.py`), then temperature-calibrated on a held-out
split (`training/fit_calibration.py`). Configs and seed live under `training/configs/`.

## Evaluation

Reported by `training/evaluate.py` and rendered by `benchmarks/report.py` (accuracy, ECE, p50/p95
latency per primitive and language).

**Full-scale run (single RTX 3060 12GB):** 9,000 English / 18,000 multilingual train / 1,500 eval
deterministic synthetic records (fully localized per language, a learnable `other` team with rich
descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual)
plus a dedicated low-rank `choice` head (r=32, near-identity init), 4 epochs, batch 16, bf16 +
gradient checkpointing.

| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
| --- | --- | --- | --- |
| English (ModernBERT-large + LoRA r=16 + choice head) | 0.859 | 0.023 | 23.3 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | 0.853 | 0.038 | 13.3 |

Per primitive (English): `choice` 0.948, `noul` 0.718, `score` 0.910; (multilingual): `choice`
0.684, `noul` 0.960, `score` 0.916. The localized, per-record-RNG data (B-1) lifted multilingual
`choice` from 0.40 to 0.68 and English overall from 0.72 to 0.86. **Raising the multilingual LoRA
rank from 16 to 64** (alpha 128) removed the cross-language capacity bottleneck: overall accuracy
0.702 → 0.853 and `es` ECE 0.170 → 0.038 (`es` accuracy 0.472 → 0.956). Two of six languages now
meet ECE ≤ 0.05 (`es` 0.038 and `pt` 0.024); `de` (0.063), `fr` (0.051), `it` (0.059) and `nl`
(ECE 0.104, accuracy 0.663) remain above target (NFR-C06 partially open).
The CUDA-graph fast path (`JEBA_FAST=1`) gives a 2.7× p50 speedup with 0 top-label flips.

**Robustness (B-4).** On a noisy view (one surface edit — typo/accents/casing — applied to 15% of
states) English drops only 0.859 → 0.854 and multilingual (r=64) 0.853 → 0.847, so the released
adapters are already robust to this noise model.

Full tables and environment are in
[`benchmarks/report.md`](https://github.com/munod/jeba/blob/main/benchmarks/report.md).

## Citation

```bibtex
@misc{jeba2026,
  title        = {jeba: a local-first System One decision engine},
  author       = {The jeba Authors},
  year         = {2026},
  howpublished = {\url{https://github.com/munod/jeba}}
}
```