docs: refresh the model card with the reseeded English numbers
Browse filesThe English row of the shared metrics table moved after the seed sweep:
accuracy 0.763 -> 0.859, ECE 0.061 -> 0.023, and the noise sentence
0.763 -> 0.760 became 0.859 -> 0.854. The multilingual numbers are
unchanged; no weights are touched in this commit.
README.md
CHANGED
|
@@ -71,12 +71,12 @@ gradient checkpointing.
|
|
| 71 |
|
| 72 |
| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
|
| 73 |
| --- | --- | --- | --- |
|
| 74 |
-
| English (ModernBERT-large + LoRA r=16 + choice head) | 0.
|
| 75 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | 0.853 | 0.038 | 13.3 |
|
| 76 |
|
| 77 |
-
Per primitive (English): `choice` 0.
|
| 78 |
0.684, `noul` 0.960, `score` 0.916. The localized, per-record-RNG data (B-1) lifted multilingual
|
| 79 |
-
`choice` from 0.40 to 0.68 and English overall from 0.72 to 0.
|
| 80 |
rank from 16 to 64** (alpha 128) removed the cross-language capacity bottleneck: overall accuracy
|
| 81 |
0.702 β 0.853 and `es` ECE 0.170 β 0.038 (`es` accuracy 0.472 β 0.956). Two of six languages now
|
| 82 |
meet ECE β€ 0.05 (`es` 0.038 and `pt` 0.024); `de` (0.063), `fr` (0.051), `it` (0.059) and `nl`
|
|
@@ -84,7 +84,7 @@ meet ECE β€ 0.05 (`es` 0.038 and `pt` 0.024); `de` (0.063), `fr` (0.051), `it`
|
|
| 84 |
The CUDA-graph fast path (`JEBA_FAST=1`) gives a 2.7Γ p50 speedup with 0 top-label flips.
|
| 85 |
|
| 86 |
**Robustness (B-4).** On a noisy view (one surface edit β typo/accents/casing β applied to 15% of
|
| 87 |
-
states) English drops only 0.
|
| 88 |
adapters are already robust to this noise model.
|
| 89 |
|
| 90 |
Full tables and environment are in
|
|
|
|
| 71 |
|
| 72 |
| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
|
| 73 |
| --- | --- | --- | --- |
|
| 74 |
+
| English (ModernBERT-large + LoRA r=16 + choice head) | 0.859 | 0.023 | 23.3 |
|
| 75 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | 0.853 | 0.038 | 13.3 |
|
| 76 |
|
| 77 |
+
Per primitive (English): `choice` 0.948, `noul` 0.718, `score` 0.910; (multilingual): `choice`
|
| 78 |
0.684, `noul` 0.960, `score` 0.916. The localized, per-record-RNG data (B-1) lifted multilingual
|
| 79 |
+
`choice` from 0.40 to 0.68 and English overall from 0.72 to 0.86. **Raising the multilingual LoRA
|
| 80 |
rank from 16 to 64** (alpha 128) removed the cross-language capacity bottleneck: overall accuracy
|
| 81 |
0.702 β 0.853 and `es` ECE 0.170 β 0.038 (`es` accuracy 0.472 β 0.956). Two of six languages now
|
| 82 |
meet ECE β€ 0.05 (`es` 0.038 and `pt` 0.024); `de` (0.063), `fr` (0.051), `it` (0.059) and `nl`
|
|
|
|
| 84 |
The CUDA-graph fast path (`JEBA_FAST=1`) gives a 2.7Γ p50 speedup with 0 top-label flips.
|
| 85 |
|
| 86 |
**Robustness (B-4).** On a noisy view (one surface edit β typo/accents/casing β applied to 15% of
|
| 87 |
+
states) English drops only 0.859 β 0.854 and multilingual (r=64) 0.853 β 0.847, so the released
|
| 88 |
adapters are already robust to this noise model.
|
| 89 |
|
| 90 |
Full tables and environment are in
|