Singaraj commited on
Commit
75ad30f
·
verified ·
1 Parent(s): bd84088

Sync model card: reproducibility scoping, seed-protocol precision

Browse files
Files changed (1) hide show
  1. README.md +11 -5
README.md CHANGED
@@ -90,11 +90,14 @@ Both models sit at the ceiling of this benchmark — FLORES+ sentences are long
90
  1,012-way retrieval saturates. Read this as evidence of zero out-of-domain degradation, not as a
91
  margin over LaBSE.
92
 
93
- Training was repeated with three random seeds; Creole→English test ndcg@10 across seeds:
94
- **0.9653 ± 0.0002** (accuracy@1 **0.9433 ± 0.0006**). The released checkpoint is seed 42, designated
95
- before results were seen.
96
 
97
- Every number is reproducible from the [training repository](https://github.com/LK-maker-007/morisien-embed).
 
 
 
98
 
99
  ## Training
100
 
@@ -135,7 +138,10 @@ Every number is reproducible from the [training repository](https://github.com/L
135
  (cosine ≈ 0.81 to the same sentence); caps-heavy text retrieves worse.
136
  - **Protocol note.** During recipe development the held-out test score was printed at the end of each
137
  training run, so recipe selection had test visibility; an independent audit bounded the resulting
138
- optimism at ≤ ~0.01 ndcg. The 3-seed replication was run after the recipe was frozen.
 
 
 
139
  - **Orthographic variation.** Training data mixes pre- and post-2011 (Lortograf Kreol Morisien)
140
  spellings; performance on older orthography is untested.
141
 
 
90
  1,012-way retrieval saturates. Read this as evidence of zero out-of-domain degradation, not as a
91
  margin over LaBSE.
92
 
93
+ The contrastive stage was repeated with three random seeds over the same deterministically mined
94
+ negative set; Creole→English test ndcg@10 across seeds: **0.9653 ± 0.0002** (accuracy@1
95
+ **0.9433 ± 0.0006**). The released checkpoint is seed 42, designated before results were seen.
96
 
97
+ Every number in the tables above is reproducible from the
98
+ [training repository](https://github.com/LK-maker-007/morisien-embed) (Matryoshka figures via
99
+ `scripts/evaluate.py --truncate-dim`). The Haitian-proximity and case-sensitivity figures under
100
+ Limitations come from an adversarial audit of the released checkpoint.
101
 
102
  ## Training
103
 
 
138
  (cosine ≈ 0.81 to the same sentence); caps-heavy text retrieves worse.
139
  - **Protocol note.** During recipe development the held-out test score was printed at the end of each
140
  training run, so recipe selection had test visibility; an independent audit bounded the resulting
141
+ optimism at ≤ ~0.01 ndcg. The 3-seed replication was run after the recipe was frozen. Leak filtering
142
+ reserves the Creole side of every evaluation pair; English/French target texts are not reserved, and
143
+ an audit found 1 of 999 benchmark passages also occurring in training as the translation of a
144
+ different Creole sentence.
145
  - **Orthographic variation.** Training data mixes pre- and post-2011 (Lortograf Kreol Morisien)
146
  spellings; performance on older orthography is untested.
147