Aleksander commited on
Commit
9bda7f0
·
verified ·
1 Parent(s): 113b1d1

Upload index.html

Browse files
Files changed (1) hide show
  1. index.html +407 -36
index.html CHANGED
@@ -124,88 +124,459 @@ footer{border-top:1px solid var(--line);padding:36px 0;color:var(--muted);font-s
124
  </div>
125
  </main>
126
 
 
127
  <article class="detail" id="model-polmath">
128
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">PolMATH</div></div></div>
129
- <div class="wrap detail-hero"><div class="eyebrow">Numerical representation research</div><h2>PolMATH</h2><p>A custom model exploring numerical value as a structured signal separate from the surface text used to write it.</p><div class="update-log"><div class="update-item"><div class="update-date">2025</div><div class="update-copy"><strong>Initial architecture and training experiments</strong><span>Structured numerical channel, numeric decoding and number↔text association tests.</span></div></div><div class="update-item"><div class="update-date">Current</div><div class="update-copy"><strong>Publication-oriented rework</strong><span>Original branch paused; cleaner protocol and English-language continuation are being prepared.</span></div></div></div></div>
 
 
 
 
 
 
 
 
 
130
  <section><div class="wrap">
131
- <div class="eyebrow">Core idea</div><h3 class="section-title">Numbers were not treated as ordinary text alone.</h3>
 
132
  <div class="prose">
133
- <p>The transformer stack, tokenizer behaviour, numerical codec, embedding fusion, auxiliary objectives, routing and numerical decoding path formed one purpose-built system.</p>
134
- <p>The numerical path could relate forms such as <strong>1939</strong>, <strong>1939.0</strong> and <strong>1.939e3</strong> through a shared structured representation while still learning semantic associations around the value.</p>
 
135
  </div>
136
- <div class="callout"><strong>Observed behaviour</strong><p>Number-word ↔ numeric-value matching, numeric-position selection and factual associations involving numerical values, including the bidirectional 1939 ↔ Second World War relation.</p></div>
137
  </div></section>
 
138
  <section><div class="wrap">
139
- <div class="eyebrow">Roadmap</div><h3 class="section-title">Paused, but not abandoned.</h3>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
140
  <div class="timeline">
141
- <div class="timeline-item"><div class="timeline-year">2025</div><div class="timeline-body"><h4>Initial PolMATH</h4><p>Custom numerical representation and multi-objective training experiments.</p></div></div>
142
- <div class="timeline-item"><div class="timeline-year">Now</div><div class="timeline-body"><h4>Paused</h4><p>Original training branch stopped while the work is reorganized.</p></div></div>
143
- <div class="timeline-item"><div class="timeline-year">Next</div><div class="timeline-body"><h4>Publication track</h4><p>Cleaner protocol and an English-language iteration intended for a more formal research release.</p></div></div>
 
144
  </div>
145
  </div></section>
146
  </article>
147
 
 
 
148
  <article class="detail" id="model-oris660">
149
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">ORIS 660M</div></div></div>
150
  <div class="wrap detail-hero">
151
- <div class="eyebrow">Structural compression and recovery</div><h2>ORIS 660M</h2>
152
- <p>A 660,131,328-parameter student derived from Bielik-1.5B-v3 by collapsing 32 transformer blocks into a selected 12-block path while preserving the teacher's width and core geometry.</p>
153
- <div class="stats"><div class="stat"><strong>660.13M</strong><span>parameters</span></div><div class="stat"><strong>32 12</strong><span>blocks</span></div><div class="stat"><strong>621.984M</strong><span>recovery labels</span></div><div class="stat"><strong>Paused</strong><span>research active</span></div></div><div class="update-log"><div class="update-item"><div class="update-date">07 Aug 2026</div><div class="update-copy"><strong>ORIS 660M branch begins</strong><span>Initial structural-reduction experiment from Bielik-1.5B-v3.</span></div></div><div class="update-item"><div class="update-date">Recovery tests</div><div class="update-copy"><strong>Selected-vs-uniform architecture comparison</strong><span>Confirmed that retained layer identity materially changes recovery trajectory.</span></div></div><div class="update-item"><div class="update-date">Later branch</div><div class="update-copy"><strong>Knowledge and data-control experiments</strong><span>Inherited-vs-random controls, diagnostic benchmarks and planned staged recovery curriculum.</span></div></div></div>
 
 
 
 
 
 
154
  </div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
155
  <section><div class="wrap">
156
- <div class="eyebrow">Versions / experiments</div><h3 class="section-title">The project is a sequence of recovery experiments.</h3>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
157
  <div class="version-grid">
158
- <div class="version-card"><small>Step 0</small><h4>Structural collapse</h4><p>12 selected layers, severe initial loss increase, much higher throughput and lower VRAM.</p></div>
159
- <div class="version-card"><small>Recovery</small><h4>Main causal LM continuation</h4><p>Direct next-token training without teacher-logit distillation.</p></div>
160
- <div class="version-card"><small>V2.5A</small><h4>Benchmark snapshot</h4><p>Returned to a non-trivial Polish benchmark regime while generation remained unstable.</p></div>
161
- <div class="version-card"><small>V3 direction</small><h4>Data-controlled recovery</h4><p>Planned 22B-token staged curriculum with knowledge and stability control.</p></div>
 
 
162
  </div>
163
  </div></section>
 
164
  <section><div class="wrap">
165
- <div class="eyebrow">Knowledge</div><h3 class="section-title">Inherited structure did not behave like a random network.</h3>
166
- <div class="stats"><div class="stat"><strong>24.7 35.0%</strong><span>fact accuracy</span></div><div class="stat"><strong>0.154 ~0.366</strong><span>teacher-rank Spearman</span></div><div class="stat"><strong>~10%</strong><span>random control accuracy</span></div><div class="stat"><strong>20 → 41.5%</strong><span>ENTITY_ONLY accuracy</span></div></div>
167
- <div class="callout"><strong>Open question</strong><p>ORIS distinguishes knowledge that remains accessible, survives in degraded form, becomes latent/inaccessible, or has to be genuinely relearned.</p></div>
 
 
 
168
  </div></section>
 
169
  <section><div class="wrap">
170
- <div class="eyebrow">Roadmap</div><h3 class="section-title">The next bottleneck is data, not just training time.</h3>
 
171
  <div class="timeline">
172
- <div class="timeline-item"><div class="timeline-year">S1 · 6B</div><div class="timeline-body"><h4>Stabilization</h4><p>Lower-risk coherent material for structural recovery.</p></div></div>
173
- <div class="timeline-item"><div class="timeline-year">S2 · 12B</div><div class="timeline-body"><h4>Knowledge recovery</h4><p>Broader domains and higher information density.</p></div></div>
174
- <div class="timeline-item"><div class="timeline-year">S3 · 4B</div><div class="timeline-body"><h4>Consolidation</h4><p>Harder distributions introduced after stabilization.</p></div></div>
175
  </div>
 
 
 
 
 
 
 
 
 
 
 
176
  </div></section>
177
  </article>
178
 
 
 
179
  <article class="detail" id="model-smallc">
180
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">ORIS Small C</div></div></div>
181
  <div class="wrap detail-hero">
182
- <div class="eyebrow">Compact Polish encoder</div><h2>ORIS Small C</h2>
183
- <p>Originally developed under the working name ORIS Bert Small C. It is a custom encoder using BERT-style MLM training, not a stock reduced BERT.</p>
184
- <div class="stats"><div class="stat"><strong>25.41M</strong><span>parameters</span></div><div class="stat"><strong>6</strong><span>layers</span></div><div class="stat"><strong>128K</strong><span>BPE vocab</span></div><div class="stat"><strong>8.00B</strong><span>pretraining tokens</span></div></div><div class="update-log"><div class="update-item"><div class="update-date">16 Aug 2026</div><div class="update-copy"><strong>Small C checkpoint completed</strong><span>8B-token from-scratch MLM pretraining completed.</span></div></div><div class="update-item"><div class="update-date">Downstream tests</div><div class="update-copy"><strong>Filtering and encoder-efficiency evaluation</strong><span>Pipeline, KLEJ-style and local throughput tests established the compact encoder's practical role.</span></div></div></div>
 
 
 
 
 
 
185
  </div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
186
  <section><div class="wrap">
187
- <div class="eyebrow">Architecture</div><h3 class="section-title">Factorized embeddings, RMSNorm and mixed local/global attention.</h3>
188
- <div class="prose"><p>Five layers use 256-token local attention; one layer performs full 1024-token global mixing. MLM logits are produced only for masked positions through the tied 128-dimensional token embedding space.</p></div>
 
 
 
 
 
 
 
 
 
189
  </div></section>
 
190
  <section><div class="wrap">
191
- <div class="eyebrow">Roadmap</div><h3 class="section-title">From pipeline tool to encoder family.</h3>
 
 
 
 
 
192
  <div class="timeline">
193
- <div class="timeline-item"><div class="timeline-year">Origin</div><div class="timeline-body"><h4>ORIS data bottleneck</h4><p>Built to accelerate local filtering, categorization and scoring.</p></div></div>
194
- <div class="timeline-item"><div class="timeline-year">Small C</div><div class="timeline-body"><h4>Concept-stage release</h4><p>From-scratch compact encoder and custom tokenizer.</p></div></div>
195
- <div class="timeline-item"><div class="timeline-year">Next</div><div class="timeline-body"><h4>Possible broader encoder line</h4><p>More architecture and downstream work before a fuller ORIS encoder release.</p></div></div>
 
196
  </div>
197
  </div></section>
198
  </article>
199
 
 
 
200
  <article class="detail" id="model-cain">
201
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">CAIN</div></div></div>
202
- <div class="wrap detail-hero"><div class="eyebrow">Persona research</div><h2>CAIN</h2><p>An unpublished Bielik 1.5B fine-tune exploring strong persona conditioning and behavioural dependence on SFT-defined identity.</p><div class="update-log"><div class="update-item"><div class="update-date">2026</div><div class="update-copy"><strong>Persona-focused SFT experiment</strong><span>Strong conversational identity, architecture-awareness data and light post-SFT reinforcement work.</span></div></div><div class="update-item"><div class="update-date">Research hold</div><div class="update-copy"><strong>Broader persona-study concept</strong><span>Not published while the experiment is reframed around persona persistence and behavioural dependence.</span></div></div></div></div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
203
  <section><div class="wrap">
204
- <div class="eyebrow">Research direction</div><h3 class="section-title">Persona as something deeper than writing style.</h3>
205
- <div class="prose"><p>CAIN was trained with a deliberately opinionated, informal persona, light RL and SFT data that included awareness of its own architecture. It tended toward antihero choices, self-sacrifice, unusual trade-offs and dry humour.</p><p>The project remains unpublished because the more interesting direction is a broader study of persona persistence and behavioural dependence in LLMs.</p></div>
 
 
 
 
 
206
  </div></section>
207
  </article>
208
 
 
209
  <article class="detail" id="model-eris">
210
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">Eris</div></div></div>
211
  <div class="wrap detail-hero">
 
124
  </div>
125
  </main>
126
 
127
+
128
  <article class="detail" id="model-polmath">
129
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">PolMATH</div></div></div>
130
+ <div class="wrap detail-hero">
131
+ <div class="eyebrow">Numerical representation research</div>
132
+ <h2>PolMATH</h2>
133
+ <p>A custom language-model experiment built around a simple question: should a number exist inside a language model only as text?</p>
134
+ <div class="update-log">
135
+ <div class="update-item"><div class="update-date">2025</div><div class="update-copy"><strong>Initial architecture and training experiments</strong><span>Custom numerical channel, numerical decoding, routing and number↔text association tests.</span></div></div>
136
+ <div class="update-item"><div class="update-date">Current</div><div class="update-copy"><strong>Publication-oriented rework</strong><span>Original branch paused; a cleaner protocol and English-language continuation are being prepared.</span></div></div>
137
+ </div>
138
+ </div>
139
+
140
  <section><div class="wrap">
141
+ <div class="eyebrow">Why it existed</div>
142
+ <h3 class="section-title">Value and notation were treated as related, but not identical.</h3>
143
  <div class="prose">
144
+ <p>PolMATH did not begin as an adapter attached to an existing pretrained model. The transformer stack, tokenizer behaviour, numerical codec, embedding fusion, auxiliary losses, routing logic and numerical decoding path were implemented as one experimental architecture.</p>
145
+ <p>Text could contain a dedicated numerical position while the actual value travelled through a structured numerical representation. That representation could encode integer and fractional digits, lengths, decimal form and exponent structure rather than collapsing every number into one scalar or leaving it entirely to BPE tokenization.</p>
146
+ <p>This made the experiment less about “doing mathematics” and more about asking whether a language model benefits from separating <em>what a number means</em> from <em>how that number happened to be written</em>.</p>
147
  </div>
 
148
  </div></section>
149
+
150
  <section><div class="wrap">
151
+ <div class="eyebrow">Architecture</div>
152
+ <h3 class="section-title">Language decides that a number belongs here. A numerical path decides which number.</h3>
153
+ <div class="version-grid">
154
+ <div class="version-card"><small>Tokenizer</small><h4>[NUM] position</h4><p>Written numbers could be normalized into a dedicated numerical position while preserving linguistic context around them.</p></div>
155
+ <div class="version-card"><small>Numeric codec</small><h4>Structured representation</h4><p>Integer digits, fractional digits, notation flags and exponent structure were represented explicitly.</p></div>
156
+ <div class="version-card"><small>Fusion</small><h4>Gated numerical injection</h4><p>Numerical features were projected into the token representation and gated per position instead of affecting every token indiscriminately.</p></div>
157
+ <div class="version-card"><small>Objectives</small><h4>Multiple numerical heads</h4><p>Categorical structure prediction and a bucket/regression path provided complementary ways of reconstructing values.</p></div>
158
+ <div class="version-card"><small>Alignment</small><h4>Numeric-only pressure</h4><p>An auxiliary objective aligned text and numerical representations at number positions and penalized numerical leakage outside them.</p></div>
159
+ <div class="version-card"><small>Generation</small><h4>Two-stage output</h4><p>The LM could first decide that a numerical token belongs in the sequence, then the numerical head could reconstruct the value.</p></div>
160
+ </div>
161
+ </div></section>
162
+
163
+ <section><div class="wrap">
164
+ <div class="eyebrow">Observed training behaviour</div>
165
+ <h3 class="section-title">The numerical channel became usable rather than decorative.</h3>
166
+ <div class="prose">
167
+ <p>The training pipeline reached the intended qualitative behaviours: matching number words to numeric values and back, deciding when a numerical position should appear despite ordinary text tokens dominating the corpus, and learning useful associations between values and facts.</p>
168
+ <p>One of the useful examples was the relation between <strong>1939</strong> and the outbreak of the Second World War. The interesting part was not only producing the number after the fact, but also retaining the association when the direction of the prompt was reversed.</p>
169
+ <p>Different written forms such as <strong>1939</strong>, <strong>1939.0</strong> and <strong>1.939e3</strong> could be learned as related expressions of the same value rather than as unrelated text fragments.</p>
170
+ </div>
171
+ <div class="callout"><strong>Important limitation</strong><p>PolMATH was not a general-purpose mathematical reasoner. Its value was narrower: it demonstrated that an explicit numerical path could coexist with ordinary language modelling and learn useful number-related behaviour.</p></div>
172
+ </div></section>
173
+
174
+ <section><div class="wrap">
175
+ <div class="eyebrow">Roadmap</div>
176
+ <h3 class="section-title">Paused, reorganized, and no longer intended to remain Polish-only.</h3>
177
  <div class="timeline">
178
+ <div class="timeline-item"><div class="timeline-year">2025</div><div class="timeline-body"><h4>Initial PolMATH</h4><p>Custom architecture and structured numerical representation experiments.</p></div></div>
179
+ <div class="timeline-item"><div class="timeline-year">2025–26</div><div class="timeline-body"><h4>Training and data experiments</h4><p>Numeric placement, notation equivalence, routing, generation and association tests.</p></div></div>
180
+ <div class="timeline-item"><div class="timeline-year">Current</div><div class="timeline-body"><h4>Original branch paused</h4><p>The work is being cleaned up rather than simply continued as another checkpoint.</p></div></div>
181
+ <div class="timeline-item"><div class="timeline-year">Next</div><div class="timeline-body"><h4>Publication-oriented continuation</h4><p>A cleaner experimental protocol and an English-language iteration are planned so the idea can be evaluated beyond a Polish-only setting.</p></div></div>
182
  </div>
183
  </div></section>
184
  </article>
185
 
186
+
187
+
188
  <article class="detail" id="model-oris660">
189
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">ORIS 660M</div></div></div>
190
  <div class="wrap detail-hero">
191
+ <div class="eyebrow">Structural compression and recovery</div>
192
+ <h2>ORIS 660M</h2>
193
+ <p>An experimental Polish base model built by heavily shortening Bielik-1.5B-v3 and then studying what survives, what collapses, and what can recover.</p>
194
+ <div class="stats"><div class="stat"><strong>660.13M</strong><span>parameters</span></div><div class="stat"><strong>32 → 12</strong><span>transformer blocks</span></div><div class="stat"><strong>621.984M</strong><span>benchmarked recovery labels</span></div><div class="stat"><strong>Paused</strong><span>research active</span></div></div>
195
+ <div class="update-log">
196
+ <div class="update-item"><div class="update-date">07 Aug 2026</div><div class="update-copy"><strong>ORIS 660M begins</strong><span>Depth-reduction experiment from Bielik-1.5B-v3.</span></div></div>
197
+ <div class="update-item"><div class="update-date">Recovery</div><div class="update-copy"><strong>Selected-vs-uniform architecture tests</strong><span>Layer identity proved important to recovery trajectory.</span></div></div>
198
+ <div class="update-item"><div class="update-date">Later work</div><div class="update-copy"><strong>Knowledge and data-control experiments</strong><span>Inherited-vs-random controls, Polish diagnostics and a staged 22B-token continuation plan.</span></div></div>
199
+ </div>
200
  </div>
201
+
202
+ <section><div class="wrap">
203
+ <div class="eyebrow">Origin</div>
204
+ <h3 class="section-title">Not a conventionally scaled-down 660M transformer.</h3>
205
+ <div class="prose">
206
+ <p>The lineage is <strong>Qwen 2.5 → Bielik-1.5B-v3 → ORIS 660M</strong>. The starting teacher had 32 transformer blocks. ORIS retained only 12 while preserving the teacher's hidden width, attention geometry, tokenizer, embeddings, feed-forward dimensions and LM head.</p>
207
+ <p>The selected blocks were <strong>0, 1, 2, 3, 4, 21, 23, 24, 25, 29, 30 and 31</strong>. They were not chosen by simply keeping every second or third layer. Small ablations and calibration tests were used to identify less-sensitive regions and compare candidate 9–12 layer structures.</p>
208
+ <p>The extreme jump from teacher layer 4 to layer 21 means that downstream blocks receive representations unlike those they originally saw during pretraining. That mismatch is one of the central features of the experiment.</p>
209
+ </div>
210
+ </div></section>
211
+
212
+ <section><div class="wrap">
213
+ <div class="eyebrow">Step 0</div>
214
+ <h3 class="section-title">The student became faster and lighter — and language modelling collapsed.</h3>
215
+ <div class="table-wrap"><table>
216
+ <thead><tr><th>Model</th><th>Parameters</th><th>Loss</th><th>Perplexity</th><th>Throughput</th><th>Peak VRAM</th></tr></thead>
217
+ <tbody>
218
+ <tr class="highlight"><td>ORIS 660M step0</td><td>660M</td><td>7.1613</td><td>1288.58</td><td>17.29k tok/s</td><td>1.54 GiB</td></tr>
219
+ <tr><td>Bielik-1.5B</td><td>1.596B</td><td>2.2692</td><td>9.67</td><td>6.69k tok/s</td><td>3.29 GiB</td></tr>
220
+ </tbody>
221
+ </table></div>
222
+ <p class="note">Fixed evaluation on 128 sequences of length 1024.</p>
223
+ <div class="callout"><strong>The point was not pruning without damage.</strong><p>The network was strongly damaged. The experiment became: can an inherited but structurally broken network reorganize itself through ordinary next-token training?</p></div>
224
+ </div></section>
225
+
226
  <section><div class="wrap">
227
+ <div class="eyebrow">Architecture selection</div>
228
+ <h3 class="section-title">The identity of retained layers materially changed recovery.</h3>
229
+ <div class="table-wrap"><table>
230
+ <thead><tr><th>Architecture</th><th>Initial validation loss</th><th>After ~2.03M training tokens</th></tr></thead>
231
+ <tbody>
232
+ <tr class="highlight"><td>Selected 12-layer ORIS</td><td>7.2737</td><td>4.3134</td></tr>
233
+ <tr><td>Uniform 12-layer control</td><td>8.9560</td><td>5.4880</td></tr>
234
+ <tr><td>Selected 10-layer candidate</td><td>—</td><td>~4.9100</td></tr>
235
+ </tbody>
236
+ </table></div>
237
+ <div class="prose" style="margin-top:24px">
238
+ <p>These experiments were intentionally small and do not establish that the chosen structure is globally optimal. They do show that equal depth does not imply equal recoverability.</p>
239
+ <p>Main recovery used standard causal language modelling. Teacher/student KL was useful for screening candidate structures, but teacher logits were not the training objective during the main continuation run.</p>
240
+ </div>
241
+ <div class="stats"><div class="stat"><strong>589.248M</strong><span>earlier recovery labels</span></div><div class="stat"><strong>32.736M</strong><span>V2.5A continuation labels</span></div><div class="stat"><strong>621.984M</strong><span>cumulative benchmarked labels</span></div><div class="stat"><strong>~1×</strong><span>roughly one recovery label per student parameter</span></div></div>
242
+ </div></section>
243
+
244
+ <section><div class="wrap">
245
+ <div class="eyebrow">Generation behaviour</div>
246
+ <h3 class="section-title">Loss recovered faster than stable generation.</h3>
247
+ <div class="prose">
248
+ <p>Continued training restored grammatical Polish, reasonable local continuations and recognizable semantic associations, but free generation remained unstable.</p>
249
+ <p>Observed issues included greedy repetition loops, topic drift, weak long-range coherence, hallucinated dates and numerical sequences, web/forum residue, metadata-like fragments, premature EOS at some checkpoints and high sensitivity to sampling strategy.</p>
250
+ <p>One of the strongest lessons was that <strong>lower loss did not translate monotonically into better generation</strong>. Syntax, factual accessibility, calibration, semantic control and free-generation stability could improve at different rates.</p>
251
+ </div>
252
+ </div></section>
253
+
254
+ <section><div class="wrap">
255
+ <div class="eyebrow">V2.5A benchmark snapshot</div>
256
+ <h3 class="section-title">A recovery snapshot, not a leaderboard claim.</h3>
257
+ <div class="table-wrap"><table>
258
+ <thead><tr><th>Task</th><th>ORIS V2.5A</th></tr></thead>
259
+ <tbody>
260
+ <tr><td>Belebele accuracy</td><td>0.2256</td></tr>
261
+ <tr><td>8tags accuracy</td><td>0.1757</td></tr>
262
+ <tr><td>PoLeMo2 in accuracy</td><td>0.4155</td></tr>
263
+ <tr><td>PoLeMo2 out accuracy</td><td>0.3684</td></tr>
264
+ <tr><td>DYK binary F1</td><td>0.1268</td></tr>
265
+ <tr><td>PSC binary F1</td><td>0.1627</td></tr>
266
+ <tr><td>PPC accuracy</td><td>0.4180</td></tr>
267
+ <tr><td>CBD macro F1</td><td>0.1291</td></tr>
268
+ <tr><td>KLEJ NER accuracy</td><td>0.1025</td></tr>
269
+ <tr class="highlight"><td>PolQA reranking</td><td>0.5157</td></tr>
270
+ </tbody>
271
+ </table></div>
272
+ </div></section>
273
+
274
+ <section><div class="wrap">
275
+ <div class="eyebrow">Later diagnostics</div>
276
+ <h3 class="section-title">Competitive signals and very obvious failure modes appeared together.</h3>
277
+ <h4>Project-specific MC / continuation probe</h4>
278
+ <div class="table-wrap"><table>
279
+ <thead><tr><th>Model</th><th>Params</th><th>Raw MC</th><th>Normalized MC</th><th>Continuation NLL</th></tr></thead>
280
+ <tbody><tr class="highlight"><td>ORIS 660M</td><td>660M</td><td>0.750</td><td>0.500</td><td>3.3525</td></tr><tr><td>Qra-1b</td><td>~1.10B</td><td>0.917</td><td>0.583</td><td>1.5715</td></tr></tbody>
281
+ </table></div>
282
+ <p class="note">Qra remained clearly stronger; the comparison tested whether ORIS had returned to a meaningful operating regime.</p>
283
+
284
+ <h4 style="margin-top:34px">SpeakLeash polish_mc diagnostic — 0-shot, 100 examples per task</h4>
285
+ <div class="table-wrap"><table>
286
+ <thead><tr><th>Model</th><th>Parameters</th><th>Accuracy</th><th>Normalized accuracy</th><th>F1</th></tr></thead>
287
+ <tbody><tr class="highlight"><td>ORIS 660M</td><td>660M</td><td>0.399</td><td>0.386</td><td>0.0116</td></tr><tr><td>APT3-1B-Base</td><td>~1B</td><td>0.330</td><td>0.300</td><td>0.3215</td></tr></tbody>
288
+ </table></div>
289
+ <div class="table-wrap"><table>
290
+ <thead><tr><th>Task</th><th>ORIS</th><th>APT3</th></tr></thead>
291
+ <tbody>
292
+ <tr><td>PoLeMo2 in accuracy</td><td><strong>0.43</strong></td><td>0.34</td></tr>
293
+ <tr><td>PoLeMo2 out accuracy</td><td>0.33</td><td><strong>0.38</strong></td></tr>
294
+ <tr><td>8tags accuracy</td><td>0.12</td><td><strong>0.20</strong></td></tr>
295
+ <tr><td>Belebele accuracy</td><td>0.23</td><td><strong>0.25</strong></td></tr>
296
+ <tr><td>KLEJ NER accuracy</td><td><strong>0.31</strong></td><td>0.05</td></tr>
297
+ <tr><td>PolQA reranking accuracy</td><td><strong>0.62</strong></td><td>0.39</td></tr>
298
+ <tr><td>PPC accuracy</td><td><strong>0.44</strong></td><td>0.42</td></tr>
299
+ <tr><td>PSC accuracy</td><td><strong>0.68</strong></td><td>0.32</td></tr>
300
+ </tbody>
301
+ </table></div>
302
+ <div class="callout"><strong>The shape of the errors mattered more than the aggregate.</strong><p>ORIS could show surprisingly high accuracy on some tasks while binary F1 collapsed because of severe class bias. That is exactly why these results are diagnostic rather than a claim of general superiority over APT3.</p></div>
303
+ </div></section>
304
+
305
+ <section><div class="wrap">
306
+ <div class="eyebrow">Knowledge recovery</div>
307
+ <h3 class="section-title">Was knowledge destroyed, degraded, inaccessible — or relearned?</h3>
308
+ <div class="prose">
309
+ <p>An inherited 12-layer ORIS student was compared against an architecture-matched model initialized from scratch. Both had comparable structural capacity; only one inherited the teacher's pretrained parameter structure.</p>
310
+ </div>
311
+ <div class="stats"><div class="stat"><strong>24.7 → 35.0%</strong><span>inherited fact-level majority accuracy</span></div><div class="stat"><strong>0.154 → ~0.366</strong><span>teacher-ranking Spearman</span></div><div class="stat"><strong>~10%</strong><span>random control accuracy</span></div><div class="stat"><strong>-0.109</strong><span>random control Spearman</span></div></div>
312
+ <div class="table-wrap"><table><thead><tr><th>ENTITY_ONLY subset</th><th>Before</th><th>After</th></tr></thead><tbody><tr class="highlight"><td>Accuracy</td><td>20.0%</td><td>41.5%</td></tr><tr><td>Mean factual margin</td><td>-0.840</td><td>+0.120</td></tr></tbody></table></div>
313
+ <div class="prose" style="margin-top:24px">
314
+ <p>The result does not prove that complete symbolic facts remain intact inside individual weights. It does suggest that the inherited network begins recovery from a qualitatively different state than a random model with the same architecture.</p>
315
+ <p>A small generation probe captured the ambiguity: “Polska jest…” remained approximately correct, while “Kopernik był…” produced <strong>“Kopernik był w kosmosie.”</strong> The statement is false, but the astronomy/space semantic neighborhood survived.</p>
316
+ </div>
317
+ <div class="version-grid">
318
+ <div class="version-card"><small>Case 1</small><h4>Accessible</h4><p>Knowledge survives and remains directly usable.</p></div>
319
+ <div class="version-card"><small>Case 2</small><h4>Degraded</h4><p>The semantic region survives while precise retrieval fails.</p></div>
320
+ <div class="version-card"><small>Case 3</small><h4>Latent</h4><p>Useful structure may survive but no longer be reachable through the altered path.</p></div>
321
+ <div class="version-card"><small>Case 4</small><h4>Relearned</h4><p>The information genuinely has to be reacquired during continued training.</p></div>
322
+ </div>
323
+ </div></section>
324
+
325
+ <section><div class="wrap">
326
+ <div class="eyebrow">Data research</div>
327
+ <h3 class="section-title">The bottleneck moved from model architecture to controlling what the recovering model sees.</h3>
328
+ <div class="prose">
329
+ <p>Later generations repeatedly exposed artifacts from web-derived training data: forum fragments, metadata, SEO text, navigation elements, date-heavy sequences, list structures, transcripts, document templates and poorly separated topic boundaries.</p>
330
+ <p>The newer pipeline therefore separates linguistic quality, structural quality, corruption, noise, difficulty, domain, information density and <strong>knowledge value</strong> rather than compressing all of them into one score.</p>
331
+ <p>A difficult scientific or legal document can have high perplexity and many rare tokens while still carrying high knowledge value. A fluent generic paragraph can look clean while contributing almost none.</p>
332
+ </div>
333
  <div class="version-grid">
334
+ <div class="version-card"><small>01</small><h4>Canonical ingest</h4><p>Stable IDs, provenance, validation, sharding and raw preservation.</p></div>
335
+ <div class="version-card"><small>02</small><h4>Repair & structure</h4><p>Encoding repair, paragraph preservation, segmentation and conservative local deduplication.</p></div>
336
+ <div class="version-card"><small>03</small><h4>Signal extraction</h4><p>Language confidence, coherence, repetition, information density, rare-token statistics and perplexity.</p></div>
337
+ <div class="version-card"><small>04</small><h4>Multi-axis classification</h4><p>Separate estimates for quality, noise, knowledge value and reconstruction decisions.</p></div>
338
+ <div class="version-card"><small>05</small><h4>Knowledge structure</h4><p>Domain hierarchies, subdomains and knowledge flags.</p></div>
339
+ <div class="version-card"><small>06</small><h4>Corpus curation</h4><p>Global exact/fuzzy deduplication, redundancy control and later curriculum balancing.</p></div>
340
  </div>
341
  </div></section>
342
+
343
  <section><div class="wrap">
344
+ <div class="eyebrow">Snowball hypothesis</div>
345
+ <h3 class="section-title">Repeated small corpus defects may become directional model errors.</h3>
346
+ <div class="prose">
347
+ <p>The working hypothesis is not that one bad example damages a foundation model. The concern is cumulative: an unstable model repeatedly sees similar defects, small representation errors reinforce each other, and later data is interpreted through an already-shifted internal state.</p>
348
+ <p>In that view, <strong>data ordering can matter as much as data inclusion</strong>.</p>
349
+ </div>
350
  </div></section>
351
+
352
  <section><div class="wrap">
353
+ <div class="eyebrow">Planned V3 continuation</div>
354
+ <h3 class="section-title">22B tokens staged by what the recovering network needs next.</h3>
355
  <div class="timeline">
356
+ <div class="timeline-item"><div class="timeline-year">S1 · 6B</div><div class="timeline-body"><h4>Stabilization and structural recovery</h4><p>Coherent, lower-risk data intended to restore reliable autoregressive behaviour.</p></div></div>
357
+ <div class="timeline-item"><div class="timeline-year">S2 · 12B</div><div class="timeline-body"><h4>Main knowledge recovery and expansion</h4><p>Broader domains, higher information density and more difficult material after stabilization.</p></div></div>
358
+ <div class="timeline-item"><div class="timeline-year">S3 · 4B</div><div class="timeline-body"><h4>Consolidation and robustness</h4><p>Harder distributions introduced under tighter control.</p></div></div>
359
  </div>
360
+ <p class="note">This curriculum is proposed, not completed.</p>
361
+ </div></section>
362
+
363
+ <section><div class="wrap">
364
+ <div class="eyebrow">Why the branch is paused</div>
365
+ <h3 class="section-title">Because simply feeding the checkpoint more web text would answer the least interesting questions.</h3>
366
+ <div class="prose">
367
+ <p>ORIS is an independent, self-funded project developed primarily on privately owned compute. A substantial portion of pruning, validation, training and evaluation was performed locally, including on an RTX 5060 Ti.</p>
368
+ <p>The existing experiments suggest that a severely depth-reduced Polish model can recover useful language-model behaviour. The harder questions are now what exactly survives structural reduction, what is genuinely relearned, whether inherited knowledge can be distinguished from reacquired knowledge, and whether recovery can be controlled through data ordering at 10B–100B+ token scale.</p>
369
+ </div>
370
+ <div class="callout"><strong>Paused does not mean failed.</strong><p>The experiment clarified the next bottleneck. Work shifted toward data infrastructure, evaluation methodology, knowledge-recovery controls and additional compute before committing the 660M branch to a much larger run.</p></div>
371
  </div></section>
372
  </article>
373
 
374
+
375
+
376
  <article class="detail" id="model-smallc">
377
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">ORIS Small C</div></div></div>
378
  <div class="wrap detail-hero">
379
+ <div class="eyebrow">Compact Polish encoder</div>
380
+ <h2>ORIS Small C</h2>
381
+ <p>Originally developed under the working name ORIS Bert Small C. It uses BERT-style masked-language pretraining, but it is not a stock scaled-down BERT.</p>
382
+ <div class="stats"><div class="stat"><strong>25.41M</strong><span>parameters</span></div><div class="stat"><strong>6</strong><span>layers</span></div><div class="stat"><strong>128K</strong><span>Polish-oriented BPE</span></div><div class="stat"><strong>8.00B</strong><span>input pretraining tokens</span></div></div>
383
+ <div class="update-log">
384
+ <div class="update-item"><div class="update-date">16 Aug 2026</div><div class="update-copy"><strong>Small C checkpoint completed</strong><span>From-scratch 8B-token MLM pretraining completed.</span></div></div>
385
+ <div class="update-item"><div class="update-date">Downstream</div><div class="update-copy"><strong>KLEJ-style and document-filtering tests</strong><span>Established both the model's limitations and its practical pipeline advantage.</span></div></div>
386
+ <div class="update-item"><div class="update-date">Efficiency</div><div class="update-copy"><strong>Local encoder benchmarking</strong><span>Measured strong throughput and VRAM advantages on RTX 5060 Ti / Blackwell-oriented workloads.</span></div></div>
387
+ </div>
388
  </div>
389
+
390
+ <section><div class="wrap">
391
+ <div class="eyebrow">Why this model exists</div>
392
+ <h3 class="section-title">A production bottleneck accidentally became an architecture project.</h3>
393
+ <div class="prose">
394
+ <p>Dataset preparation for ORIS 660M became a throughput bottleneck, so a smaller encoder was built specifically for local filtering, categorization and scoring on NVIDIA Blackwell hardware.</p>
395
+ <p>What began as an internal pipeline tool became a compact Polish encoder intended primarily as a backbone for task-specific fine-tuning.</p>
396
+ </div>
397
+ <div class="callout"><strong>Not a sentence-embedding model out of the box.</strong><p>Raw mean-pooled embeddings are strongly anisotropic and are not recommended for zero-shot semantic search without additional contrastive or task-specific fine-tuning.</p></div>
398
+ </div></section>
399
+
400
+ <section><div class="wrap">
401
+ <div class="eyebrow">Architecture</div>
402
+ <h3 class="section-title">The “BERT” part describes the training style more than the architecture.</h3>
403
+ <div class="table-wrap"><table>
404
+ <thead><tr><th>Property</th><th>Value</th></tr></thead>
405
+ <tbody>
406
+ <tr><td>Model type</td><td>Custom Transformer encoder</td></tr>
407
+ <tr><td>Hidden size</td><td>384</td></tr>
408
+ <tr><td>Token embedding size</td><td>128</td></tr>
409
+ <tr><td>Attention heads</td><td>6</td></tr>
410
+ <tr><td>Context length</td><td>1024</td></tr>
411
+ <tr class="highlight"><td>Attention layout</td><td>256, 256, 1024, 256, 256, 256</td></tr>
412
+ <tr><td>Normalization</td><td>RMSNorm</td></tr>
413
+ <tr><td>Objective</td><td>Masked Language Modeling</td></tr>
414
+ <tr><td>Initialization</td><td>From random initialization</td></tr>
415
+ </tbody>
416
+ </table></div>
417
+ <div class="prose" style="margin-top:24px">
418
+ <p>Five of six layers use local 256-token attention. The third layer performs full 1024-token attention and acts as the main global mixing layer.</p>
419
+ <p>The vocabulary is intentionally large, but the token embedding dimension is only 128 while the encoder hidden size is 384. This factorization keeps the 128K vocabulary from dominating the parameter budget.</p>
420
+ <p>During MLM pretraining, vocabulary logits are calculated only for masked positions. Those hidden states are projected from 384 dimensions back into the 128-dimensional token space and scored using the tied token-embedding matrix.</p>
421
+ </div>
422
+ </div></section>
423
+
424
+ <section><div class="wrap">
425
+ <div class="eyebrow">Pretraining</div>
426
+ <h3 class="section-title">8.00B input tokens from scratch on a single local GPU.</h3>
427
+ <div class="table-wrap"><table>
428
+ <thead><tr><th>Setting</th><th>Value</th></tr></thead>
429
+ <tbody>
430
+ <tr><td>Sequence length</td><td>1024</td></tr>
431
+ <tr><td>Micro-batch size</td><td>16</td></tr>
432
+ <tr><td>Gradient accumulation</td><td>8</td></tr>
433
+ <tr><td>Tokens / optimizer update</td><td>131,072</td></tr>
434
+ <tr><td>Optimizer updates</td><td>61,036</td></tr>
435
+ <tr><td>Peak learning rate</td><td>3e-4</td></tr>
436
+ <tr><td>Mask probability</td><td>15%</td></tr>
437
+ <tr><td>Precision</td><td>BF16 autocast</td></tr>
438
+ <tr><td>GPU</td><td>RTX 5060 Ti 16GB</td></tr>
439
+ <tr><td>Training time</td><td>~13.44 h</td></tr>
440
+ <tr class="highlight"><td>Average throughput</td><td>~165.4K input tok/s</td></tr>
441
+ </tbody>
442
+ </table></div>
443
+ <div class="prose" style="margin-top:24px">
444
+ <p>Training data was primarily Polish MADLAD with an auxiliary mixture containing Wikipedia, OpenSubtitles PL, balanced NKJP and Polish legal/judicial text. The auxiliary pool represented roughly 15% of generated training sequences.</p>
445
+ <p>The tokenizer is a custom 128K BPE with NFKC normalization and Metaspace pre-tokenization. It is Polish-oriented but includes a broad Unicode alphabet for noisy web text.</p>
446
+ </div>
447
+ </div></section>
448
+
449
+ <section><div class="wrap">
450
+ <div class="eyebrow">Sentence-space limitation</div>
451
+ <h3 class="section-title">The raw encoder space is highly anisotropic under mean pooling.</h3>
452
+ <div class="table-wrap"><table>
453
+ <thead><tr><th>Model</th><th>Mean cosine for unrelated texts</th></tr></thead>
454
+ <tbody><tr class="highlight"><td>ORIS Small C</td><td>~0.987</td></tr><tr><td>HerBERT</td><td>~0.90</td></tr><tr><td>PolDense</td><td>~0.26</td></tr></tbody>
455
+ </table></div>
456
+ <p class="note">Internal diagnostic only; not a general encoder-quality benchmark.</p>
457
+ </div></section>
458
+
459
+ <section><div class="wrap">
460
+ <div class="eyebrow">Polish downstream benchmarks</div>
461
+ <h3 class="section-title">Competitive on several tasks, clearly weaker on others — at about one quarter of PolBERTa's size.</h3>
462
+ <div class="table-wrap"><table>
463
+ <thead><tr><th>Task</th><th>Metric</th><th>ORIS Small C</th><th>PolBERTa base</th></tr></thead>
464
+ <tbody>
465
+ <tr><td>NKJP-NER</td><td>Macro-F1</td><td>75.52</td><td><strong>84.36</strong></td></tr>
466
+ <tr class="highlight"><td>CDSC-E</td><td>Accuracy</td><td><strong>91.30</strong></td><td>91.00</td></tr>
467
+ <tr><td>CDSC-R</td><td>Spearman</td><td>88.18</td><td><strong>88.97</strong></td></tr>
468
+ <tr class="highlight"><td>CBD</td><td>F1(+)</td><td><strong>50.24</strong></td><td>43.75</td></tr>
469
+ <tr><td>PolEmo2.0-IN</td><td>Accuracy</td><td>83.33</td><td><strong>85.32</strong></td></tr>
470
+ <tr class="highlight"><td>PolEmo2.0-OUT</td><td>Accuracy</td><td><strong>65.59</strong></td><td>63.77</td></tr>
471
+ <tr><td>DYK</td><td>F1(+)</td><td>37.86</td><td><strong>46.31</strong></td></tr>
472
+ <tr><td>PSC</td><td>Macro-F1</td><td>57.28</td><td><strong>85.87</strong></td></tr>
473
+ <tr><td>AR</td><td>MAE ↓</td><td>0.5929</td><td><strong>0.5753</strong></td></tr>
474
+ </tbody>
475
+ </table></div>
476
+ <p class="note">Local evaluation using the same fixed procedure for ORIS Small C and PolBERTa base; not an official KLEJ leaderboard submission.</p>
477
+ </div></section>
478
+
479
+ <section><div class="wrap">
480
+ <div class="eyebrow">Document filtering</div>
481
+ <h3 class="section-title">The workload it was originally built for is where the design makes the most sense.</h3>
482
+ <div class="table-wrap"><table>
483
+ <thead><tr><th>Metric</th><th>mmBERT-base</th><th>ORIS Small C</th></tr></thead>
484
+ <tbody>
485
+ <tr class="highlight"><td>Decision Macro-F1</td><td>0.4334</td><td><strong>0.5015</strong></td></tr>
486
+ <tr><td>Decision accuracy</td><td>0.5185</td><td><strong>0.6296</strong></td></tr>
487
+ <tr><td>Training time</td><td>583.2 s</td><td><strong>124.4 s</strong></td></tr>
488
+ <tr><td>Peak VRAM</td><td>5.83 GiB</td><td><strong>0.52 GiB</strong></td></tr>
489
+ </tbody>
490
+ </table></div>
491
+ <h4 style="margin-top:32px">Production-style full pipeline</h4>
492
+ <div class="table-wrap"><table>
493
+ <thead><tr><th>Metric</th><th>mmBERT-base</th><th>ORIS Small C</th></tr></thead>
494
+ <tbody>
495
+ <tr class="highlight"><td>Full pipeline time</td><td>21.732 s</td><td><strong>4.408 s</strong></td></tr>
496
+ <tr><td>Documents / second</td><td>11.78</td><td><strong>58.07</strong></td></tr>
497
+ <tr><td>Mean latency / document</td><td>84.89 ms</td><td><strong>17.22 ms</strong></td></tr>
498
+ <tr><td>Peak VRAM</td><td>1.806 GiB</td><td><strong>0.252 GiB</strong></td></tr>
499
+ </tbody>
500
+ </table></div>
501
+ <div class="callout"><strong>~4.93× more documents per second in this workload.</strong><p>This is a task-specific pipeline result, not a claim that Small C is universally superior to larger encoders.</p></div>
502
+ </div></section>
503
+
504
  <section><div class="wrap">
505
+ <div class="eyebrow">Encoder efficiency</div>
506
+ <h3 class="section-title">The compact architecture also showed a large raw forward-pass advantage.</h3>
507
+ <div class="table-wrap"><table>
508
+ <thead><tr><th>Setting</th><th>ORIS Small C</th><th>mmBERT-small</th><th>Advantage</th></tr></thead>
509
+ <tbody>
510
+ <tr><td>Batch 1, 128 tokens</td><td><strong>3.828 ms</strong></td><td>17.153 ms</td><td><strong>4.48×</strong></td></tr>
511
+ <tr><td>Batch 1, 1024 tokens</td><td><strong>3.940 ms</strong></td><td>17.229 ms</td><td><strong>4.37×</strong></td></tr>
512
+ <tr class="highlight"><td>Batch 8, 1024 tokens</td><td><strong>1.25M tok/s</strong></td><td>159K tok/s</td><td><strong>7.85×</strong></td></tr>
513
+ </tbody>
514
+ </table></div>
515
+ <p class="note">Different tokenizers mean cross-model tokens/s should be interpreted carefully.</p>
516
  </div></section>
517
+
518
  <section><div class="wrap">
519
+ <div class="eyebrow">Roadmap and limitations</div>
520
+ <h3 class="section-title">Small C is useful precisely because its limits are explicit.</h3>
521
+ <div class="prose">
522
+ <p>Raw mean-pooled sentence representations have poor cosine-space separation. Retrieval, ranking and sentence similarity need additional fine-tuning. Larger Polish encoders remain stronger on several benchmarks, and five local-attention layers mean cross-window communication relies heavily on one global mixing layer.</p>
523
+ <p>The model is therefore best understood as a compact MLM-pretrained backbone for supervised tasks, filtering and feature extraction — not as a universal zero-shot embedding model.</p>
524
+ </div>
525
  <div class="timeline">
526
+ <div class="timeline-item"><div class="timeline-year">Origin</div><div class="timeline-body"><h4>ORIS 660M pipeline bottleneck</h4><p>Built to make local filtering and scoring fast enough to matter.</p></div></div>
527
+ <div class="timeline-item"><div class="timeline-year">16 Aug 2026</div><div class="timeline-body"><h4>Small C checkpoint completed</h4><p>8.00B-token from-scratch pretraining completed.</p></div></div>
528
+ <div class="timeline-item"><div class="timeline-year">Current</div><div class="timeline-body"><h4>Gated research release</h4><p>Used as a practical encoder backbone while downstream and representation-space behaviour are evaluated.</p></div></div>
529
+ <div class="timeline-item"><div class="timeline-year">Possible next</div><div class="timeline-body"><h4>Broader ORIS encoder line</h4><p>Further architecture, training and evaluation work may turn the concept-stage checkpoint into a fuller encoder family.</p></div></div>
530
  </div>
531
  </div></section>
532
  </article>
533
 
534
+
535
+
536
  <article class="detail" id="model-cain">
537
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">CAIN</div></div></div>
538
+ <div class="wrap detail-hero">
539
+ <div class="eyebrow">Persona research branch</div>
540
+ <h2>CAIN</h2>
541
+ <p>An unpublished Bielik 1.5B chat fine-tune built around a deliberately strong conversational identity rather than a generic assistant persona.</p>
542
+ <div class="update-log">
543
+ <div class="update-item"><div class="update-date">2026</div><div class="update-copy"><strong>Persona-focused SFT experiment</strong><span>Strong identity, architecture-awareness data and light post-SFT reinforcement work.</span></div></div>
544
+ <div class="update-item"><div class="update-date">Research hold</div><div class="update-copy"><strong>Reframed as a broader research question</strong><span>The interesting part became persona persistence and behavioural dependence rather than the checkpoint itself.</span></div></div>
545
+ </div>
546
+ </div>
547
+
548
+ <section><div class="wrap">
549
+ <div class="eyebrow">Why it existed</div>
550
+ <h3 class="section-title">Not “make a chatbot nicer”, but test how much identity can survive outside the examples that taught it.</h3>
551
+ <div class="prose">
552
+ <p>CAIN was fine-tuned on Bielik 1.5B for chat inference with a deliberately specific persona. Its SFT data included awareness of its own architecture and a less sanitized conversational style. Post-SFT reinforcement was present, but comparatively light rather than a deep RL stage.</p>
553
+ <p>The character tended toward antihero choices, self-sacrifice, unusual trade-offs and non-standard opinions. The humour was intentionally dry and informal; at times it described itself in the spirit of someone you could sit down with over a beer.</p>
554
+ <p>The interesting observation was not simply that a model can imitate a style. The broader research direction is whether an SFT-defined persona can become a persistent behavioural prior that changes choices, preferences and conversational decisions even when the prompt does not explicitly ask for the persona.</p>
555
+ </div>
556
+ </div></section>
557
+
558
+ <section><div class="wrap">
559
+ <div class="eyebrow">What is not claimed</div>
560
+ <h3 class="section-title">Persona is not consciousness, and a funny model is not evidence of an inner person.</h3>
561
+ <div class="prose">
562
+ <p>The project is intentionally framed as behaviour and conditioning research. Architecture-awareness in the training data means the model could talk about its own design; it does not establish self-awareness in a stronger sense.</p>
563
+ <p>Likewise, reduced guardrails and unusual opinions are properties of the training setup and observed behaviour, not proof of independent beliefs.</p>
564
+ </div>
565
+ <div class="callout"><strong>Why unpublished?</strong><p>The checkpoint itself is less interesting than the experiment it suggests: controlled comparisons of different personas, how stable they remain under distribution shift, and where style ends and decision-level behavioural dependence begins.</p></div>
566
+ </div></section>
567
+
568
  <section><div class="wrap">
569
+ <div class="eyebrow">Possible research direction</div>
570
+ <h3 class="section-title">From one strong persona to controlled persona persistence experiments.</h3>
571
+ <div class="timeline">
572
+ <div class="timeline-item"><div class="timeline-year">CAIN</div><div class="timeline-body"><h4>Single strong-persona SFT</h4><p>Exploratory chat model with intentionally distinctive preferences and humour.</p></div></div>
573
+ <div class="timeline-item"><div class="timeline-year">Next</div><div class="timeline-body"><h4>Matched persona controls</h4><p>Train multiple personas under comparable data budgets and evaluate persistence outside their training templates.</p></div></div>
574
+ <div class="timeline-item"><div class="timeline-year">Later</div><div class="timeline-body"><h4>Behavioural dependence</h4><p>Separate surface style, preference shifts, decision consistency and susceptibility to persona removal or replacement.</p></div></div>
575
+ </div>
576
  </div></section>
577
  </article>
578
 
579
+
580
  <article class="detail" id="model-eris">
581
  <div class="detail-top"><div class="wrap detail-bar"><button class="back" onclick="showIndex()">← All models</button><div class="detail-name">Eris</div></div></div>
582
  <div class="wrap detail-hero">