Aleksander commited on
Commit
c8d6a18
·
verified ·
1 Parent(s): a2f5229

Update index.html

Browse files
Files changed (1) hide show
  1. index.html +1446 -18
index.html CHANGED
@@ -1,19 +1,1447 @@
1
  <!doctype html>
2
- <html>
3
- <head>
4
- <meta charset="utf-8" />
5
- <meta name="viewport" content="width=device-width" />
6
- <title>My static Space</title>
7
- <link rel="stylesheet" href="style.css" />
8
- </head>
9
- <body>
10
- <div class="card">
11
- <h1>Welcome to your static Space!</h1>
12
- <p>You can modify this app directly by editing <i>index.html</i> in the Files and versions tab.</p>
13
- <p>
14
- Also don't forget to check the
15
- <a href="https://huggingface.co/docs/hub/spaces" target="_blank">Spaces documentation</a>.
16
- </p>
17
- </div>
18
- </body>
19
- </html>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  <!doctype html>
2
+ <html lang="en">
3
+ <head>
4
+ <meta charset="utf-8" />
5
+ <meta
6
+ name="viewport"
7
+ content="width=device-width, initial-scale=1"
8
+ />
9
+
10
+ <meta
11
+ name="description"
12
+ content="ORIS LLM Polish language model research project focused on architecture compression, recovery pretraining, data engineering and evaluation."
13
+ />
14
+
15
+ <title>ORIS LLM Polish Language Model Research</title>
16
+
17
+ <link
18
+ rel="stylesheet"
19
+ href="style.css"
20
+ />
21
+ </head>
22
+
23
+ <body>
24
+
25
+ <div class="background-grid"></div>
26
+ <div class="background-glow glow-one"></div>
27
+ <div class="background-glow glow-two"></div>
28
+
29
+ <header class="topbar">
30
+ <div class="container nav-inner">
31
+
32
+ <a
33
+ class="brand"
34
+ href="#top"
35
+ >
36
+ <span class="brand-mark">O</span>
37
+
38
+ <span class="brand-text">
39
+ ORIS
40
+ <small>LLM RESEARCH</small>
41
+ </span>
42
+ </a>
43
+
44
+ <nav class="nav-links">
45
+ <a href="#mission">Mission</a>
46
+ <a href="#architecture">Architecture</a>
47
+ <a href="#training">Training</a>
48
+ <a href="#evaluation">Evaluation</a>
49
+ <a href="#generations">Generations</a>
50
+ <a href="#data">Data</a>
51
+ <a href="#status">Status</a>
52
+ </nav>
53
+
54
+ </div>
55
+ </header>
56
+
57
+
58
+ <main id="top">
59
+
60
+ <!-- =====================================================
61
+ HERO
62
+ ====================================================== -->
63
+
64
+ <section class="hero">
65
+ <div class="container hero-grid">
66
+
67
+ <div class="hero-main">
68
+
69
+ <div class="eyebrow">
70
+ POLISH LANGUAGE MODEL RESEARCH
71
+ </div>
72
+
73
+ <h1>
74
+ ORIS
75
+ <span>660M</span>
76
+ </h1>
77
+
78
+ <p class="hero-lead">
79
+ An experimental Polish language model
80
+ developed as part of the ORIS family.
81
+ The project explores non-standard trade-offs
82
+ between architecture, data quality,
83
+ compute efficiency and practical language
84
+ model behavior.
85
+ </p>
86
+
87
+ <div class="hero-actions">
88
+ <a
89
+ class="button button-primary"
90
+ href="#current-results"
91
+ >
92
+ Current results
93
+ </a>
94
+
95
+ <a
96
+ class="button button-secondary"
97
+ href="#status"
98
+ >
99
+ Project status
100
+ </a>
101
+ </div>
102
+
103
+ <p class="hero-note">
104
+ Research project · Work in progress ·
105
+ Weights not released
106
+ </p>
107
+
108
+ </div>
109
+
110
+
111
+ <div class="hero-panel">
112
+
113
+ <div class="panel-label">
114
+ CURRENT MODEL
115
+ </div>
116
+
117
+ <div class="big-stat">
118
+ 660.13M
119
+ <small>parameters</small>
120
+ </div>
121
+
122
+ <div class="stat-grid">
123
+
124
+ <div>
125
+ <strong>12</strong>
126
+ <span>layers</span>
127
+ </div>
128
+
129
+ <div>
130
+ <strong>1536</strong>
131
+ <span>hidden size</span>
132
+ </div>
133
+
134
+ <div>
135
+ <strong>8960</strong>
136
+ <span>FFN size</span>
137
+ </div>
138
+
139
+ <div>
140
+ <strong>12 / 2</strong>
141
+ <span>Q / KV heads</span>
142
+ </div>
143
+
144
+ </div>
145
+
146
+ <div class="hero-lineage">
147
+ <span>Qwen 2.5</span>
148
+ <b>→</b>
149
+ <span>Bielik v3</span>
150
+ <b>→</b>
151
+ <span class="accent">ORIS 660M</span>
152
+ </div>
153
+
154
+ </div>
155
+
156
+ </div>
157
+ </section>
158
+
159
+
160
+ <!-- =====================================================
161
+ QUICK FACTS
162
+ ====================================================== -->
163
+
164
+ <section class="section compact-section">
165
+ <div class="container">
166
+
167
+ <div class="quick-grid">
168
+
169
+ <article class="metric-card">
170
+ <span class="metric-label">
171
+ Project started
172
+ </span>
173
+
174
+ <strong>
175
+ 07.08.2026
176
+ </strong>
177
+ </article>
178
+
179
+
180
+ <article class="metric-card">
181
+ <span class="metric-label">
182
+ Teacher model
183
+ </span>
184
+
185
+ <strong>
186
+ Bielik-1.5B-v3
187
+ </strong>
188
+ </article>
189
+
190
+
191
+ <article class="metric-card">
192
+ <span class="metric-label">
193
+ Depth reduction
194
+ </span>
195
+
196
+ <strong>
197
+ 32L → 12L
198
+ </strong>
199
+ </article>
200
+
201
+
202
+ <article class="metric-card">
203
+ <span class="metric-label">
204
+ Recovery observed
205
+ </span>
206
+
207
+ <strong>
208
+ 170.23M
209
+ </strong>
210
+
211
+ <small>
212
+ label tokens
213
+ </small>
214
+ </article>
215
+
216
+ </div>
217
+
218
+ </div>
219
+ </section>
220
+
221
+
222
+ <!-- =====================================================
223
+ MISSION
224
+ ====================================================== -->
225
+
226
+ <section
227
+ id="mission"
228
+ class="section"
229
+ >
230
+ <div class="container narrow">
231
+
232
+ <div class="section-heading">
233
+ <span>01 / MISSION</span>
234
+
235
+ <h2>
236
+ ORIS was not created simply
237
+ to be a “better model”.
238
+ </h2>
239
+ </div>
240
+
241
+ <div class="prose">
242
+
243
+ <p class="lead-paragraph">
244
+ The purpose of the ORIS family is to explore
245
+ unusual gaps in the Polish language model
246
+ ecosystem and establish new reference points
247
+ where the usual “bigger is better” comparison
248
+ is not enough.
249
+ </p>
250
+
251
+ <p>
252
+ ORIS treats the model itself as a research object.
253
+ The project is interested not only in final scores,
254
+ but also in the process: which parts of an
255
+ architecture can be removed, which representations
256
+ survive pruning, how a model recovers language
257
+ behavior after aggressive compression, how little
258
+ recovery compute can still produce meaningful
259
+ results, and how strongly data quality affects
260
+ a relatively small model.
261
+ </p>
262
+
263
+ <p>
264
+ The goal is not to claim superiority over other
265
+ Polish models. Instead, ORIS aims to share
266
+ measurements, failed experiments, architectural
267
+ decisions, limitations and practical observations
268
+ that may be useful when building future models.
269
+ </p>
270
+
271
+ <blockquote>
272
+ ORIS is intended to be both a model family
273
+ and a record of what worked, what failed,
274
+ and why.
275
+ </blockquote>
276
+
277
+ </div>
278
+
279
+ </div>
280
+ </section>
281
+
282
+
283
+ <!-- =====================================================
284
+ LINEAGE
285
+ ====================================================== -->
286
+
287
+ <section class="section dark-section">
288
+ <div class="container">
289
+
290
+ <div class="section-heading">
291
+ <span>02 / LINEAGE</span>
292
+
293
+ <h2>
294
+ Qwen 2.5 → Bielik v3 → ORIS 660M
295
+ </h2>
296
+ </div>
297
+
298
+ <div class="lineage">
299
+
300
+ <div class="lineage-node">
301
+ <small>FOUNDATION</small>
302
+ <strong>Qwen 2.5</strong>
303
+ </div>
304
+
305
+ <div class="lineage-arrow">
306
+
307
+ </div>
308
+
309
+ <div class="lineage-node">
310
+ <small>POLISH BASE</small>
311
+ <strong>Bielik-1.5B-v3</strong>
312
+ </div>
313
+
314
+ <div class="lineage-arrow">
315
+
316
+ </div>
317
+
318
+ <div class="lineage-node active">
319
+ <small>ORIS BRANCH</small>
320
+ <strong>ORIS 660M</strong>
321
+ </div>
322
+
323
+ </div>
324
+
325
+ <div class="prose wide-prose">
326
+
327
+ <p>
328
+ ORIS 660M was created by aggressively reducing
329
+ the depth of Bielik-1.5B-v3 while preserving
330
+ the original width of the representation,
331
+ attention geometry, embeddings, tokenizer
332
+ and language-model head.
333
+ </p>
334
+
335
+ <p>
336
+ The resulting model is not a conventionally
337
+ narrow 660M transformer. It preserves the
338
+ internal width of the much larger teacher,
339
+ but performs the transformation through
340
+ only twelve decoder blocks instead of thirty-two.
341
+ </p>
342
+
343
+ </div>
344
+
345
+ </div>
346
+ </section>
347
+
348
+
349
+ <!-- =====================================================
350
+ ARCHITECTURE
351
+ ====================================================== -->
352
+
353
+ <section
354
+ id="architecture"
355
+ class="section"
356
+ >
357
+ <div class="container">
358
+
359
+ <div class="section-heading">
360
+ <span>03 / ARCHITECTURE</span>
361
+
362
+ <h2>
363
+ A wide 660M model created primarily
364
+ through depth pruning.
365
+ </h2>
366
+ </div>
367
+
368
+ <div class="architecture-grid">
369
+
370
+ <div class="architecture-card">
371
+ <span>Total parameters</span>
372
+ <strong>660,131,328</strong>
373
+ </div>
374
+
375
+ <div class="architecture-card">
376
+ <span>Decoder layers</span>
377
+ <strong>12</strong>
378
+ </div>
379
+
380
+ <div class="architecture-card">
381
+ <span>Hidden size</span>
382
+ <strong>1536</strong>
383
+ </div>
384
+
385
+ <div class="architecture-card">
386
+ <span>Intermediate size</span>
387
+ <strong>8960</strong>
388
+ </div>
389
+
390
+ <div class="architecture-card">
391
+ <span>Attention heads</span>
392
+ <strong>12</strong>
393
+ </div>
394
+
395
+ <div class="architecture-card">
396
+ <span>KV heads</span>
397
+ <strong>2</strong>
398
+ </div>
399
+
400
+ <div class="architecture-card">
401
+ <span>Head dimension</span>
402
+ <strong>128</strong>
403
+ </div>
404
+
405
+ <div class="architecture-card">
406
+ <span>Vocabulary</span>
407
+ <strong>32,000</strong>
408
+ </div>
409
+
410
+ </div>
411
+
412
+
413
+ <div class="split-panel">
414
+
415
+ <div>
416
+ <div class="panel-label">
417
+ RETAINED TEACHER BLOCKS
418
+ </div>
419
+
420
+ <div class="layer-list">
421
+ <span>0</span>
422
+ <span>1</span>
423
+ <span>2</span>
424
+ <span>3</span>
425
+ <span>4</span>
426
+ <span>21</span>
427
+ <span>23</span>
428
+ <span>24</span>
429
+ <span>25</span>
430
+ <span>29</span>
431
+ <span>30</span>
432
+ <span>31</span>
433
+ </div>
434
+ </div>
435
+
436
+
437
+ <div class="prose">
438
+
439
+ <p>
440
+ The model does not simply keep the first
441
+ twelve layers of the teacher. The retained
442
+ path spans early, middle-late and final
443
+ regions of the original network.
444
+ </p>
445
+
446
+ <p>
447
+ Width, FFN size, attention geometry,
448
+ embeddings and LM head were preserved.
449
+ The main architectural intervention
450
+ was the reduction:
451
+ </p>
452
+
453
+ <div class="big-arrow">
454
+ 32 layers
455
+ <span>→</span>
456
+ 12 layers
457
+ </div>
458
+
459
+ </div>
460
+
461
+ </div>
462
+
463
+
464
+ <div class="compute-panel">
465
+
466
+ <div class="panel-label">
467
+ PARAMETER DISTRIBUTION
468
+ </div>
469
+
470
+ <div class="compute-grid">
471
+
472
+ <div>
473
+ <strong>~46.82M</strong>
474
+ <span>parameters / block</span>
475
+ </div>
476
+
477
+ <div>
478
+ <strong>~41.31M</strong>
479
+ <span>MLP / block</span>
480
+ </div>
481
+
482
+ <div>
483
+ <strong>~5.51M</strong>
484
+ <span>attention / block</span>
485
+ </div>
486
+
487
+ <div>
488
+ <strong>~88%</strong>
489
+ <span>block parameters in MLP</span>
490
+ </div>
491
+
492
+ </div>
493
+
494
+ </div>
495
+
496
+ </div>
497
+ </section>
498
+
499
+
500
+ <!-- =====================================================
501
+ INITIAL DAMAGE
502
+ ====================================================== -->
503
+
504
+ <section class="section dark-section">
505
+ <div class="container narrow">
506
+
507
+ <div class="section-heading">
508
+ <span>04 / STEP ZERO</span>
509
+
510
+ <h2>
511
+ Pruning preserved structure,
512
+ but heavily damaged language modeling.
513
+ </h2>
514
+ </div>
515
+
516
+ <div class="prose">
517
+
518
+ <p>
519
+ The initial 660M student was not expected
520
+ to behave like a ready language model.
521
+ Removing twenty transformer blocks severely
522
+ disrupted the sequence of transformations
523
+ learned by the teacher.
524
+ </p>
525
+
526
+ </div>
527
+
528
+
529
+ <div class="comparison-table">
530
+
531
+ <div class="comparison-row comparison-header">
532
+ <span>Model</span>
533
+ <span>Loss</span>
534
+ <span>Perplexity</span>
535
+ <span>Throughput</span>
536
+ <span>Peak VRAM</span>
537
+ </div>
538
+
539
+ <div class="comparison-row">
540
+ <strong>ORIS 660M step0</strong>
541
+ <span>7.1613</span>
542
+ <span>1288.58</span>
543
+ <span>17.29k tok/s</span>
544
+ <span>1.54 GiB</span>
545
+ </div>
546
+
547
+ <div class="comparison-row">
548
+ <strong>Bielik-1.5B</strong>
549
+ <span>2.2692</span>
550
+ <span>9.67</span>
551
+ <span>6.69k tok/s</span>
552
+ <span>3.29 GiB</span>
553
+ </div>
554
+
555
+ </div>
556
+
557
+ <p class="caption">
558
+ Fixed 128 × 1024 evaluation sequences.
559
+ These measurements describe the initial
560
+ post-pruning state and its teacher reference.
561
+ </p>
562
+
563
+ </div>
564
+ </section>
565
+
566
+
567
+ <!-- =====================================================
568
+ RECOVERY
569
+ ====================================================== -->
570
+
571
+ <section
572
+ id="training"
573
+ class="section"
574
+ >
575
+ <div class="container">
576
+
577
+ <div class="section-heading">
578
+ <span>05 / RECOVERY TRAINING</span>
579
+
580
+ <h2>
581
+ The first objective was recovery,
582
+ not scaling.
583
+ </h2>
584
+ </div>
585
+
586
+
587
+ <div class="training-grid">
588
+
589
+ <article class="training-card">
590
+
591
+ <span class="training-label">
592
+ Sequence length
593
+ </span>
594
+
595
+ <strong>1024</strong>
596
+
597
+ </article>
598
+
599
+
600
+ <article class="training-card">
601
+
602
+ <span class="training-label">
603
+ Micro batch
604
+ </span>
605
+
606
+ <strong>2</strong>
607
+
608
+ </article>
609
+
610
+
611
+ <article class="training-card">
612
+
613
+ <span class="training-label">
614
+ Gradient accumulation
615
+ </span>
616
+
617
+ <strong>256</strong>
618
+
619
+ </article>
620
+
621
+
622
+ <article class="training-card">
623
+
624
+ <span class="training-label">
625
+ Label tokens / update
626
+ </span>
627
+
628
+ <strong>523,776</strong>
629
+
630
+ </article>
631
+
632
+
633
+ <article class="training-card">
634
+
635
+ <span class="training-label">
636
+ Precision
637
+ </span>
638
+
639
+ <strong>BF16</strong>
640
+
641
+ </article>
642
+
643
+
644
+ <article class="training-card">
645
+
646
+ <span class="training-label">
647
+ Optimizer
648
+ </span>
649
+
650
+ <strong>Fused AdamW</strong>
651
+
652
+ </article>
653
+
654
+
655
+ <article class="training-card">
656
+
657
+ <span class="training-label">
658
+ Max learning rate
659
+ </span>
660
+
661
+ <strong>2e-5</strong>
662
+
663
+ </article>
664
+
665
+
666
+ <article class="training-card">
667
+
668
+ <span class="training-label">
669
+ Training throughput
670
+ </span>
671
+
672
+ <strong>~9.1k tok/s</strong>
673
+
674
+ </article>
675
+
676
+ </div>
677
+
678
+
679
+ <div
680
+ id="current-results"
681
+ class="milestone"
682
+ >
683
+
684
+ <div>
685
+ <span class="panel-label">
686
+ LATEST PRESERVED MILESTONE
687
+ </span>
688
+
689
+ <h3>
690
+ Checkpoint 325
691
+ </h3>
692
+ </div>
693
+
694
+ <div class="milestone-stats">
695
+
696
+ <div>
697
+ <strong>170,227,200</strong>
698
+ <span>label tokens consumed</span>
699
+ </div>
700
+
701
+ <div>
702
+ <strong>190,720</strong>
703
+ <span>documents read</span>
704
+ </div>
705
+
706
+ <div>
707
+ <strong>190,720</strong>
708
+ <span>EOS tokens added</span>
709
+ </div>
710
+
711
+ <div>
712
+ <strong>3.7726</strong>
713
+ <span>last training-loss window</span>
714
+ </div>
715
+
716
+ </div>
717
+
718
+ </div>
719
+
720
+
721
+ <div class="warning-box">
722
+ <strong>
723
+ Important:
724
+ </strong>
725
+
726
+ <p>
727
+ The 3.7726 value is a training-window loss
728
+ on the clean continuation corpus.
729
+ It must not be directly compared with the
730
+ historical step0 fixed-evaluation loss of 7.1613.
731
+ </p>
732
+ </div>
733
+
734
+ </div>
735
+ </section>
736
+
737
+
738
+ <!-- =====================================================
739
+ DATA PROBLEM
740
+ ====================================================== -->
741
+
742
+ <section
743
+ id="data"
744
+ class="section dark-section"
745
+ >
746
+ <div class="container">
747
+
748
+ <div class="section-heading">
749
+ <span>06 / DATA</span>
750
+
751
+ <h2>
752
+ The bottleneck moved from model training
753
+ to data infrastructure.
754
+ </h2>
755
+ </div>
756
+
757
+ <div class="prose wide-prose">
758
+
759
+ <p class="lead-paragraph">
760
+ ORIS 660M exposed a problem that is easy
761
+ to underestimate when experiments begin
762
+ at small scale: training hundreds of millions
763
+ of tokens is straightforward compared with
764
+ reliably managing tens or hundreds of billions
765
+ of tokens.
766
+ </p>
767
+
768
+ <p>
769
+ Early recovery data still contained many
770
+ structures inherited from web crawls:
771
+ boilerplate, duplicated fragments, listing pages,
772
+ SEO text, navigation elements, disconnected
773
+ topics, encoding damage, forum residue,
774
+ article metadata and repeated templates.
775
+ </p>
776
+
777
+ <p>
778
+ Continuing to scale training without addressing
779
+ those representations would simply make the
780
+ model better at reproducing the same defects.
781
+ The project therefore changed direction
782
+ temporarily from model compute to data compute.
783
+ </p>
784
+
785
+ </div>
786
+
787
+
788
+ <div class="data-pipeline">
789
+
790
+ <div class="data-stage keep">
791
+ <strong>KEEP</strong>
792
+ <span>
793
+ usable coherent documents
794
+ </span>
795
+ </div>
796
+
797
+ <div class="data-stage clean">
798
+ <strong>CLEAN</strong>
799
+ <span>
800
+ recoverable core + removable noise
801
+ </span>
802
+ </div>
803
+
804
+ <div class="data-stage split">
805
+ <strong>SPLIT</strong>
806
+ <span>
807
+ multiple recoverable document regions
808
+ </span>
809
+ </div>
810
+
811
+ <div class="data-stage drop">
812
+ <strong>DROP</strong>
813
+ <span>
814
+ no sufficiently useful core
815
+ </span>
816
+ </div>
817
+
818
+ </div>
819
+
820
+
821
+ <div class="prose wide-prose">
822
+
823
+ <p>
824
+ The V2 data pipeline combines deterministic
825
+ structural signals with semantic analysis.
826
+ It measures repetition, fragmentation,
827
+ cross-document templates, paragraph coherence,
828
+ boundary candidates, corrupted characters
829
+ and learned mmBERT signals.
830
+ </p>
831
+
832
+ <p>
833
+ The goal is not to reduce the corpus as much
834
+ as possible. The goal is to preserve useful
835
+ Polish text while separating raw source material
836
+ from training-ready units.
837
+ </p>
838
+
839
+ </div>
840
+
841
+
842
+ <div class="data-number-grid">
843
+
844
+ <div>
845
+ <strong>1.39M+</strong>
846
+ <span>
847
+ documents observed in one V2 shard run
848
+ </span>
849
+ </div>
850
+
851
+ <div>
852
+ <strong>1.55B+</strong>
853
+ <span>
854
+ approximate raw tokens in that run
855
+ </span>
856
+ </div>
857
+
858
+ <div>
859
+ <strong>~17.2</strong>
860
+ <span>
861
+ documents / second
862
+ </span>
863
+ </div>
864
+
865
+ <div>
866
+ <strong>99.77%</strong>
867
+ <span>
868
+ token retention in conservative KEEP hygiene
869
+ </span>
870
+ </div>
871
+
872
+ </div>
873
+
874
+ </div>
875
+ </section>
876
+
877
+
878
+ <!-- =====================================================
879
+ EVALUATION
880
+ ====================================================== -->
881
+
882
+ <section
883
+ id="evaluation"
884
+ class="section"
885
+ >
886
+ <div class="container">
887
+
888
+ <div class="section-heading">
889
+ <span>07 / EVALUATION</span>
890
+
891
+ <h2>
892
+ Diagnostic results, not leaderboard claims.
893
+ </h2>
894
+ </div>
895
+
896
+ <div class="prose wide-prose">
897
+
898
+ <p>
899
+ ORIS 660M has not completed a full pretraining
900
+ schedule and is not presented as a finished model.
901
+ Current evaluations are intended to understand
902
+ recovery behavior rather than establish a final rank.
903
+ </p>
904
+
905
+ <p>
906
+ The current checkpoint was compared against
907
+ larger Polish base models on small controlled
908
+ probes and on a limited subset of the
909
+ Open PL LLM Leaderboard multiple-choice suite.
910
+ </p>
911
+
912
+ </div>
913
+
914
+
915
+ <div class="benchmark-title">
916
+ Custom diagnostic benchmark
917
+ </div>
918
+
919
+ <div class="comparison-table">
920
+
921
+ <div class="comparison-row comparison-header">
922
+ <span>Model</span>
923
+ <span>Parameters</span>
924
+ <span>Raw MC acc.</span>
925
+ <span>Normalized acc.</span>
926
+ <span>Continuation NLL</span>
927
+ </div>
928
+
929
+ <div class="comparison-row highlighted">
930
+ <strong>ORIS 660M step325</strong>
931
+ <span>660.13M</span>
932
+ <span>0.750</span>
933
+ <span>0.500</span>
934
+ <span>3.3525</span>
935
+ </div>
936
+
937
+ <div class="comparison-row">
938
+ <strong>Qra-1b</strong>
939
+ <span>~1.10B</span>
940
+ <span>0.917</span>
941
+ <span>0.583</span>
942
+ <span>1.5715</span>
943
+ </div>
944
+
945
+ </div>
946
+
947
+
948
+ <p class="caption">
949
+ Project-specific lightweight probe.
950
+ Not an official benchmark suite.
951
+ </p>
952
+
953
+
954
+ <div class="benchmark-title">
955
+ Open PL LLM Leaderboard task subset
956
+ </div>
957
+
958
+ <div class="comparison-table">
959
+
960
+ <div class="comparison-row comparison-header">
961
+ <span>Model</span>
962
+ <span>Parameters</span>
963
+ <span>Accuracy</span>
964
+ <span>Acc. norm</span>
965
+ <span>F1</span>
966
+ </div>
967
+
968
+ <div class="comparison-row highlighted">
969
+ <strong>ORIS 660M step325</strong>
970
+ <span>660.13M</span>
971
+ <span>0.399</span>
972
+ <span>0.386</span>
973
+ <span>0.0116</span>
974
+ </div>
975
+
976
+ <div class="comparison-row">
977
+ <strong>APT3-1B-Base</strong>
978
+ <span>~1B</span>
979
+ <span>0.330</span>
980
+ <span>0.300</span>
981
+ <span>0.3215</span>
982
+ </div>
983
+
984
+ </div>
985
+
986
+ <p class="caption">
987
+ `polish_mc`, 0-shot, limited to 100 examples
988
+ per task. These are diagnostic subset results,
989
+ not official leaderboard scores.
990
+ </p>
991
+
992
+
993
+ <div class="warning-box benchmark-warning">
994
+
995
+ <strong>
996
+ Aggregate accuracy is not the whole story.
997
+ </strong>
998
+
999
+ <p>
1000
+ ORIS showed severe class bias on some binary
1001
+ tasks. For example, high raw accuracy could
1002
+ coexist with near-zero F1. The model therefore
1003
+ cannot be described as generally outperforming
1004
+ APT3 based on this diagnostic run.
1005
+ </p>
1006
+
1007
+ </div>
1008
+
1009
+
1010
+ <div class="task-grid">
1011
+
1012
+ <div>
1013
+ <span>PoLeMo2 in</span>
1014
+ <strong>0.43</strong>
1015
+ </div>
1016
+
1017
+ <div>
1018
+ <span>PoLeMo2 out</span>
1019
+ <strong>0.33</strong>
1020
+ </div>
1021
+
1022
+ <div>
1023
+ <span>Belebele MC</span>
1024
+ <strong>0.23</strong>
1025
+ </div>
1026
+
1027
+ <div>
1028
+ <span>PolQA reranking</span>
1029
+ <strong>0.62</strong>
1030
+ </div>
1031
+
1032
+ <div>
1033
+ <span>PPC</span>
1034
+ <strong>0.44</strong>
1035
+ </div>
1036
+
1037
+ <div>
1038
+ <span>PSC</span>
1039
+ <strong>0.68</strong>
1040
+ </div>
1041
+
1042
+ </div>
1043
+
1044
+ </div>
1045
+ </section>
1046
+
1047
+
1048
+ <!-- =====================================================
1049
+ GENERATIONS
1050
+ ====================================================== -->
1051
+
1052
+ <section
1053
+ id="generations"
1054
+ class="section dark-section"
1055
+ >
1056
+ <div class="container">
1057
+
1058
+ <div class="section-heading">
1059
+ <span>08 / GENERATION BEHAVIOR</span>
1060
+
1061
+ <h2>
1062
+ Language ability returned faster
1063
+ than generation stability.
1064
+ </h2>
1065
+ </div>
1066
+
1067
+ <div class="prose wide-prose">
1068
+
1069
+ <p>
1070
+ The current model can produce fluent Polish
1071
+ locally and sampling can yield surprisingly
1072
+ coherent paragraphs. However, greedy decoding
1073
+ still exposes strong autoregressive instability.
1074
+ </p>
1075
+
1076
+ <p>
1077
+ Repetition loops, semantic drift and remnants
1078
+ of web-style representations remain clearly visible.
1079
+ </p>
1080
+
1081
+ </div>
1082
+
1083
+
1084
+ <div class="generation-card">
1085
+
1086
+ <div class="generation-header">
1087
+ <span>PROMPT</span>
1088
+
1089
+ <strong>
1090
+ Inflacja oznacza wzrost
1091
+ </strong>
1092
+ </div>
1093
+
1094
+ <div class="generation-body">
1095
+
1096
+ <span class="generation-mode">
1097
+ GREEDY
1098
+ </span>
1099
+
1100
+ <p>
1101
+ o 1,5 proc. W 2013 r. w Polsce...
1102
+ w porównaniu do 1,5 proc...
1103
+ w porównaniu do 1,5 proc...
1104
+ </p>
1105
+
1106
+ </div>
1107
+
1108
+ <div class="generation-comment">
1109
+ Repetition collapse remains clearly visible.
1110
+ </div>
1111
+
1112
+ </div>
1113
+
1114
+
1115
+ <div class="generation-card">
1116
+
1117
+ <div class="generation-header">
1118
+ <span>PROMPT</span>
1119
+
1120
+ <strong>
1121
+ Kiedy rano otworzył okno, zobaczył
1122
+ </strong>
1123
+ </div>
1124
+
1125
+ <div class="generation-body">
1126
+
1127
+ <span class="generation-mode sampled-mode">
1128
+ SAMPLED
1129
+ </span>
1130
+
1131
+ <p>
1132
+ ...someone lying on the floor.
1133
+ She became frightened and ran to the window...
1134
+ </p>
1135
+
1136
+ </div>
1137
+
1138
+ <div class="generation-comment">
1139
+ Sampling can recover a recognizable narrative
1140
+ structure, but the model still drifts and mixes
1141
+ unrelated representations.
1142
+ </div>
1143
+
1144
+ </div>
1145
+
1146
+
1147
+ <div class="generation-card">
1148
+
1149
+ <div class="generation-header">
1150
+ <span>PROMPT</span>
1151
+
1152
+ <strong>
1153
+ Model językowy przewiduje kolejne tokeny
1154
+ na podstawie
1155
+ </strong>
1156
+ </div>
1157
+
1158
+ <div class="generation-body">
1159
+
1160
+ <span class="generation-mode">
1161
+ GREEDY
1162
+ </span>
1163
+
1164
+ <p>
1165
+ ...tokens. Tokens are also available
1166
+ as QR codes. Tokens are also available...
1167
+ </p>
1168
+
1169
+ </div>
1170
+
1171
+ <div class="generation-comment">
1172
+ Local syntax is functional, while global semantics
1173
+ and repetition control are still incomplete.
1174
+ </div>
1175
+
1176
+ </div>
1177
+
1178
+
1179
+ <div class="behavior-grid">
1180
+
1181
+ <div class="positive">
1182
+ <strong>Recovered</strong>
1183
+
1184
+ <ul>
1185
+ <li>Polish syntax</li>
1186
+ <li>local sentence fluency</li>
1187
+ <li>basic factual associations</li>
1188
+ <li>usable sampled continuations</li>
1189
+ </ul>
1190
+ </div>
1191
+
1192
+
1193
+ <div class="negative">
1194
+ <strong>Still unstable</strong>
1195
+
1196
+ <ul>
1197
+ <li>greedy repetition</li>
1198
+ <li>topic drift</li>
1199
+ <li>long-range coherence</li>
1200
+ <li>binary class calibration</li>
1201
+ <li>web-pattern residue</li>
1202
+ </ul>
1203
+ </div>
1204
+
1205
+ </div>
1206
+
1207
+ </div>
1208
+ </section>
1209
+
1210
+
1211
+ <!-- =====================================================
1212
+ WHY PAUSED
1213
+ ====================================================== -->
1214
+
1215
+ <section
1216
+ id="status"
1217
+ class="section"
1218
+ >
1219
+ <div class="container narrow">
1220
+
1221
+ <div class="section-heading">
1222
+ <span>09 / CURRENT STATUS</span>
1223
+
1224
+ <h2>
1225
+ ORIS 660M is currently paused,
1226
+ not finished.
1227
+ </h2>
1228
+ </div>
1229
+
1230
+ <div class="status-box">
1231
+
1232
+ <div class="status-indicator">
1233
+ <span></span>
1234
+ PAUSED / RESEARCH ACTIVE
1235
+ </div>
1236
+
1237
+ <p>
1238
+ ORIS 660M reached a useful recovery milestone,
1239
+ but continuing the same training immediately
1240
+ would not answer the most important remaining
1241
+ questions.
1242
+ </p>
1243
+
1244
+ <p>
1245
+ The current limitation is not simply GPU training
1246
+ throughput. The project now requires a reliable
1247
+ strategy for selecting, cleaning, reconstructing,
1248
+ deduplicating and independently evaluating
1249
+ datasets at a scale approaching or exceeding
1250
+ <strong>100B tokens</strong>.
1251
+ </p>
1252
+
1253
+ <p>
1254
+ Several components of that strategy have already
1255
+ been prototyped, including semantic document
1256
+ classification, structural reconstruction,
1257
+ conservative cleaning and persistent metadata.
1258
+ The complete large-scale pipeline has not yet
1259
+ been fully deployed.
1260
+ </p>
1261
+
1262
+ <p>
1263
+ Additional compute infrastructure is also planned.
1264
+ Until both data throughput and evaluation
1265
+ throughput are adequate, pushing ORIS 660M
1266
+ through a much larger training run would provide
1267
+ weaker evidence than improving the experimental
1268
+ foundation first.
1269
+ </p>
1270
+
1271
+ </div>
1272
+
1273
+ </div>
1274
+ </section>
1275
+
1276
+
1277
+ <!-- =====================================================
1278
+ ROADMAP
1279
+ ====================================================== -->
1280
+
1281
+ <section class="section dark-section">
1282
+ <div class="container">
1283
+
1284
+ <div class="section-heading">
1285
+ <span>10 / ROADMAP</span>
1286
+
1287
+ <h2>
1288
+ What comes next.
1289
+ </h2>
1290
+ </div>
1291
+
1292
+ <div class="roadmap">
1293
+
1294
+ <div class="roadmap-item done">
1295
+ <span>01</span>
1296
+
1297
+ <div>
1298
+ <strong>
1299
+ 660M architecture
1300
+ </strong>
1301
+
1302
+ <p>
1303
+ Depth-pruned ORIS branch established.
1304
+ </p>
1305
+ </div>
1306
+ </div>
1307
+
1308
+
1309
+ <div class="roadmap-item done">
1310
+ <span>02</span>
1311
+
1312
+ <div>
1313
+ <strong>
1314
+ Recovery training
1315
+ </strong>
1316
+
1317
+ <p>
1318
+ Initial language recovery demonstrated.
1319
+ </p>
1320
+ </div>
1321
+ </div>
1322
+
1323
+
1324
+ <div class="roadmap-item active">
1325
+ <span>03</span>
1326
+
1327
+ <div>
1328
+ <strong>
1329
+ Data infrastructure
1330
+ </strong>
1331
+
1332
+ <p>
1333
+ Build scalable Polish corpus selection,
1334
+ cleaning and reconstruction.
1335
+ </p>
1336
+ </div>
1337
+ </div>
1338
+
1339
+
1340
+ <div class="roadmap-item">
1341
+ <span>04</span>
1342
+
1343
+ <div>
1344
+ <strong>
1345
+ Deduplication and curriculum
1346
+ </strong>
1347
+
1348
+ <p>
1349
+ Finalize training-ready corpus policies.
1350
+ </p>
1351
+ </div>
1352
+ </div>
1353
+
1354
+
1355
+ <div class="roadmap-item">
1356
+ <span>05</span>
1357
+
1358
+ <div>
1359
+ <strong>
1360
+ Large-scale continuation
1361
+ </strong>
1362
+
1363
+ <p>
1364
+ Resume ORIS training with a significantly
1365
+ larger controlled dataset.
1366
+ </p>
1367
+ </div>
1368
+ </div>
1369
+
1370
+
1371
+ <div class="roadmap-item">
1372
+ <span>06</span>
1373
+
1374
+ <div>
1375
+ <strong>
1376
+ Full evaluation
1377
+ </strong>
1378
+
1379
+ <p>
1380
+ Polish benchmarks, held-out language
1381
+ modeling, repetition, calibration
1382
+ and generation stability.
1383
+ </p>
1384
+ </div>
1385
+ </div>
1386
+
1387
+ </div>
1388
+
1389
+ </div>
1390
+ </section>
1391
+
1392
+
1393
+ <!-- =====================================================
1394
+ FINAL
1395
+ ====================================================== -->
1396
+
1397
+ <section class="section final-section">
1398
+ <div class="container narrow">
1399
+
1400
+ <div class="final-mark">
1401
+ ORIS
1402
+ </div>
1403
+
1404
+ <h2>
1405
+ What remains when a large model
1406
+ loses most of its depth?
1407
+ </h2>
1408
+
1409
+ <p>
1410
+ ORIS 660M is the first measurable answer
1411
+ produced by this project.
1412
+ </p>
1413
+
1414
+ <p>
1415
+ It is not a final model.
1416
+ It is a research checkpoint.
1417
+ </p>
1418
+
1419
+ <div class="final-meta">
1420
+ <span>Started 07.08.2026</span>
1421
+ <span>Polish LLM research</span>
1422
+ <span>Work in progress</span>
1423
+ </div>
1424
+
1425
+ </div>
1426
+ </section>
1427
+
1428
+ </main>
1429
+
1430
+
1431
+ <footer>
1432
+ <div class="container footer-inner">
1433
+
1434
+ <span>
1435
+ ORIS LLM
1436
+ </span>
1437
+
1438
+ <span>
1439
+ Architecture compression · Recovery ·
1440
+ Data engineering · Evaluation
1441
+ </span>
1442
+
1443
+ </div>
1444
+ </footer>
1445
+
1446
+ </body>
1447
+ </html>