bertfil commited on
Commit
321f1bb
Β·
verified Β·
1 Parent(s): 3f5c322

Backfill veritas replication logbook

Browse files
README.md CHANGED
@@ -1,10 +1,18 @@
1
- ---
2
- title: BhahZSDowo
3
- emoji: πŸ“Š
4
- colorFrom: blue
5
- colorTo: yellow
6
- sdk: static
7
- pinned: false
8
- ---
9
-
10
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: "Repro - Benchmarking at the Edge of Comprehension"
3
+ emoji: πŸ”¬
4
+ colorFrom: blue
5
+ colorTo: green
6
+ sdk: static
7
+ pinned: false
8
+ tags:
9
+ - trackio
10
+ - trackio-logbook
11
+ - open-experiment
12
+ - icml2026-repro
13
+ - paper-BhahZSDowo
14
+ ---
15
+
16
+ # Repro - Benchmarking at the Edge of Comprehension
17
+
18
+ An open experiment logbook, published with [Trackio](https://github.com/gradio-app/trackio).
bucket-icon.svg ADDED
index.html CHANGED
@@ -1,19 +1,54 @@
1
- <!doctype html>
2
- <html>
3
- <head>
4
- <meta charset="utf-8" />
5
- <meta name="viewport" content="width=device-width" />
6
- <title>My static Space</title>
7
- <link rel="stylesheet" href="style.css" />
8
- </head>
9
- <body>
10
- <div class="card">
11
- <h1>Welcome to your static Space!</h1>
12
- <p>You can modify this app directly by editing <i>index.html</i> in the Files and versions tab.</p>
13
- <p>
14
- Also don't forget to check the
15
- <a href="https://huggingface.co/docs/hub/spaces" target="_blank">Spaces documentation</a>.
16
- </p>
17
- </div>
18
- </body>
19
- </html>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <!doctype html>
2
+ <html lang="en">
3
+ <head>
4
+ <meta charset="utf-8" />
5
+ <meta name="viewport" content="width=device-width, initial-scale=1" />
6
+ <title>Repro - Benchmarking at the Edge of Comprehension</title>
7
+ <link rel="stylesheet" href="./logbook.css" />
8
+ </head>
9
+ <body>
10
+ <div id="app">
11
+ <aside id="sidebar">
12
+ <div id="book-head">
13
+ <img id="book-wordmark" src="./trackio-wordmark-dark.png" alt="" />
14
+ <div id="book-title" class="sr-only">Logbook</div>
15
+ </div>
16
+ <nav id="tree"></nav>
17
+ <div id="sidebar-foot" hidden>
18
+ <button id="connect-btn" type="button">
19
+ <span class="ico">β“˜</span> Collaborate with your agent
20
+ </button>
21
+ </div>
22
+ </aside>
23
+ <main id="content">
24
+ <div id="page"></div>
25
+ </main>
26
+ </div>
27
+
28
+ <div id="modal" hidden>
29
+ <div class="modal-backdrop"></div>
30
+ <div class="modal-card" role="dialog" aria-modal="true">
31
+ <div class="modal-head">
32
+ <div class="modal-title">
33
+ <img class="modal-logo" src="./trackio-logo.png" alt="" />
34
+ Collaborate with your agent
35
+ </div>
36
+ <div class="modal-actions">
37
+ <button id="copy-agent" class="btn">Copy for agent</button>
38
+ <button id="modal-close" class="btn icon" aria-label="Close">Γ—</button>
39
+ </div>
40
+ </div>
41
+ <div class="modal-body">
42
+ <p class="modal-intro">
43
+ Point your coding agent at this logbook. It reads a compact,
44
+ token-efficient version β€” and if you've given it write access to this
45
+ Space, it can add findings that sync back automatically.
46
+ </p>
47
+ <ol id="connect-steps"></ol>
48
+ </div>
49
+ </div>
50
+ </div>
51
+
52
+ <script src="./logbook.js"></script>
53
+ </body>
54
+ </html>
logbook.css ADDED
@@ -0,0 +1,1602 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ :root {
2
+ --bg: #ffffff;
3
+ --paper: #fdfcf9;
4
+ --panel: #ffffff;
5
+ --ink: #1f2937;
6
+ --muted: #6b7280;
7
+ --line: #e5e7eb;
8
+ --accent: #f97316;
9
+ --accent-strong: #ea580c;
10
+ --accent-soft: #fff7ed;
11
+ --accent-line: rgba(249, 115, 22, 0.16);
12
+ --grid-line: rgba(31, 41, 55, 0.045);
13
+ --code-bg: #f3f4f6;
14
+ --radius: 12px;
15
+ --serif: ui-serif, "Iowan Old Style", "Palatino Linotype", Georgia, serif;
16
+ --sans: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, Helvetica, Arial,
17
+ sans-serif;
18
+ --mono: "SFMono-Regular", "Cascadia Mono", "JetBrains Mono", Menlo, Consolas,
19
+ ui-monospace, monospace;
20
+ }
21
+
22
+ * {
23
+ box-sizing: border-box;
24
+ }
25
+
26
+ html,
27
+ body {
28
+ margin: 0;
29
+ padding: 0;
30
+ }
31
+
32
+ html {
33
+ scroll-behavior: smooth;
34
+ }
35
+
36
+ body {
37
+ background: var(--bg);
38
+ color: var(--ink);
39
+ font-family: var(--sans);
40
+ font-size: 13px;
41
+ line-height: 1.65;
42
+ -webkit-font-smoothing: antialiased;
43
+ }
44
+
45
+ #app {
46
+ display: flex;
47
+ min-height: 100vh;
48
+ }
49
+
50
+ /* ---- sidebar (composition-book cover) ---- */
51
+ #sidebar {
52
+ width: 280px;
53
+ flex: 0 0 280px;
54
+ background: #17181c;
55
+ color: #e7e7ea;
56
+ position: sticky;
57
+ top: 0;
58
+ height: 100vh;
59
+ overflow-y: auto;
60
+ padding: 22px 16px;
61
+ display: flex;
62
+ flex-direction: column;
63
+ }
64
+
65
+ #book-head {
66
+ display: flex;
67
+ align-items: center;
68
+ gap: 10px;
69
+ padding: 8px;
70
+ margin-bottom: 12px;
71
+ border-radius: 10px;
72
+ cursor: pointer;
73
+ transition: background 0.12s;
74
+ }
75
+ #book-head:hover {
76
+ background: rgba(255, 255, 255, 0.05);
77
+ }
78
+ #book-wordmark {
79
+ width: 154px;
80
+ height: auto;
81
+ object-fit: contain;
82
+ }
83
+ .sr-only {
84
+ position: absolute;
85
+ width: 1px;
86
+ height: 1px;
87
+ padding: 0;
88
+ margin: -1px;
89
+ overflow: hidden;
90
+ clip: rect(0, 0, 0, 0);
91
+ white-space: nowrap;
92
+ border: 0;
93
+ }
94
+
95
+ #tree {
96
+ flex: 1;
97
+ padding-top: 8px;
98
+ }
99
+
100
+ #tree a {
101
+ display: block;
102
+ padding: 6px 10px;
103
+ border-radius: 8px;
104
+ color: #c3c4cb;
105
+ text-decoration: none;
106
+ font-size: 14px;
107
+ transition: background 0.12s, color 0.12s;
108
+ }
109
+
110
+ #tree a:hover {
111
+ background: rgba(255, 255, 255, 0.06);
112
+ color: #ffffff;
113
+ }
114
+
115
+ #tree a.active {
116
+ background: rgba(249, 115, 22, 0.16);
117
+ color: #fdba74;
118
+ font-weight: 600;
119
+ }
120
+
121
+ #tree a .tree-mark {
122
+ color: #6b6d76;
123
+ }
124
+
125
+ #tree a:hover .tree-mark,
126
+ #tree a.active .tree-mark {
127
+ color: inherit;
128
+ opacity: 0.6;
129
+ }
130
+
131
+ #tree .depth-1 {
132
+ padding-left: 22px;
133
+ }
134
+ #tree .depth-2 {
135
+ padding-left: 34px;
136
+ }
137
+ #tree .depth-3 {
138
+ padding-left: 46px;
139
+ }
140
+
141
+
142
+ /* ---- content ---- */
143
+ #content {
144
+ flex: 1;
145
+ min-width: 0;
146
+ padding: 48px 40px 120px;
147
+ background-color: var(--paper);
148
+ background-image:
149
+ linear-gradient(var(--grid-line) 1px, transparent 1px),
150
+ linear-gradient(90deg, var(--grid-line) 1px, transparent 1px);
151
+ background-size: 26px 26px;
152
+ background-position: center top;
153
+ }
154
+
155
+ #page {
156
+ width: 100%;
157
+ min-width: 0;
158
+ max-width: 1052px;
159
+ margin: 0 auto;
160
+ }
161
+
162
+ .page-section {
163
+ scroll-margin-top: 40px;
164
+ padding: 0 0 35px;
165
+ margin: 0 0 32px;
166
+ }
167
+
168
+ .page-section:last-child {
169
+ margin-bottom: 0;
170
+ }
171
+
172
+ .page-layout {
173
+ display: grid;
174
+ grid-template-columns: minmax(0, 760px) 248px;
175
+ gap: 44px;
176
+ align-items: start;
177
+ }
178
+
179
+ .page-body {
180
+ min-width: 0;
181
+ }
182
+
183
+ .resource-anchor {
184
+ display: block;
185
+ height: 0;
186
+ overflow: hidden;
187
+ }
188
+
189
+ /* ---- pinned notes ---- */
190
+ .pinned-notes {
191
+ margin: 30px 0 0;
192
+ }
193
+ .pinned-notes-list .cell {
194
+ margin: 0;
195
+ border-color: rgba(249, 115, 22, 0.55);
196
+ }
197
+ .pinned-notes-list .cell + .cell {
198
+ margin-top: 12px;
199
+ }
200
+ .cell.pinned-source {
201
+ border-color: rgba(249, 115, 22, 0.55);
202
+ }
203
+ .book-intro.has-pinned-notes {
204
+ border-bottom: none;
205
+ padding-bottom: 22px;
206
+ margin-bottom: 30px;
207
+ }
208
+ .book-intro.book-intro-tight {
209
+ border-bottom: none;
210
+ padding-bottom: 4px;
211
+ margin-bottom: 20px;
212
+ }
213
+
214
+ #page h1 {
215
+ font-family: var(--serif);
216
+ font-size: 34px;
217
+ line-height: 1.15;
218
+ letter-spacing: -0.02em;
219
+ margin: 0 0 8px;
220
+ overflow-wrap: anywhere;
221
+ }
222
+
223
+ #page .page-section:not(.book-intro) h1 {
224
+ font-size: 26px;
225
+ }
226
+
227
+ #page h2 {
228
+ font-family: var(--serif);
229
+ font-size: 24px;
230
+ margin: 36px 0 10px;
231
+ }
232
+
233
+ #page h3 {
234
+ font-size: 17px;
235
+ font-weight: 700;
236
+ margin: 26px 0 2px;
237
+ letter-spacing: -0.01em;
238
+ }
239
+
240
+ #page h3::before {
241
+ content: "";
242
+ display: inline-block;
243
+ width: 7px;
244
+ height: 7px;
245
+ border-radius: 2px;
246
+ background: var(--accent);
247
+ margin-right: 10px;
248
+ vertical-align: middle;
249
+ transform: translateY(-1px);
250
+ }
251
+
252
+ #page p {
253
+ margin: 10px 0;
254
+ }
255
+
256
+ #page blockquote {
257
+ margin: 14px 0;
258
+ padding: 2px 16px;
259
+ border-left: 3px solid #fdba74;
260
+ color: var(--muted);
261
+ }
262
+
263
+ #page hr {
264
+ display: none;
265
+ }
266
+
267
+ #page code {
268
+ font-family: var(--mono);
269
+ font-size: 0.86em;
270
+ background: var(--code-bg);
271
+ padding: 2px 6px;
272
+ border-radius: 6px;
273
+ }
274
+
275
+ #page pre {
276
+ max-width: 100%;
277
+ background: var(--code-bg);
278
+ border: 1px solid var(--line);
279
+ border-radius: var(--radius);
280
+ padding: 14px 16px;
281
+ overflow-x: auto;
282
+ }
283
+ #page pre code {
284
+ background: none;
285
+ padding: 0;
286
+ font-size: 11.5px;
287
+ }
288
+
289
+ /* ---- code blocks + collapsible accordion ---- */
290
+ #page pre.hl {
291
+ background: #17181c;
292
+ border: none;
293
+ color: #e7e7ea;
294
+ font-size: 13px;
295
+ line-height: 1.58;
296
+ }
297
+ #page pre.hl code {
298
+ color: inherit;
299
+ font-family: var(--mono);
300
+ }
301
+ .code-accordion {
302
+ border: 1px solid rgba(249, 115, 22, 0.2);
303
+ border-radius: 8px;
304
+ overflow: hidden;
305
+ margin: 12px 0;
306
+ background: #17181c;
307
+ }
308
+ .code-accordion summary {
309
+ list-style: none;
310
+ cursor: pointer;
311
+ display: flex;
312
+ align-items: center;
313
+ gap: 9px;
314
+ padding: 9px 12px;
315
+ font-family: var(--mono);
316
+ font-size: 11.5px;
317
+ font-weight: 700;
318
+ color: #e7e7ea;
319
+ background: #1e2027;
320
+ user-select: none;
321
+ overflow-wrap: anywhere;
322
+ }
323
+ .code-accordion summary::-webkit-details-marker {
324
+ display: none;
325
+ }
326
+ .code-accordion summary::after {
327
+ content: "β–Έ";
328
+ margin-left: auto;
329
+ color: var(--accent);
330
+ transition: transform 0.12s;
331
+ transform: rotate(180deg);
332
+ }
333
+ .code-accordion[open] summary::after {
334
+ transform: rotate(90deg);
335
+ }
336
+ .code-accordion .code-ico {
337
+ color: var(--accent);
338
+ font-weight: 700;
339
+ }
340
+ .code-accordion pre.hl {
341
+ margin: 0;
342
+ border-radius: 0;
343
+ border: none;
344
+ border-top: 1px solid rgba(249, 115, 22, 0.16);
345
+ }
346
+ .tok-comment {
347
+ color: #7a7d87;
348
+ font-style: italic;
349
+ }
350
+ .tok-string {
351
+ color: #a5d6a7;
352
+ }
353
+ .tok-keyword {
354
+ color: #fdba74;
355
+ }
356
+ .tok-number {
357
+ color: #7fd0e0;
358
+ }
359
+
360
+ #page a {
361
+ color: var(--accent);
362
+ }
363
+
364
+ #page ul {
365
+ padding-left: 20px;
366
+ }
367
+
368
+ .ts {
369
+ font-family: var(--mono);
370
+ font-size: 12px;
371
+ color: var(--muted);
372
+ background: none;
373
+ padding: 0;
374
+ }
375
+
376
+ /* ---- notebook-style cells ---- */
377
+ .cell {
378
+ max-width: 100%;
379
+ border: 1px solid var(--line);
380
+ border-radius: 10px;
381
+ background: rgba(255, 255, 255, 0.86);
382
+ margin: 18px 0;
383
+ overflow: hidden;
384
+ box-shadow: 0 2px 10px rgba(31, 41, 55, 0.035);
385
+ }
386
+ .cell-head {
387
+ display: flex;
388
+ justify-content: space-between;
389
+ gap: 16px;
390
+ align-items: center;
391
+ padding: 14px 18px;
392
+ background: rgba(255, 255, 255, 0.92);
393
+ border-bottom: 1px solid var(--line);
394
+ }
395
+ .cell-head.no-title {
396
+ justify-content: flex-end;
397
+ padding-top: 10px;
398
+ padding-bottom: 10px;
399
+ }
400
+ .cell-title {
401
+ flex: 1;
402
+ min-width: 0;
403
+ font-size: 13px;
404
+ font-weight: 650;
405
+ color: var(--ink);
406
+ line-height: 1.35;
407
+ overflow-wrap: anywhere;
408
+ }
409
+ .cell-meta {
410
+ flex: 0 0 auto;
411
+ display: flex;
412
+ align-items: center;
413
+ gap: 10px;
414
+ font-family: var(--sans);
415
+ font-size: 13px;
416
+ color: var(--muted);
417
+ }
418
+ .cell-open {
419
+ flex: 0 0 auto;
420
+ font-family: var(--mono);
421
+ font-size: 12px;
422
+ color: var(--accent);
423
+ text-decoration: none;
424
+ }
425
+ .cell-open:hover {
426
+ color: var(--accent-strong);
427
+ }
428
+ .cell-body {
429
+ min-width: 0;
430
+ padding: 14px 18px 18px;
431
+ }
432
+ .cell.dashboard .cell-body {
433
+ padding: 0;
434
+ }
435
+ #page .cell-body h1,
436
+ #page .cell-body h2 {
437
+ font-family: var(--sans);
438
+ font-size: 17px;
439
+ font-weight: 700;
440
+ letter-spacing: -0.01em;
441
+ line-height: 1.35;
442
+ margin: 22px 0 6px;
443
+ }
444
+ #page .cell-body > :first-child {
445
+ margin-top: 0;
446
+ }
447
+ #page .cell-body > :last-child {
448
+ margin-bottom: 0;
449
+ }
450
+ .cell.code .cell-head {
451
+ background: #fbfbfc;
452
+ }
453
+ .figure-fit {
454
+ position: relative;
455
+ overflow: hidden;
456
+ min-height: 160px;
457
+ border: 1px solid var(--line);
458
+ border-radius: 8px;
459
+ background: #fff;
460
+ }
461
+ .figure-fit[hidden] {
462
+ display: none;
463
+ }
464
+ .figure-fit:fullscreen,
465
+ .figure-fit:-webkit-full-screen {
466
+ width: 100%;
467
+ height: 100%;
468
+ border: none;
469
+ border-radius: 0;
470
+ }
471
+ .figure-frame {
472
+ display: block;
473
+ width: 100%;
474
+ min-height: 160px;
475
+ border: none;
476
+ background: #fff;
477
+ }
478
+ .figure-frame[hidden],
479
+ .figure-raw[hidden] {
480
+ display: none;
481
+ }
482
+ .fig-switch {
483
+ position: relative;
484
+ display: inline-flex;
485
+ flex: 0 0 auto;
486
+ border: 1px solid var(--line);
487
+ border-radius: 999px;
488
+ background: var(--code-bg);
489
+ padding: 2px;
490
+ }
491
+ .fig-switch button {
492
+ position: relative;
493
+ z-index: 1;
494
+ flex: 1;
495
+ min-width: 62px;
496
+ border: none;
497
+ background: none;
498
+ font-family: var(--sans);
499
+ font-size: 12px;
500
+ font-weight: 600;
501
+ color: var(--muted);
502
+ padding: 3px 12px;
503
+ border-radius: 999px;
504
+ cursor: pointer;
505
+ transition: color 0.15s;
506
+ }
507
+ .fig-switch button.active {
508
+ color: var(--accent-strong);
509
+ }
510
+ .fig-switch-thumb {
511
+ position: absolute;
512
+ top: 2px;
513
+ bottom: 2px;
514
+ left: 2px;
515
+ width: calc(50% - 2px);
516
+ border-radius: 999px;
517
+ background: var(--panel);
518
+ border: 1px solid rgba(249, 115, 22, 0.35);
519
+ box-shadow: 0 1px 4px rgba(31, 41, 55, 0.08);
520
+ transition: transform 0.18s ease;
521
+ }
522
+ .fig-switch.raw .fig-switch-thumb {
523
+ transform: translateX(100%);
524
+ }
525
+ #page .figure-raw pre {
526
+ margin: 0;
527
+ max-height: 420px;
528
+ overflow: auto;
529
+ font-family: var(--mono);
530
+ font-size: 13px;
531
+ line-height: 1.55;
532
+ background: var(--code-bg);
533
+ border: 1px solid var(--line);
534
+ border-radius: 8px;
535
+ padding: 12px 14px;
536
+ }
537
+ /* ---- figure fullscreen ---- */
538
+ .cell-fullscreen {
539
+ position: relative;
540
+ display: inline-flex;
541
+ flex: 0 0 auto;
542
+ }
543
+ .cell-fullscreen-btn {
544
+ display: inline-flex;
545
+ align-items: center;
546
+ justify-content: center;
547
+ width: 26px;
548
+ height: 26px;
549
+ padding: 0;
550
+ border: 1px solid var(--line);
551
+ border-radius: 999px;
552
+ background: var(--code-bg);
553
+ color: var(--muted);
554
+ cursor: pointer;
555
+ transition: color 0.15s, border-color 0.15s, background 0.15s;
556
+ }
557
+ .cell-fullscreen-btn:hover {
558
+ color: var(--accent-strong);
559
+ border-color: rgba(249, 115, 22, 0.35);
560
+ background: var(--accent-soft);
561
+ }
562
+ .cell-fullscreen-btn svg {
563
+ width: 14px;
564
+ height: 14px;
565
+ }
566
+ /* ---- copyable snippets ---- */
567
+ .snippet {
568
+ position: relative;
569
+ }
570
+ .copy-snippet {
571
+ position: absolute;
572
+ top: 7px;
573
+ right: 8px;
574
+ width: 24px;
575
+ height: 24px;
576
+ border: none;
577
+ border-radius: 6px;
578
+ background: rgba(255, 255, 255, 0.08);
579
+ color: #9a9da8;
580
+ font-size: 12px;
581
+ line-height: 1;
582
+ cursor: pointer;
583
+ opacity: 0;
584
+ transition: opacity 0.12s, color 0.12s, background 0.12s;
585
+ }
586
+ .snippet:hover .copy-snippet,
587
+ .jp-out:hover .copy-snippet,
588
+ .figure-raw:hover .copy-snippet,
589
+ .code-accordion summary:hover .copy-snippet {
590
+ opacity: 1;
591
+ }
592
+ .copy-snippet:hover {
593
+ color: #ffffff;
594
+ background: rgba(255, 255, 255, 0.16);
595
+ }
596
+ .copy-snippet.copied {
597
+ color: #52d08a;
598
+ opacity: 1;
599
+ }
600
+ .code-accordion .code-name {
601
+ user-select: text;
602
+ cursor: text;
603
+ }
604
+ .jp-out,
605
+ .figure-raw {
606
+ position: relative;
607
+ }
608
+ .jp-out .copy-snippet,
609
+ .figure-raw .copy-snippet {
610
+ background: var(--code-bg);
611
+ color: var(--muted);
612
+ border: 1px solid var(--line);
613
+ }
614
+ .jp-out .copy-snippet:hover,
615
+ .figure-raw .copy-snippet:hover {
616
+ color: var(--accent-strong);
617
+ background: var(--panel);
618
+ }
619
+
620
+ /* ---- jupyter-style code cells ---- */
621
+ .jp {
622
+ border: 1px solid var(--line);
623
+ border-radius: 10px;
624
+ overflow: hidden;
625
+ margin: 12px 0;
626
+ background: var(--panel);
627
+ }
628
+ .jp-gutter {
629
+ flex: 0 0 46px;
630
+ padding: 13px 0 0 13px;
631
+ font-family: var(--mono);
632
+ font-size: 10.5px;
633
+ letter-spacing: 0.07em;
634
+ text-transform: uppercase;
635
+ font-weight: 600;
636
+ user-select: none;
637
+ }
638
+ .jp-in {
639
+ display: flex;
640
+ background: #17181c;
641
+ }
642
+ .jp-in .jp-gutter {
643
+ color: #6f727d;
644
+ }
645
+ .jp-in-body {
646
+ flex: 1;
647
+ min-width: 0;
648
+ }
649
+ #page .jp-in-body pre.hl {
650
+ margin: 0;
651
+ border: none;
652
+ border-radius: 0;
653
+ background: none;
654
+ padding: 12px 16px 12px 0;
655
+ }
656
+ .jp-in-body .code-accordion {
657
+ margin: 0;
658
+ border: none;
659
+ border-top: 1px solid rgba(255, 255, 255, 0.09);
660
+ border-radius: 0;
661
+ background: none;
662
+ }
663
+ .jp-in-body .code-accordion summary {
664
+ background: none;
665
+ padding: 9px 16px 9px 0;
666
+ }
667
+ .jp-in-body .code-accordion pre.hl {
668
+ border-top: 1px solid rgba(255, 255, 255, 0.09);
669
+ }
670
+ .jp-meta {
671
+ padding: 5px 14px;
672
+ font-family: var(--mono);
673
+ font-size: 11.5px;
674
+ color: var(--muted);
675
+ background: #fbfbfc;
676
+ border-top: 1px solid var(--line);
677
+ }
678
+ .jp-out {
679
+ display: flex;
680
+ border-top: 1px solid var(--line);
681
+ background: var(--panel);
682
+ }
683
+ .jp-out .jp-gutter {
684
+ color: var(--accent-strong);
685
+ }
686
+ .jp-out-body {
687
+ flex: 1;
688
+ min-width: 0;
689
+ }
690
+ #page .jp-out-pre {
691
+ min-width: 0;
692
+ margin: 0;
693
+ border: none;
694
+ border-radius: 0;
695
+ background: none;
696
+ color: var(--ink);
697
+ font-family: var(--mono);
698
+ font-size: 13px;
699
+ line-height: 1.55;
700
+ padding: 12px 16px 12px 0;
701
+ white-space: pre;
702
+ overflow-x: auto;
703
+ overflow-y: auto;
704
+ max-height: 26em;
705
+ }
706
+ .jp-artifacts {
707
+ display: flex;
708
+ flex-direction: column;
709
+ }
710
+ .jp-out-body .jp-out-pre + .jp-artifacts {
711
+ border-top: 1px solid var(--line);
712
+ }
713
+ .out-artifact {
714
+ display: flex;
715
+ align-items: baseline;
716
+ gap: 8px;
717
+ padding: 9px 16px 9px 0;
718
+ text-decoration: none;
719
+ color: inherit;
720
+ }
721
+ .out-artifact + .out-artifact {
722
+ border-top: 1px solid var(--line);
723
+ }
724
+ a.out-artifact:hover .out-artifact-name {
725
+ color: var(--accent-strong);
726
+ }
727
+ .out-artifact-ico {
728
+ flex: 0 0 auto;
729
+ font-size: 13px;
730
+ }
731
+ .out-artifact-name {
732
+ font-family: var(--mono);
733
+ font-size: 12.5px;
734
+ font-weight: 600;
735
+ color: var(--ink);
736
+ overflow: hidden;
737
+ text-overflow: ellipsis;
738
+ white-space: nowrap;
739
+ }
740
+ .out-artifact-meta {
741
+ flex: 0 0 auto;
742
+ margin-left: auto;
743
+ padding-left: 12px;
744
+ font-size: 12px;
745
+ color: var(--muted);
746
+ white-space: nowrap;
747
+ }
748
+ .out-artifact-state.open {
749
+ color: var(--accent);
750
+ font-weight: 600;
751
+ }
752
+ .trackio-embed {
753
+ border: 1px solid var(--line);
754
+ border-radius: var(--radius);
755
+ overflow: hidden;
756
+ background: var(--panel);
757
+ }
758
+ .trackio-cell-meta {
759
+ display: flex;
760
+ gap: 6px;
761
+ flex-wrap: wrap;
762
+ justify-content: flex-end;
763
+ }
764
+
765
+ /* ---- unfurl cards ---- */
766
+ .unfurl {
767
+ display: block;
768
+ border: 1px solid var(--line);
769
+ border-radius: var(--radius);
770
+ background: var(--panel);
771
+ margin: 12px 0;
772
+ overflow: hidden;
773
+ text-decoration: none;
774
+ color: inherit;
775
+ transition: border-color 0.14s, box-shadow 0.14s;
776
+ }
777
+ .unfurl:hover {
778
+ border-color: #cfcbe6;
779
+ box-shadow: 0 4px 18px rgba(30, 20, 80, 0.06);
780
+ }
781
+
782
+ .unfurl-body {
783
+ padding: 13px 16px;
784
+ display: flex;
785
+ gap: 12px;
786
+ align-items: flex-start;
787
+ }
788
+
789
+ .unfurl-ico {
790
+ font-size: 20px;
791
+ line-height: 1.3;
792
+ flex: 0 0 auto;
793
+ }
794
+
795
+ .unfurl-main {
796
+ min-width: 0;
797
+ flex: 1;
798
+ }
799
+
800
+ .unfurl-kind {
801
+ font-family: var(--mono);
802
+ font-size: 10.5px;
803
+ text-transform: uppercase;
804
+ letter-spacing: 0.08em;
805
+ color: var(--accent);
806
+ font-weight: 600;
807
+ }
808
+
809
+ .unfurl-title {
810
+ font-weight: 650;
811
+ font-size: 15px;
812
+ margin: 1px 0 2px;
813
+ white-space: nowrap;
814
+ overflow: hidden;
815
+ text-overflow: ellipsis;
816
+ }
817
+
818
+ .unfurl-desc {
819
+ color: var(--muted);
820
+ font-size: 13.5px;
821
+ line-height: 1.45;
822
+ }
823
+
824
+ .unfurl-meta {
825
+ margin-top: 6px;
826
+ display: flex;
827
+ flex-wrap: wrap;
828
+ gap: 6px;
829
+ }
830
+
831
+ .chip {
832
+ font-size: 11.5px;
833
+ background: var(--code-bg);
834
+ border-radius: 999px;
835
+ padding: 2px 9px;
836
+ color: var(--muted);
837
+ font-family: var(--mono);
838
+ }
839
+
840
+ .unfurl-raw {
841
+ font-family: var(--mono);
842
+ font-size: 11px;
843
+ color: var(--muted);
844
+ border-top: 1px solid var(--line);
845
+ padding: 7px 16px;
846
+ white-space: nowrap;
847
+ overflow: hidden;
848
+ text-overflow: ellipsis;
849
+ }
850
+
851
+ .unfurl.embed {
852
+ padding: 0;
853
+ overflow: hidden;
854
+ }
855
+ .embed-head {
856
+ display: flex;
857
+ align-items: center;
858
+ gap: 10px;
859
+ padding: 10px 14px;
860
+ border-bottom: 1px solid var(--line);
861
+ }
862
+ .embed-head .unfurl-kind {
863
+ flex: 0 0 auto;
864
+ }
865
+ .embed-title {
866
+ flex: 1;
867
+ min-width: 0;
868
+ font-weight: 650;
869
+ font-size: 14px;
870
+ color: var(--ink);
871
+ text-decoration: none;
872
+ white-space: nowrap;
873
+ overflow: hidden;
874
+ text-overflow: ellipsis;
875
+ }
876
+ .embed-title:hover {
877
+ color: var(--accent);
878
+ }
879
+ .embed-open {
880
+ flex: 0 0 auto;
881
+ font-family: var(--mono);
882
+ font-size: 12px;
883
+ color: var(--accent);
884
+ text-decoration: none;
885
+ }
886
+ .embed-frame {
887
+ display: block;
888
+ width: 100%;
889
+ height: 560px;
890
+ border: 0;
891
+ background: var(--code-bg);
892
+ }
893
+
894
+ .dashboard-shell {
895
+ display: block;
896
+ }
897
+ .dashboard-shell .dashboard-frame {
898
+ display: block;
899
+ width: 100%;
900
+ height: 900px;
901
+ border: 0;
902
+ background: var(--code-bg);
903
+ }
904
+
905
+ .unfurl.image {
906
+ padding: 0;
907
+ }
908
+ .unfurl.image img {
909
+ display: block;
910
+ width: 100%;
911
+ height: auto;
912
+ max-height: 460px;
913
+ object-fit: contain;
914
+ background: var(--code-bg);
915
+ }
916
+
917
+ .artifact-chip {
918
+ border: 1px solid var(--line);
919
+ background: var(--panel);
920
+ border-radius: var(--radius);
921
+ padding: 10px 14px;
922
+ margin: 8px 0;
923
+ font-size: 14px;
924
+ }
925
+ .cell.dashboard .artifact-chip {
926
+ margin: 14px 18px 18px;
927
+ }
928
+ .artifact-chip code {
929
+ color: var(--accent);
930
+ }
931
+
932
+ /* ---- task board ---- */
933
+ .board-wrap {
934
+ overflow-x: auto;
935
+ border: 1px solid var(--line);
936
+ border-radius: var(--radius);
937
+ margin: 12px 0 20px;
938
+ background: var(--panel);
939
+ }
940
+ table.board {
941
+ border-collapse: collapse;
942
+ width: 100%;
943
+ font-size: 14px;
944
+ }
945
+ table.board th,
946
+ table.board td {
947
+ text-align: left;
948
+ padding: 9px 14px;
949
+ border-bottom: 1px solid var(--line);
950
+ vertical-align: top;
951
+ }
952
+ table.board thead th {
953
+ background: var(--accent-soft);
954
+ font-size: 12px;
955
+ text-transform: uppercase;
956
+ letter-spacing: 0.05em;
957
+ color: #9a4a12;
958
+ font-weight: 600;
959
+ border-bottom: 1px solid var(--line);
960
+ }
961
+ table.board tbody tr:last-child td {
962
+ border-bottom: none;
963
+ }
964
+ table.board .col-check {
965
+ text-align: center;
966
+ width: 92px;
967
+ white-space: nowrap;
968
+ }
969
+ table.board tr.section-row td {
970
+ background: var(--accent-soft);
971
+ text-align: center;
972
+ font-weight: 700;
973
+ font-size: 13px;
974
+ color: var(--accent-strong);
975
+ padding: 7px 14px;
976
+ letter-spacing: 0.02em;
977
+ }
978
+ .box {
979
+ display: inline-flex;
980
+ align-items: center;
981
+ justify-content: center;
982
+ width: 18px;
983
+ height: 18px;
984
+ border: 1.5px solid #cfcbe0;
985
+ border-radius: 5px;
986
+ font-size: 12px;
987
+ color: #fff;
988
+ line-height: 1;
989
+ }
990
+ .box.on {
991
+ background: var(--accent);
992
+ border-color: var(--accent);
993
+ }
994
+ .who-chip {
995
+ display: inline-block;
996
+ padding: 3px 12px;
997
+ border-radius: 999px;
998
+ font-size: 12.5px;
999
+ font-weight: 600;
1000
+ white-space: nowrap;
1001
+ }
1002
+ .who-chip.muted {
1003
+ background: var(--code-bg);
1004
+ color: var(--muted);
1005
+ font-weight: 500;
1006
+ }
1007
+
1008
+ /* ---- status badges + clickable rows ---- */
1009
+ table.board .col-status {
1010
+ width: 130px;
1011
+ white-space: nowrap;
1012
+ }
1013
+ .badge {
1014
+ display: inline-block;
1015
+ padding: 3px 11px;
1016
+ border-radius: 999px;
1017
+ font-size: 12px;
1018
+ font-weight: 600;
1019
+ letter-spacing: 0.01em;
1020
+ }
1021
+ .badge.gray {
1022
+ background: var(--code-bg);
1023
+ color: var(--muted);
1024
+ }
1025
+ .badge.amber {
1026
+ background: var(--accent-soft);
1027
+ color: #b45309;
1028
+ }
1029
+ .badge.green {
1030
+ background: #e6f7ee;
1031
+ color: #1a8a55;
1032
+ }
1033
+ .badge.red {
1034
+ background: #fde8ec;
1035
+ color: #c62a4b;
1036
+ }
1037
+ table.board tr.linked-row {
1038
+ cursor: pointer;
1039
+ }
1040
+ table.board tr.linked-row:hover td {
1041
+ background: var(--accent-soft);
1042
+ }
1043
+ table.board tr.linked-row a {
1044
+ color: var(--ink);
1045
+ font-weight: 600;
1046
+ text-decoration: none;
1047
+ }
1048
+ table.board tr.linked-row:hover a {
1049
+ color: var(--accent-strong);
1050
+ }
1051
+
1052
+ /* ---- agent read hint ---- */
1053
+ .agent-hint {
1054
+ display: flex;
1055
+ align-items: center;
1056
+ flex-wrap: wrap;
1057
+ gap: 8px;
1058
+ margin: 4px 0 22px;
1059
+ font-size: 12.5px;
1060
+ color: var(--muted);
1061
+ }
1062
+ #page .agent-hint code {
1063
+ background: var(--code-bg);
1064
+ padding: 2px 9px;
1065
+ border-radius: 6px;
1066
+ font-family: var(--mono);
1067
+ font-size: 12px;
1068
+ font-weight: 500;
1069
+ color: var(--ink);
1070
+ }
1071
+ .agent-hint .copy {
1072
+ flex: 0 0 auto;
1073
+ background: none;
1074
+ color: var(--muted);
1075
+ border: 1px solid var(--line);
1076
+ border-radius: 6px;
1077
+ width: 22px;
1078
+ height: 22px;
1079
+ font-size: 11px;
1080
+ line-height: 1;
1081
+ cursor: pointer;
1082
+ transition: color 0.12s, border-color 0.12s;
1083
+ }
1084
+ .agent-hint .copy:hover {
1085
+ color: var(--accent-strong);
1086
+ border-color: var(--accent);
1087
+ }
1088
+ .agent-hint .copy.copied {
1089
+ color: #1a8a55;
1090
+ border-color: #1a8a55;
1091
+ }
1092
+ .agent-hint-note {
1093
+ margin-left: auto;
1094
+ font-size: 12px;
1095
+ color: var(--muted);
1096
+ }
1097
+
1098
+ /* ---- logbook summary stats ---- */
1099
+ .logbook-stats {
1100
+ display: flex;
1101
+ flex-wrap: wrap;
1102
+ gap: 12px;
1103
+ margin: 0 0 28px;
1104
+ }
1105
+ .stat-tile {
1106
+ position: relative;
1107
+ display: inline-flex;
1108
+ align-items: center;
1109
+ gap: 11px;
1110
+ border: 1px solid var(--line);
1111
+ background: var(--panel);
1112
+ border-radius: var(--radius);
1113
+ padding: 12px 23px;
1114
+ font: inherit;
1115
+ text-align: left;
1116
+ cursor: pointer;
1117
+ transition: border-color 0.12s, box-shadow 0.12s;
1118
+ }
1119
+ .stat-tile:hover:not([disabled]) {
1120
+ border-color: rgba(249, 115, 22, 0.45);
1121
+ box-shadow: 0 3px 12px rgba(31, 41, 55, 0.06);
1122
+ }
1123
+ .stat-tile:focus-visible {
1124
+ outline: 2px solid var(--accent);
1125
+ outline-offset: 2px;
1126
+ }
1127
+ .stat-tile[disabled] {
1128
+ cursor: default;
1129
+ opacity: 0.7;
1130
+ }
1131
+ .stat-tile.open {
1132
+ border-color: rgba(249, 115, 22, 0.6);
1133
+ box-shadow: 0 3px 12px rgba(31, 41, 55, 0.08);
1134
+ }
1135
+ .stat-icon {
1136
+ width: 24px;
1137
+ height: 24px;
1138
+ flex: 0 0 24px;
1139
+ object-fit: contain;
1140
+ align-self: center;
1141
+ }
1142
+ .stat-text {
1143
+ display: flex;
1144
+ align-items: baseline;
1145
+ gap: 8px;
1146
+ white-space: nowrap;
1147
+ line-height: 1;
1148
+ }
1149
+ .stat-num {
1150
+ font-family: var(--mono);
1151
+ font-size: 20px;
1152
+ font-weight: 600;
1153
+ line-height: 1;
1154
+ color: var(--accent-strong);
1155
+ }
1156
+ .stat-label {
1157
+ font-size: 15px;
1158
+ line-height: 1;
1159
+ color: var(--muted);
1160
+ }
1161
+ .stat-caret {
1162
+ margin-left: 2px;
1163
+ font-size: 10px;
1164
+ color: var(--muted);
1165
+ align-self: center;
1166
+ transition: transform 0.12s;
1167
+ }
1168
+ .stat-tile.open .stat-caret {
1169
+ transform: rotate(180deg);
1170
+ }
1171
+ .stat-popover {
1172
+ position: absolute;
1173
+ top: 100%;
1174
+ left: 0;
1175
+ margin-top: 6px;
1176
+ min-width: 300px;
1177
+ max-width: min(460px, 92vw);
1178
+ max-height: 340px;
1179
+ overflow-y: auto;
1180
+ z-index: 20;
1181
+ background: var(--panel);
1182
+ border: 1px solid var(--line);
1183
+ border-radius: var(--radius);
1184
+ box-shadow: 0 8px 28px rgba(31, 41, 55, 0.12);
1185
+ padding: 6px;
1186
+ }
1187
+ .stat-popover[hidden] {
1188
+ display: none;
1189
+ }
1190
+ .stat-pop-head {
1191
+ padding: 6px 10px 8px;
1192
+ font-size: 11.5px;
1193
+ font-weight: 700;
1194
+ letter-spacing: 0.03em;
1195
+ text-transform: uppercase;
1196
+ color: var(--muted);
1197
+ }
1198
+ .stat-row {
1199
+ display: flex;
1200
+ align-items: flex-start;
1201
+ gap: 10px;
1202
+ padding: 9px 11px;
1203
+ border-radius: 9px;
1204
+ border: 1px solid transparent;
1205
+ text-decoration: none;
1206
+ color: inherit;
1207
+ cursor: pointer;
1208
+ }
1209
+ .stat-row:hover {
1210
+ border-color: rgba(249, 115, 22, 0.4);
1211
+ background: var(--accent-soft);
1212
+ }
1213
+ .stat-row-ico {
1214
+ font-size: 15px;
1215
+ line-height: 1.3;
1216
+ flex: 0 0 auto;
1217
+ }
1218
+ .stat-row-main {
1219
+ min-width: 0;
1220
+ flex: 1;
1221
+ }
1222
+ .stat-row-title {
1223
+ font-family: var(--mono);
1224
+ font-size: 12.5px;
1225
+ font-weight: 600;
1226
+ color: var(--ink);
1227
+ overflow: hidden;
1228
+ text-overflow: ellipsis;
1229
+ white-space: nowrap;
1230
+ }
1231
+ .stat-row-meta {
1232
+ margin-top: 2px;
1233
+ font-size: 12px;
1234
+ color: var(--muted);
1235
+ }
1236
+ .stat-row-state.open {
1237
+ color: var(--accent);
1238
+ font-weight: 600;
1239
+ border-radius: 5px;
1240
+ padding: 1px 5px;
1241
+ margin: -1px -2px;
1242
+ }
1243
+ .stat-row-state.open:hover {
1244
+ background: rgba(249, 115, 22, 0.14);
1245
+ text-decoration: underline;
1246
+ }
1247
+ .art-ico {
1248
+ width: 1em;
1249
+ height: 1em;
1250
+ object-fit: contain;
1251
+ vertical-align: -0.15em;
1252
+ }
1253
+
1254
+ /* ---- scroll-to-resource highlight ---- */
1255
+ .res-flash {
1256
+ animation: res-flash 1.5s ease;
1257
+ border-radius: 8px;
1258
+ }
1259
+ @keyframes res-flash {
1260
+ 0%,
1261
+ 25% {
1262
+ box-shadow: 0 0 0 3px var(--accent);
1263
+ }
1264
+ 100% {
1265
+ box-shadow: 0 0 0 3px rgba(249, 115, 22, 0);
1266
+ }
1267
+ }
1268
+
1269
+ /* ---- inline resource chips ---- */
1270
+ #page .res-chip {
1271
+ display: inline-flex;
1272
+ align-items: center;
1273
+ gap: 5px;
1274
+ max-width: 100%;
1275
+ padding: 0 9px 0 6px;
1276
+ margin: 0 1px;
1277
+ border: 1px solid var(--line);
1278
+ border-radius: 999px;
1279
+ background: var(--panel);
1280
+ font-family: var(--mono);
1281
+ font-size: 0.78em;
1282
+ font-weight: 600;
1283
+ color: var(--ink);
1284
+ text-decoration: none;
1285
+ white-space: nowrap;
1286
+ overflow: hidden;
1287
+ text-overflow: ellipsis;
1288
+ vertical-align: middle;
1289
+ line-height: 1.65;
1290
+ transform: translateY(-0.08em);
1291
+ transition: border-color 0.12s, background 0.12s, color 0.12s;
1292
+ }
1293
+ .res-chip-ico {
1294
+ font-size: 1.05em;
1295
+ line-height: 1;
1296
+ }
1297
+ #page .res-chip:hover,
1298
+ #page .res-chip.res-hl {
1299
+ border-color: var(--accent);
1300
+ background: var(--accent-soft);
1301
+ color: var(--accent-strong);
1302
+ }
1303
+ #page a.res-link.res-hl {
1304
+ background: var(--accent-soft);
1305
+ border-radius: 4px;
1306
+ }
1307
+ .rail-item.res-hl {
1308
+ border-color: var(--accent);
1309
+ background: var(--accent-soft);
1310
+ box-shadow: 0 3px 12px rgba(249, 115, 22, 0.14);
1311
+ }
1312
+ .rail-item.res-hl .rail-title {
1313
+ color: var(--accent-strong);
1314
+ }
1315
+ .rail-item.rail-local {
1316
+ cursor: default;
1317
+ }
1318
+ .artifact-chip.res-hl {
1319
+ border-color: var(--accent);
1320
+ background: var(--accent-soft);
1321
+ }
1322
+
1323
+ /* ---- contextual resources rail ---- */
1324
+ .context-rail {
1325
+ position: relative;
1326
+ width: 248px;
1327
+ }
1328
+ .context-rail[hidden] {
1329
+ display: none;
1330
+ }
1331
+ .rail-kind {
1332
+ display: flex;
1333
+ align-items: center;
1334
+ gap: 5px;
1335
+ font-family: var(--mono);
1336
+ font-size: 10px;
1337
+ text-transform: uppercase;
1338
+ letter-spacing: 0.08em;
1339
+ font-weight: 600;
1340
+ color: var(--accent);
1341
+ margin-bottom: 4px;
1342
+ }
1343
+ .rail-item {
1344
+ position: absolute;
1345
+ left: 0;
1346
+ right: 0;
1347
+ display: block;
1348
+ border: 1px solid var(--line);
1349
+ border-radius: 10px;
1350
+ background: var(--panel);
1351
+ padding: 9px 12px;
1352
+ margin-bottom: 8px;
1353
+ text-decoration: none;
1354
+ color: inherit;
1355
+ transition: border-color 0.14s, box-shadow 0.14s;
1356
+ }
1357
+ .rail-item:hover {
1358
+ border-color: rgba(249, 115, 22, 0.45);
1359
+ box-shadow: 0 3px 12px rgba(31, 41, 55, 0.06);
1360
+ }
1361
+ .rail-title {
1362
+ font-family: var(--mono);
1363
+ font-size: 12.5px;
1364
+ font-weight: 600;
1365
+ color: var(--ink);
1366
+ overflow-wrap: anywhere;
1367
+ line-height: 1.4;
1368
+ }
1369
+ .rail-item:hover .rail-title {
1370
+ color: var(--accent-strong);
1371
+ }
1372
+ .rail-meta {
1373
+ font-size: 11.5px;
1374
+ color: var(--muted);
1375
+ margin-top: 2px;
1376
+ }
1377
+
1378
+ @media (max-width: 1400px) {
1379
+ .page-layout {
1380
+ display: block;
1381
+ }
1382
+ .context-rail {
1383
+ width: 100%;
1384
+ margin-top: 28px;
1385
+ position: static;
1386
+ min-height: 0 !important;
1387
+ display: grid;
1388
+ grid-template-columns: repeat(auto-fit, minmax(220px, 1fr));
1389
+ gap: 10px;
1390
+ }
1391
+ .context-rail[hidden] {
1392
+ display: none;
1393
+ }
1394
+ .context-rail .rail-item {
1395
+ position: static;
1396
+ margin-bottom: 0;
1397
+ }
1398
+ }
1399
+
1400
+ /* ---- connect footer + modal ---- */
1401
+ #sidebar-foot {
1402
+ margin-top: auto;
1403
+ padding-top: 14px;
1404
+ border-top: 1px solid rgba(255, 255, 255, 0.1);
1405
+ }
1406
+
1407
+ #connect-btn {
1408
+ width: 100%;
1409
+ display: flex;
1410
+ align-items: center;
1411
+ gap: 8px;
1412
+ background: rgba(255, 255, 255, 0.05);
1413
+ color: #c3c4cb;
1414
+ border: 1px solid rgba(255, 255, 255, 0.12);
1415
+ border-radius: 9px;
1416
+ padding: 9px 12px;
1417
+ font-size: 13.5px;
1418
+ font-family: var(--sans);
1419
+ cursor: pointer;
1420
+ transition: background 0.12s, color 0.12s, border-color 0.12s;
1421
+ }
1422
+ #connect-btn:hover {
1423
+ background: rgba(249, 115, 22, 0.14);
1424
+ border-color: rgba(249, 115, 22, 0.4);
1425
+ color: #fdba74;
1426
+ }
1427
+ #connect-btn .ico {
1428
+ font-size: 15px;
1429
+ }
1430
+
1431
+ #modal[hidden] {
1432
+ display: none;
1433
+ }
1434
+ #modal {
1435
+ position: fixed;
1436
+ inset: 0;
1437
+ z-index: 100;
1438
+ display: flex;
1439
+ align-items: center;
1440
+ justify-content: center;
1441
+ padding: 24px;
1442
+ }
1443
+ .modal-backdrop {
1444
+ position: absolute;
1445
+ inset: 0;
1446
+ background: rgba(20, 18, 30, 0.5);
1447
+ backdrop-filter: blur(2px);
1448
+ }
1449
+ .modal-card {
1450
+ position: relative;
1451
+ background: var(--panel);
1452
+ border-radius: 16px;
1453
+ width: 100%;
1454
+ max-width: 620px;
1455
+ max-height: 85vh;
1456
+ overflow-y: auto;
1457
+ box-shadow: 0 24px 70px rgba(20, 15, 50, 0.28);
1458
+ }
1459
+ .modal-head {
1460
+ display: flex;
1461
+ align-items: center;
1462
+ justify-content: space-between;
1463
+ gap: 12px;
1464
+ padding: 18px 22px;
1465
+ border-bottom: 1px solid var(--line);
1466
+ position: sticky;
1467
+ top: 0;
1468
+ background: var(--panel);
1469
+ }
1470
+ .modal-title {
1471
+ display: flex;
1472
+ align-items: center;
1473
+ gap: 10px;
1474
+ font-family: var(--serif);
1475
+ font-size: 21px;
1476
+ letter-spacing: -0.01em;
1477
+ }
1478
+ .modal-logo {
1479
+ width: 26px;
1480
+ height: 26px;
1481
+ object-fit: contain;
1482
+ }
1483
+ .modal-actions {
1484
+ display: flex;
1485
+ align-items: center;
1486
+ gap: 8px;
1487
+ }
1488
+ .btn {
1489
+ font-family: var(--sans);
1490
+ font-size: 13.5px;
1491
+ font-weight: 600;
1492
+ border: 1px solid var(--line);
1493
+ background: var(--panel);
1494
+ color: var(--ink);
1495
+ border-radius: 9px;
1496
+ padding: 8px 13px;
1497
+ cursor: pointer;
1498
+ transition: background 0.12s, border-color 0.12s, color 0.12s;
1499
+ }
1500
+ .btn:hover {
1501
+ border-color: var(--accent);
1502
+ color: var(--accent-strong);
1503
+ }
1504
+ .btn.copied {
1505
+ border-color: #1a8a55;
1506
+ color: #1a8a55;
1507
+ }
1508
+ .btn.icon {
1509
+ font-size: 18px;
1510
+ line-height: 1;
1511
+ padding: 6px 11px;
1512
+ font-weight: 400;
1513
+ }
1514
+ .modal-body {
1515
+ padding: 20px 22px 26px;
1516
+ }
1517
+ .modal-intro {
1518
+ margin: 0 0 20px;
1519
+ color: var(--muted);
1520
+ line-height: 1.55;
1521
+ }
1522
+ #connect-steps {
1523
+ list-style: none;
1524
+ margin: 0;
1525
+ padding: 0;
1526
+ }
1527
+ #connect-steps li {
1528
+ margin-bottom: 18px;
1529
+ }
1530
+ .step-title {
1531
+ font-weight: 600;
1532
+ font-size: 14.5px;
1533
+ margin-bottom: 8px;
1534
+ }
1535
+ .codeblock {
1536
+ display: flex;
1537
+ align-items: center;
1538
+ gap: 8px;
1539
+ background: #17181c;
1540
+ border-radius: 10px;
1541
+ padding: 11px 12px 11px 15px;
1542
+ }
1543
+ .codeblock code {
1544
+ flex: 1;
1545
+ min-width: 0;
1546
+ overflow-x: auto;
1547
+ white-space: nowrap;
1548
+ font-family: var(--mono);
1549
+ font-size: 13px;
1550
+ color: #f0efff;
1551
+ background: none;
1552
+ padding: 0;
1553
+ }
1554
+ .codeblock .copy {
1555
+ flex: 0 0 auto;
1556
+ background: rgba(255, 255, 255, 0.08);
1557
+ color: #c3c4cb;
1558
+ border: 1px solid rgba(255, 255, 255, 0.14);
1559
+ border-radius: 7px;
1560
+ width: 30px;
1561
+ height: 30px;
1562
+ font-size: 14px;
1563
+ cursor: pointer;
1564
+ transition: background 0.12s, color 0.12s;
1565
+ }
1566
+ .codeblock .copy:hover {
1567
+ background: rgba(249, 115, 22, 0.2);
1568
+ color: #fdba74;
1569
+ }
1570
+ .codeblock .copy.copied {
1571
+ color: #52d08a;
1572
+ }
1573
+
1574
+ @media (max-width: 720px) {
1575
+ #app {
1576
+ flex-direction: column;
1577
+ }
1578
+ #sidebar {
1579
+ width: 100%;
1580
+ flex: none;
1581
+ height: auto;
1582
+ position: static;
1583
+ }
1584
+ #content {
1585
+ display: block;
1586
+ width: 100%;
1587
+ padding: 28px 20px 80px;
1588
+ overflow-x: hidden;
1589
+ }
1590
+ #page {
1591
+ width: 100%;
1592
+ max-width: 100%;
1593
+ }
1594
+ #page h1 {
1595
+ font-size: 30px;
1596
+ }
1597
+ .cell-head {
1598
+ align-items: flex-start;
1599
+ flex-direction: column;
1600
+ gap: 4px;
1601
+ }
1602
+ }
logbook.js ADDED
@@ -0,0 +1,2275 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ (function () {
2
+ "use strict";
3
+
4
+ let MANIFEST = null;
5
+ const PAGE_CACHE = {};
6
+ const UNFURL_CACHE = {};
7
+ const LIVE_RELOAD_MS = 1500;
8
+ const FIGURE_FRAME_WINDOWS = new Set();
9
+ let FIGURE_NAVIGATION_READY = false;
10
+
11
+ function esc(s) {
12
+ return String(s)
13
+ .replace(/&/g, "&amp;")
14
+ .replace(/</g, "&lt;")
15
+ .replace(/>/g, "&gt;")
16
+ .replace(/"/g, "&quot;")
17
+ .replace(/'/g, "&#39;");
18
+ }
19
+
20
+ function flattenTree(node, depth, acc) {
21
+ acc.push({ node: node, depth: depth });
22
+ (node.children || []).forEach((c) => flattenTree(c, depth + 1, acc));
23
+ return acc;
24
+ }
25
+
26
+ function findNode(node, slug) {
27
+ if (node.slug === slug) return node;
28
+ for (const c of node.children || []) {
29
+ const hit = findNode(c, slug);
30
+ if (hit) return hit;
31
+ }
32
+ return null;
33
+ }
34
+
35
+ /* -------------------- minimal markdown -------------------- */
36
+
37
+ function inline(text) {
38
+ let t = esc(text);
39
+ t = t.replace(/`([^`]+)`/g, (_, c) => `<code>${c}</code>`);
40
+ t = t.replace(/\*\*([^*]+)\*\*/g, (_, c) => `<strong>${c}</strong>`);
41
+ t = t.replace(/\[([^\]]+)\]\(([^)]+)\)/g, (_, txt, url) => {
42
+ const safe = esc(url);
43
+ const attrs = /^https?:/.test(url) ? ' target="_blank" rel="noopener"' : "";
44
+ const item = /^https?:/.test(url) ? classifyResource(url) : null;
45
+ const data = item
46
+ ? ` class="res-link" data-res-url="${esc(item.url)}"`
47
+ : "";
48
+ return `<a href="${safe}"${attrs}${data}>${txt}</a>`;
49
+ });
50
+ t = t.replace(/(^|[\s(])(https?:\/\/[^\s<>)"'`]+)/g, (m, pre, url) => {
51
+ let rest = "";
52
+ const cut = url.search(/&quot;|&#39;|&lt;|&gt;/);
53
+ if (cut !== -1) {
54
+ rest = url.slice(cut);
55
+ url = url.slice(0, cut);
56
+ }
57
+ const trailing = (url.match(/[.,;:!?`]+$/) || [""])[0];
58
+ const clean = trailing ? url.slice(0, -trailing.length) : url;
59
+ if (!clean) return m;
60
+ const item = classifyResource(clean);
61
+ if (item) return `${pre}${resChipHtml(item)}${trailing}${rest}`;
62
+ return `${pre}<a href="${clean}" target="_blank" rel="noopener">${clean}</a>${trailing}${rest}`;
63
+ });
64
+ return t;
65
+ }
66
+
67
+ function resChipHtml(item) {
68
+ return (
69
+ `<a class="res-chip" href="${esc(item.url)}" target="_blank" ` +
70
+ `rel="noopener" data-res-url="${esc(item.url)}">` +
71
+ `<span class="res-chip-ico">${RESOURCE_ICONS[item.kind]}</span>` +
72
+ `${esc(item.id)}</a>`
73
+ );
74
+ }
75
+
76
+ const URL_ONLY = /^(https?:\/\/[^\s]+)$/;
77
+ const DETECTED_URL =
78
+ /(https?:\/\/[^\s<>)\]"'`]+|trackio-local-dashboard:\/\/[^\s<>)\]"'`]+|trackio-artifact:\/\/[^\s<>)\]"'`]+|trackio-local-path:\/\/[^\s<>)\]"'`]+)/g;
79
+
80
+ function renderMarkdown(md, container) {
81
+ const cellRe = /(^|\n)---\n<!-- trackio-cell\n([\s\S]*?)\n-->\n([\s\S]*?)(?=\n---\n<!-- trackio-cell\n|\s*$)/g;
82
+ const tokens = [];
83
+ let pos = 0;
84
+ let found = false;
85
+ let match;
86
+ while ((match = cellRe.exec(md))) {
87
+ found = true;
88
+ tokens.push({
89
+ kind: "md",
90
+ text: md.slice(pos, match.index + match[1].length),
91
+ });
92
+ tokens.push({
93
+ kind: "cell",
94
+ meta: parseCellMeta(match[2]),
95
+ body: match[3],
96
+ });
97
+ pos = match.index + match[0].length;
98
+ }
99
+ tokens.push({ kind: "md", text: found ? md.slice(pos) : md });
100
+
101
+ for (let i = 0; i < tokens.length; i++) {
102
+ const t = tokens[i];
103
+ if (t.kind === "md") {
104
+ renderMarkdownPlain(t.text, container);
105
+ continue;
106
+ }
107
+ if (t.consumed) continue;
108
+ if (t.meta.type === "code") {
109
+ const arts = [];
110
+ for (let j = i + 1; j < tokens.length; j++) {
111
+ const n = tokens[j];
112
+ if (n.kind === "md") {
113
+ if (n.text.trim() === "") continue;
114
+ break;
115
+ }
116
+ if (n.meta.type === "artifact") {
117
+ arts.push(n);
118
+ n.consumed = true;
119
+ continue;
120
+ }
121
+ break;
122
+ }
123
+ renderCell(t.meta, t.body, container, arts);
124
+ } else {
125
+ renderCell(t.meta, t.body, container);
126
+ }
127
+ }
128
+ }
129
+
130
+ function parseCellMeta(raw) {
131
+ try {
132
+ return JSON.parse(raw);
133
+ } catch (e) {
134
+ return { type: "markdown", title: "Note" };
135
+ }
136
+ }
137
+
138
+ function renderMarkdownPlain(md, container) {
139
+ const lines = md.replace(/<!--[\s\S]*?-->/g, "").split("\n");
140
+ let i = 0;
141
+ let para = [];
142
+
143
+ function flushPara() {
144
+ if (!para.length) return;
145
+ const joined = para.join(" ").trim();
146
+ para = [];
147
+ if (!joined) return;
148
+ if (/^trackio-artifact:\/\/\S+$/.test(joined)) return;
149
+ if (/^trackio-local-path:\/\/\S+$/.test(joined)) return;
150
+ if (joined.indexOf("πŸ“¦ Artifact") !== -1) {
151
+ const div = document.createElement("div");
152
+ div.className = "artifact-chip";
153
+ div.innerHTML = ARTIFACT_ICON_IMG + inline(joined.replace(/πŸ“¦\s*/, ""));
154
+ container.appendChild(div);
155
+ return;
156
+ }
157
+ if (URL_ONLY.test(joined) || IMG_PATH.test(joined)) {
158
+ const el = renderStandaloneUrl(joined);
159
+ if (el) container.appendChild(el);
160
+ return;
161
+ }
162
+ const p = document.createElement("p");
163
+ p.innerHTML = inline(joined);
164
+ container.appendChild(p);
165
+ }
166
+
167
+ while (i < lines.length) {
168
+ const line = lines[i];
169
+ const trimmed = line.trim();
170
+
171
+ if (trimmed === "") {
172
+ flushPara();
173
+ i++;
174
+ continue;
175
+ }
176
+ const fence = trimmed.match(/^(`{3,}|~{3,})(.*)$/);
177
+ if (fence) {
178
+ flushPara();
179
+ const marker = fence[1][0];
180
+ const closeRe = new RegExp("^" + marker + "{" + fence[1].length + ",}\\s*$");
181
+ const info = fence[2].trim();
182
+ const buf = [];
183
+ i++;
184
+ while (i < lines.length && !closeRe.test(lines[i].trim())) {
185
+ buf.push(lines[i]);
186
+ i++;
187
+ }
188
+ i++;
189
+ const lang = (info.split(/\s+/)[0] || "").toLowerCase();
190
+ const tm = info.match(/title=(\S+)/);
191
+ container.appendChild(
192
+ renderCode(buf.join("\n"), lang, tm ? tm[1] : null)
193
+ );
194
+ continue;
195
+ }
196
+ if (trimmed === "---") {
197
+ flushPara();
198
+ container.appendChild(document.createElement("hr"));
199
+ i++;
200
+ continue;
201
+ }
202
+ const h = trimmed.match(/^(#{1,4})\s+(.*)$/);
203
+ if (h) {
204
+ flushPara();
205
+ const el = document.createElement("h" + h[1].length);
206
+ el.innerHTML = inline(h[2]);
207
+ container.appendChild(el);
208
+ i++;
209
+ continue;
210
+ }
211
+ if (
212
+ trimmed.startsWith("|") &&
213
+ i + 1 < lines.length &&
214
+ /^\|?[\s:|-]*-{2,}[\s:|-]*\|?$/.test(lines[i + 1].trim())
215
+ ) {
216
+ flushPara();
217
+ const rows = [];
218
+ while (i < lines.length && lines[i].trim().startsWith("|")) {
219
+ rows.push(parseRow(lines[i].trim()));
220
+ i++;
221
+ }
222
+ renderTable(rows, container);
223
+ continue;
224
+ }
225
+ if (trimmed.startsWith("> ")) {
226
+ flushPara();
227
+ const bq = document.createElement("blockquote");
228
+ bq.innerHTML = inline(trimmed.slice(2));
229
+ container.appendChild(bq);
230
+ i++;
231
+ continue;
232
+ }
233
+ if (/^`[^`]+`$/.test(trimmed)) {
234
+ flushPara();
235
+ const el = document.createElement("div");
236
+ el.className = "ts";
237
+ el.textContent = trimmed.replace(/`/g, "");
238
+ container.appendChild(el);
239
+ i++;
240
+ continue;
241
+ }
242
+ if (trimmed.startsWith("- ")) {
243
+ flushPara();
244
+ const items = [];
245
+ while (i < lines.length && lines[i].trim().startsWith("- ")) {
246
+ items.push(lines[i].trim().slice(2).trim());
247
+ i++;
248
+ }
249
+ renderList(items, container);
250
+ continue;
251
+ }
252
+ para.push(trimmed);
253
+ i++;
254
+ }
255
+ flushPara();
256
+ }
257
+
258
+ function renderCell(meta, body, container, artifacts) {
259
+ const cell = document.createElement("section");
260
+ cell.className = `cell ${meta.type || "markdown"}`;
261
+ if (meta.id) cell.dataset.cellId = meta.id;
262
+ if (isPinned(meta)) cell.classList.add("pinned-source");
263
+
264
+ const head = document.createElement("div");
265
+ head.className = "cell-head";
266
+ const rawTitle = (meta.title || "").trim();
267
+ const title = rawTitle && rawTitle.toLowerCase() !== "untitled" ? esc(rawTitle) : "";
268
+ const when = meta.created_at ? `<span>${esc(formatTime(meta.created_at))}</span>` : "";
269
+ head.innerHTML =
270
+ (title ? `<div class="cell-title">${title}</div>` : "") +
271
+ `<div class="cell-meta">${when}</div>`;
272
+ if (!title) head.classList.add("no-title");
273
+ cell.appendChild(head);
274
+
275
+ const bodyEl = document.createElement("div");
276
+ bodyEl.className = "cell-body";
277
+ if (meta.type === "code") {
278
+ renderCodeCell(body, bodyEl, artifacts);
279
+ } else if (meta.type === "figure") {
280
+ cell.dataset.resUrl = `trackio-figure://${(meta.title || "Figure").trim()}`;
281
+ renderFigureCell(body, bodyEl, head);
282
+ } else if (meta.type === "artifact") {
283
+ renderMarkdownPlain(body, bodyEl);
284
+ const chip = bodyEl.querySelector(".artifact-chip");
285
+ const uri = body.match(
286
+ /(trackio-artifact:\/\/\S+|trackio-local-path:\/\/\S+|https:\/\/huggingface\.co\/buckets\/[^\s<)]+#\S+)/
287
+ );
288
+ if (chip && uri) chip.dataset.resUrl = uri[1];
289
+ } else if (meta.type === "dashboard") {
290
+ const sp = body.match(/https:\/\/huggingface\.co\/spaces\/[^\s<>)"'`]+/);
291
+ cell.dataset.resUrl = sp
292
+ ? sp[0]
293
+ : `trackio-local-dashboard://${(meta.dashboard_project || "").trim()}`;
294
+ renderDashboardCell(meta, body, bodyEl, head);
295
+ } else {
296
+ const cleaned = stripDuplicateTitle(body, meta.title);
297
+ renderMarkdownPlain(cleaned, bodyEl);
298
+ renderDetectedEmbeds(cleaned, bodyEl);
299
+ }
300
+ cell.appendChild(bodyEl);
301
+ container.appendChild(cell);
302
+ return cell;
303
+ }
304
+
305
+ function isPinned(meta) {
306
+ return Boolean(meta && (meta.pinned === true || meta.pinned === "true"));
307
+ }
308
+
309
+ function stripDuplicateTitle(body, title) {
310
+ if (!title) return body;
311
+ const m = body.match(/^\s*#{1,6}\s+([^\n]+)\n?/);
312
+ if (!m) return body;
313
+ const norm = (s) =>
314
+ s
315
+ .toLowerCase()
316
+ .replace(/[*_`#]/g, "")
317
+ .replace(/\s+/g, " ")
318
+ .trim();
319
+ return norm(m[1]) === norm(title) ? body.slice(m[0].length) : body;
320
+ }
321
+
322
+ function formatTime(iso) {
323
+ const d = new Date(iso);
324
+ if (Number.isNaN(d.getTime())) return iso;
325
+ return d.toLocaleString(undefined, {
326
+ month: "short",
327
+ day: "numeric",
328
+ hour: "2-digit",
329
+ minute: "2-digit",
330
+ });
331
+ }
332
+
333
+ function parseFences(text) {
334
+ const fenceRe = /(`{3,4}|~{3,4})([^\n]*)\n([\s\S]*?)\n\1/g;
335
+ const parts = [];
336
+ let pos = 0;
337
+ let match;
338
+ while ((match = fenceRe.exec(text))) {
339
+ if (match.index > pos) {
340
+ parts.push({ kind: "text", text: text.slice(pos, match.index) });
341
+ }
342
+ const info = match[2].trim();
343
+ const lang = (info.split(/\s+/)[0] || "").toLowerCase();
344
+ const titleMatch = info.match(/title=(\S+)/);
345
+ parts.push({
346
+ kind: lang === "result" || lang === "output" ? "output" : "code",
347
+ lang,
348
+ title: titleMatch ? titleMatch[1] : null,
349
+ text: match[3],
350
+ });
351
+ pos = match.index + match[0].length;
352
+ }
353
+ if (pos < text.length) parts.push({ kind: "text", text: text.slice(pos) });
354
+ return parts;
355
+ }
356
+
357
+ function fitFigureFrame(frame, wrap) {
358
+ let doc;
359
+ try {
360
+ doc = frame.contentDocument;
361
+ } catch (e) {
362
+ return;
363
+ }
364
+ if (!doc || !doc.body) return;
365
+ frame.style.transform = "none";
366
+ frame.style.width = "100%";
367
+ frame.style.height = "auto";
368
+ frame.style.position = "";
369
+ frame.style.left = "";
370
+ frame.style.top = "";
371
+ const avail = wrap.clientWidth;
372
+ const isFullscreen =
373
+ document.fullscreenElement === wrap ||
374
+ document.webkitFullscreenElement === wrap;
375
+ const availHeight = isFullscreen ? wrap.clientHeight : Infinity;
376
+ const cw = Math.max(doc.body.scrollWidth, doc.documentElement.scrollWidth, 1);
377
+ const ch = Math.max(doc.body.scrollHeight, doc.documentElement.scrollHeight, 1);
378
+ const scale = Math.min(avail / cw, availHeight / ch);
379
+ if (avail && scale < 1 - 1e-3) {
380
+ frame.style.width = `${cw}px`;
381
+ frame.style.height = `${ch}px`;
382
+ frame.style.transformOrigin = "top left";
383
+ frame.style.transform = `scale(${scale})`;
384
+ if (isFullscreen) {
385
+ frame.style.position = "absolute";
386
+ frame.style.left = `${Math.max(0, (avail - cw * scale) / 2)}px`;
387
+ frame.style.top = `${Math.max(0, (availHeight - ch * scale) / 2)}px`;
388
+ wrap.style.height = "100%";
389
+ } else {
390
+ wrap.style.height = `${Math.ceil(ch * scale)}px`;
391
+ }
392
+ } else {
393
+ frame.style.width = "100%";
394
+ frame.style.height = `${ch}px`;
395
+ wrap.style.height = isFullscreen ? "100%" : `${ch}px`;
396
+ }
397
+ }
398
+
399
+ function attachFigureFit(frame, wrap) {
400
+ const refit = () => fitFigureFrame(frame, wrap);
401
+ frame.addEventListener("load", refit);
402
+ if (window.ResizeObserver) {
403
+ const ro = new ResizeObserver(() => refit());
404
+ ro.observe(wrap);
405
+ }
406
+ }
407
+
408
+ function renderFigureCell(text, container, head) {
409
+ const parts = parseFences(text);
410
+ const htmlPart = parts.find((part) => part.lang === "html");
411
+ const rawPart = parts.find((part) => part.lang === "raw");
412
+ if (!htmlPart || !htmlPart.text.trim()) {
413
+ const empty = document.createElement("p");
414
+ empty.className = "muted";
415
+ empty.textContent = "No figure HTML.";
416
+ container.appendChild(empty);
417
+ return;
418
+ }
419
+ const frame = document.createElement("iframe");
420
+ frame.className = "figure-frame";
421
+ frame.sandbox = "allow-scripts allow-same-origin";
422
+ frame.loading = "lazy";
423
+ frame.srcdoc = htmlPart.text;
424
+ registerFigureNavigation(frame);
425
+ const figWrap = document.createElement("div");
426
+ figWrap.className = "figure-fit";
427
+ figWrap.appendChild(frame);
428
+ attachFigureFit(frame, figWrap);
429
+ if (head) {
430
+ const metaEl = head.querySelector(".cell-meta");
431
+ if (metaEl)
432
+ metaEl.insertBefore(buildFullscreenControl(figWrap, frame), metaEl.firstChild);
433
+ }
434
+ if (!rawPart || !rawPart.text.trim()) {
435
+ container.appendChild(figWrap);
436
+ return;
437
+ }
438
+ const sw = document.createElement("div");
439
+ sw.className = "fig-switch";
440
+ const thumb = document.createElement("span");
441
+ thumb.className = "fig-switch-thumb";
442
+ const figBtn = document.createElement("button");
443
+ figBtn.type = "button";
444
+ figBtn.className = "active";
445
+ figBtn.textContent = "Figure";
446
+ const rawBtn = document.createElement("button");
447
+ rawBtn.type = "button";
448
+ rawBtn.textContent = "Raw";
449
+ sw.appendChild(thumb);
450
+ sw.appendChild(figBtn);
451
+ sw.appendChild(rawBtn);
452
+ const rawView = document.createElement("div");
453
+ rawView.className = "figure-raw";
454
+ rawView.hidden = true;
455
+ const pre = document.createElement("pre");
456
+ const code = document.createElement("code");
457
+ code.textContent = rawPart.text;
458
+ pre.appendChild(code);
459
+ rawView.appendChild(pre);
460
+ rawView.appendChild(copySnippetBtn(rawPart.text));
461
+ const select = (showRaw) => {
462
+ sw.classList.toggle("raw", showRaw);
463
+ figBtn.classList.toggle("active", !showRaw);
464
+ rawBtn.classList.toggle("active", showRaw);
465
+ figWrap.hidden = showRaw;
466
+ rawView.hidden = !showRaw;
467
+ };
468
+ figBtn.addEventListener("click", () => select(false));
469
+ rawBtn.addEventListener("click", () => select(true));
470
+ if (head) {
471
+ head.insertBefore(sw, head.querySelector(".cell-meta"));
472
+ } else {
473
+ container.appendChild(sw);
474
+ }
475
+ container.appendChild(figWrap);
476
+ container.appendChild(rawView);
477
+ }
478
+
479
+ // Poster embeds can send `{ type: "trackio-logbook:navigate", target: "..." }`
480
+ // from their iframe. Only accept messages from figure frames we created, and
481
+ // only route to pages that are present in this logbook's manifest.
482
+ function registerFigureNavigation(frame) {
483
+ const registerFrameWindow = () => {
484
+ if (frame.contentWindow) FIGURE_FRAME_WINDOWS.add(frame.contentWindow);
485
+ };
486
+ // `srcdoc` replaces the initial about:blank document. Register after that
487
+ // navigation as well, so messages come from the live figure document.
488
+ frame.addEventListener("load", registerFrameWindow);
489
+ registerFrameWindow();
490
+ if (FIGURE_NAVIGATION_READY) return;
491
+ FIGURE_NAVIGATION_READY = true;
492
+ window.addEventListener("message", (event) => {
493
+ if (!FIGURE_FRAME_WINDOWS.has(event.source)) return;
494
+ const message = event.data;
495
+ if (!message || message.type !== "trackio-logbook:navigate") return;
496
+ const target = String(message.target || "").replace(/^#?\//, "");
497
+ if (!target || !MANIFEST || !findNode(MANIFEST.root, target)) return;
498
+ const hash = "#/" + target;
499
+ if (location.hash === hash) scrollToHash();
500
+ else location.hash = hash;
501
+ });
502
+ }
503
+
504
+ const FULLSCREEN_ICON =
505
+ '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" ' +
506
+ 'stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true">' +
507
+ '<path d="M8 3H3v5M16 3h5v5M21 16v5h-5M3 16v5h5"/>' +
508
+ '<path d="M3 8 8 3M16 3l5 5M21 16l-5 5M8 21l-5-5"/></svg>';
509
+
510
+ // Figures are rendered in same-origin iframes, so fullscreen the fitted
511
+ // wrapper rather than the iframe document. This uses the browser's native
512
+ // fullscreen UI and preserves the figure's existing responsive sizing.
513
+ function buildFullscreenControl(figWrap, frame) {
514
+ const wrap = document.createElement("span");
515
+ wrap.className = "cell-fullscreen";
516
+ const btn = document.createElement("button");
517
+ btn.type = "button";
518
+ btn.className = "cell-fullscreen-btn";
519
+ btn.setAttribute("aria-label", "Open figure in fullscreen");
520
+ btn.title = "Open figure in fullscreen";
521
+ btn.innerHTML = FULLSCREEN_ICON;
522
+ wrap.appendChild(btn);
523
+
524
+ btn.addEventListener("click", async () => {
525
+ const request = figWrap.requestFullscreen || figWrap.webkitRequestFullscreen;
526
+ if (!request) return;
527
+ try {
528
+ await request.call(figWrap);
529
+ } catch (_) {
530
+ // Fullscreen can be disabled by the embedding browser or policy.
531
+ }
532
+ });
533
+ document.addEventListener("fullscreenchange", () => {
534
+ if (document.fullscreenElement === figWrap) fitFigureFrame(frame, figWrap);
535
+ });
536
+ return wrap;
537
+ }
538
+
539
+ function extractUrls(text) {
540
+ const seen = new Set();
541
+ const urls = [];
542
+ let match;
543
+ while ((match = DETECTED_URL.exec(text))) {
544
+ const url = match[1].replace(/[.,;:!?'"`]+$/, "");
545
+ if (!seen.has(url)) {
546
+ seen.add(url);
547
+ urls.push(url);
548
+ }
549
+ }
550
+ DETECTED_URL.lastIndex = 0;
551
+ return urls;
552
+ }
553
+
554
+ const IMG_URL = /(\.(png|jpe?g|gif|svg|webp)(\?|$)|\/artifact_blob\/)/i;
555
+
556
+ function renderDetectedEmbeds(text, container) {
557
+ extractUrls(text).forEach((url) => {
558
+ if (url.startsWith("trackio-local-dashboard://")) {
559
+ const div = document.createElement("div");
560
+ div.className = "artifact-chip";
561
+ div.dataset.resUrl = url;
562
+ div.innerHTML =
563
+ "🎯 <strong>Local Trackio dashboard</strong> β€” publish the logbook to share it";
564
+ container.appendChild(div);
565
+ } else if (IMG_URL.test(url)) {
566
+ container.appendChild(renderImage(url));
567
+ } else if (/huggingface\.co\/spaces\//.test(url)) {
568
+ maybeEmbedTrackioSpace(url, container);
569
+ }
570
+ });
571
+ }
572
+
573
+ function renderStandaloneUrl(url) {
574
+ if (IMG_URL.test(url) || IMG_PATH.test(url)) return renderImage(url);
575
+ const item = classifyResource(url);
576
+ if (item) {
577
+ const marker = document.createElement("span");
578
+ marker.className = "resource-anchor";
579
+ marker.dataset.resUrl = item.url;
580
+ marker.setAttribute("aria-hidden", "true");
581
+ return marker;
582
+ }
583
+ const p = document.createElement("p");
584
+ p.innerHTML = inline(url);
585
+ return p;
586
+ }
587
+
588
+ function renderImage(url) {
589
+ const a = document.createElement("a");
590
+ a.className = "unfurl image";
591
+ a.href = url;
592
+ a.target = "_blank";
593
+ a.rel = "noopener";
594
+ const img = document.createElement("img");
595
+ img.loading = "lazy";
596
+ img.src = url;
597
+ img.alt = "artifact image";
598
+ a.appendChild(img);
599
+ return a;
600
+ }
601
+
602
+ function maybeEmbedTrackioSpace(url, container) {
603
+ const id = url.split("/spaces/")[1].split(/[?#]/)[0].replace(/\/$/, "");
604
+ const holder = document.createElement("div");
605
+ container.appendChild(holder);
606
+ getJSON(`https://huggingface.co/api/spaces/${id}`).then((d) => {
607
+ const tags = (d && d.tags) || [];
608
+ if (tags.some((t) => String(t).toLowerCase() === "trackio")) {
609
+ renderTrackioSpaceEmbed(holder, url, id);
610
+ } else {
611
+ holder.remove();
612
+ }
613
+ });
614
+ }
615
+
616
+ function jpGutter(label) {
617
+ const g = document.createElement("div");
618
+ g.className = "jp-gutter";
619
+ g.textContent = label;
620
+ return g;
621
+ }
622
+
623
+ function renderOutArtifact(info) {
624
+ const remote = !info.local && !!info.url;
625
+ const el = document.createElement(remote ? "a" : "div");
626
+ el.className = "out-artifact";
627
+ if (remote) {
628
+ el.href = info.url;
629
+ el.target = "_blank";
630
+ el.rel = "noopener";
631
+ }
632
+ el.dataset.resUrl = info.resUrl;
633
+ const parts = [info.type, info.size].filter(Boolean).map(esc);
634
+ const state = remote
635
+ ? `<span class="out-artifact-state open">Open β†—</span>`
636
+ : `<span class="out-artifact-state">publish to share</span>`;
637
+ const meta = parts.length ? `${parts.join(" Β· ")} Β· ${state}` : state;
638
+ el.innerHTML =
639
+ `<span class="out-artifact-ico">${ARTIFACT_ICON_IMG}</span>` +
640
+ `<span class="out-artifact-name">${esc(info.name)}</span>` +
641
+ `<span class="out-artifact-meta">${meta}</span>`;
642
+ return el;
643
+ }
644
+
645
+ function renderCodeCell(body, container, artifacts) {
646
+ const parts = parseFences(body);
647
+ const block = document.createElement("div");
648
+ block.className = "jp";
649
+ const input = document.createElement("div");
650
+ input.className = "jp-in";
651
+ const inputBody = document.createElement("div");
652
+ inputBody.className = "jp-in-body";
653
+ input.appendChild(jpGutter("In"));
654
+ input.appendChild(inputBody);
655
+ let metaEl = null;
656
+ let outputEl = null;
657
+ let outBody = null;
658
+ const ensureOut = () => {
659
+ if (outputEl) return;
660
+ outputEl = document.createElement("div");
661
+ outputEl.className = "jp-out";
662
+ outputEl.appendChild(jpGutter("Out"));
663
+ outBody = document.createElement("div");
664
+ outBody.className = "jp-out-body";
665
+ outputEl.appendChild(outBody);
666
+ };
667
+ const embedTexts = [];
668
+ parts.forEach((part) => {
669
+ if (part.kind === "text") {
670
+ const text = part.text.trim();
671
+ if (!text) return;
672
+ if (/^exit\s+\S+(\s|Β·)/.test(text)) {
673
+ metaEl = document.createElement("div");
674
+ metaEl.className = "jp-meta";
675
+ metaEl.textContent = text.replace(
676
+ /\s*Β·\s*[A-Z][a-z]{2} \d{1,2}, \d{4}.*$/,
677
+ ""
678
+ );
679
+ } else {
680
+ renderMarkdownPlain(text, container);
681
+ embedTexts.push(text);
682
+ }
683
+ return;
684
+ }
685
+ if (part.kind === "output") {
686
+ ensureOut();
687
+ const pre = document.createElement("pre");
688
+ pre.className = "jp-out-pre";
689
+ const c = document.createElement("code");
690
+ c.textContent = part.text;
691
+ pre.appendChild(c);
692
+ outBody.appendChild(pre);
693
+ outputEl.appendChild(copySnippetBtn(part.text));
694
+ embedTexts.push(part.text);
695
+ return;
696
+ }
697
+ inputBody.appendChild(renderCode(part.text, part.lang, part.title));
698
+ });
699
+ if (artifacts && artifacts.length) {
700
+ ensureOut();
701
+ const artWrap = document.createElement("div");
702
+ artWrap.className = "jp-artifacts";
703
+ artifacts.forEach((a) => {
704
+ artWrap.appendChild(
705
+ renderOutArtifact(artifactInfoFromCell(a.meta, a.body))
706
+ );
707
+ });
708
+ outBody.appendChild(artWrap);
709
+ }
710
+ if (inputBody.childNodes.length > 0) block.appendChild(input);
711
+ if (metaEl) block.appendChild(metaEl);
712
+ if (outputEl) block.appendChild(outputEl);
713
+ if (block.childNodes.length) container.appendChild(block);
714
+ embedTexts.forEach((text) => renderDetectedEmbeds(text, container));
715
+ }
716
+
717
+ function parseRow(line) {
718
+ let s = line.trim();
719
+ if (s.startsWith("|")) s = s.slice(1);
720
+ if (s.endsWith("|")) s = s.slice(0, -1);
721
+ return s.split(/(?<!\\)\|/).map((c) => c.replace(/\\\|/g, "|").trim());
722
+ }
723
+
724
+ const TRUTHY = ["x", "βœ“", "βœ”", "yes", "done", "true", "[x]"];
725
+ const CHIP_COLORS = [
726
+ ["#e7f0ff", "#2158d0"],
727
+ ["#fde8ec", "#c62a4b"],
728
+ ["#e6f7ee", "#1a8a55"],
729
+ ["#fdf0e0", "#b26a12"],
730
+ ["#efe9ff", "#5b3bd6"],
731
+ ["#e6f6f8", "#127b88"],
732
+ ];
733
+
734
+ function chipColor(name) {
735
+ let h = 0;
736
+ for (let i = 0; i < name.length; i++) h = (h * 31 + name.charCodeAt(i)) >>> 0;
737
+ return CHIP_COLORS[h % CHIP_COLORS.length];
738
+ }
739
+
740
+ const STATUS_MAP = {
741
+ "": ["Planned", "gray"],
742
+ planned: ["Planned", "gray"],
743
+ todo: ["Planned", "gray"],
744
+ "to do": ["Planned", "gray"],
745
+ backlog: ["Planned", "gray"],
746
+ "in progress": ["In progress", "amber"],
747
+ "in-progress": ["In progress", "amber"],
748
+ wip: ["In progress", "amber"],
749
+ running: ["In progress", "amber"],
750
+ active: ["In progress", "amber"],
751
+ done: ["Done", "green"],
752
+ complete: ["Done", "green"],
753
+ completed: ["Done", "green"],
754
+ blocked: ["Blocked", "red"],
755
+ failed: ["Failed", "red"],
756
+ abandoned: ["Abandoned", "gray"],
757
+ };
758
+
759
+ function statusBadge(val) {
760
+ const [label, tone] = STATUS_MAP[val.toLowerCase()] || [val || "β€”", "gray"];
761
+ return `<span class="badge ${tone}">${esc(label)}</span>`;
762
+ }
763
+
764
+ function renderTable(rows, container) {
765
+ if (rows.length < 2) return;
766
+ const header = rows[0];
767
+ const body = rows.slice(2);
768
+ const roles = header.map((h) => {
769
+ const t = h.toLowerCase();
770
+ if (t.includes("status") || t.includes("state")) return "status";
771
+ if (t.includes("progress") || t.includes("complete") || t.includes("done"))
772
+ return "check";
773
+ if (t === "who" || t.includes("assign") || t.includes("owner")) return "who";
774
+ return "text";
775
+ });
776
+ const table = document.createElement("table");
777
+ table.className = "board";
778
+ const thead = document.createElement("thead");
779
+ const htr = document.createElement("tr");
780
+ header.forEach((h, c) => {
781
+ const th = document.createElement("th");
782
+ th.textContent = h;
783
+ if (roles[c] === "check") th.className = "col-check";
784
+ htr.appendChild(th);
785
+ });
786
+ thead.appendChild(htr);
787
+ table.appendChild(thead);
788
+ const tbody = document.createElement("tbody");
789
+ body.forEach((cells) => {
790
+ const nonEmpty = cells.filter((x) => x !== "").length;
791
+ if (header.length > 1 && nonEmpty === 1 && cells[0]) {
792
+ const tr = document.createElement("tr");
793
+ tr.className = "section-row";
794
+ const td = document.createElement("td");
795
+ td.colSpan = header.length;
796
+ td.innerHTML = inline(cells[0]);
797
+ tr.appendChild(td);
798
+ tbody.appendChild(tr);
799
+ return;
800
+ }
801
+ const tr = document.createElement("tr");
802
+ header.forEach((_, c) => {
803
+ const td = document.createElement("td");
804
+ const val = (cells[c] || "").trim();
805
+ if (roles[c] === "status") {
806
+ td.className = "col-status";
807
+ td.innerHTML = statusBadge(val);
808
+ } else if (roles[c] === "check") {
809
+ td.className = "col-check";
810
+ const on = TRUTHY.indexOf(val.toLowerCase()) !== -1;
811
+ td.innerHTML = `<span class="box ${on ? "on" : ""}">${on ? "βœ“" : ""}</span>`;
812
+ } else if (roles[c] === "who") {
813
+ if (!val || /^to assign$/i.test(val)) {
814
+ td.innerHTML = `<span class="who-chip muted">${esc(val || "β€”")}</span>`;
815
+ } else {
816
+ const [bg, fg] = chipColor(val);
817
+ td.innerHTML = `<span class="who-chip" style="background:${bg};color:${fg}">${esc(val)}</span>`;
818
+ }
819
+ } else {
820
+ td.innerHTML = inline(val);
821
+ }
822
+ tr.appendChild(td);
823
+ });
824
+ const link = tr.querySelector('a[href^="#/"]');
825
+ if (link) {
826
+ tr.classList.add("linked-row");
827
+ tr.addEventListener("click", (e) => {
828
+ if (e.target.tagName !== "A") location.hash = link.getAttribute("href");
829
+ });
830
+ }
831
+ tbody.appendChild(tr);
832
+ });
833
+ table.appendChild(tbody);
834
+ const wrap = document.createElement("div");
835
+ wrap.className = "board-wrap";
836
+ wrap.appendChild(table);
837
+ container.appendChild(wrap);
838
+ }
839
+
840
+ const HL_RULES = {
841
+ python: [
842
+ ["comment", /#[^\n]*/],
843
+ ["string", /'''[\s\S]*?'''|"""[\s\S]*?"""|'(?:\\.|[^'\\])*'|"(?:\\.|[^"\\])*"/],
844
+ [
845
+ "keyword",
846
+ /\b(?:def|class|return|if|elif|else|for|while|import|from|as|with|try|except|finally|raise|in|not|and|or|is|None|True|False|lambda|yield|global|nonlocal|assert|pass|break|continue|async|await|print)\b/,
847
+ ],
848
+ ["number", /\b\d[\d_.eE+-]*\b/],
849
+ ],
850
+ bash: [
851
+ ["comment", /#[^\n]*/],
852
+ ["string", /'(?:\\.|[^'\\])*'|"(?:\\.|[^"\\])*"/],
853
+ ["keyword", /\b(?:if|then|else|fi|for|in|do|done|while|case|esac|function|export|source|echo|cd|return|local)\b/],
854
+ ["number", /(?<=\s)-{1,2}[a-zA-Z][\w-]*/],
855
+ ],
856
+ json: [
857
+ ["string", /"(?:\\.|[^"\\])*"/],
858
+ ["keyword", /\b(?:true|false|null)\b/],
859
+ ["number", /-?\b\d[\d.eE+-]*\b/],
860
+ ],
861
+ yaml: [
862
+ ["comment", /#[^\n]*/],
863
+ ["string", /'(?:\\.|[^'\\])*'|"(?:\\.|[^"\\])*"/],
864
+ ["keyword", /\b(?:true|false|null|yes|no)\b/],
865
+ ["number", /-?\b\d[\d.eE+-]*\b/],
866
+ ],
867
+ };
868
+ HL_RULES.javascript = HL_RULES.python;
869
+ HL_RULES.typescript = HL_RULES.python;
870
+ HL_RULES.sql = [
871
+ ["comment", /--[^\n]*/],
872
+ ["string", /'(?:\\.|[^'\\])*'/],
873
+ [
874
+ "keyword",
875
+ /\b(?:SELECT|FROM|WHERE|JOIN|LEFT|RIGHT|INNER|OUTER|ON|GROUP|BY|ORDER|LIMIT|INSERT|INTO|VALUES|UPDATE|SET|DELETE|CREATE|TABLE|AS|AND|OR|NOT|NULL|COUNT|DISTINCT|IN)\b/i,
876
+ ],
877
+ ["number", /\b\d[\d.]*\b/],
878
+ ];
879
+
880
+ function highlightCode(code, lang) {
881
+ const rules = HL_RULES[lang];
882
+ if (!rules) return esc(code);
883
+ const combined = new RegExp(rules.map((r) => "(" + r[1].source + ")").join("|"), "g");
884
+ let out = "";
885
+ let last = 0;
886
+ let m;
887
+ while ((m = combined.exec(code))) {
888
+ if (m[0] === "") {
889
+ combined.lastIndex++;
890
+ continue;
891
+ }
892
+ out += esc(code.slice(last, m.index));
893
+ let gi = 1;
894
+ while (gi < m.length && m[gi] === undefined) gi++;
895
+ out += `<span class="tok-${rules[gi - 1][0]}">${esc(m[0])}</span>`;
896
+ last = m.index + m[0].length;
897
+ }
898
+ out += esc(code.slice(last));
899
+ return out;
900
+ }
901
+
902
+ function copySnippetBtn(text) {
903
+ const btn = document.createElement("button");
904
+ btn.type = "button";
905
+ btn.className = "copy-snippet";
906
+ btn.title = "Copy";
907
+ btn.textContent = "⧉";
908
+ btn.addEventListener("click", (e) => {
909
+ e.preventDefault();
910
+ e.stopPropagation();
911
+ copyText(text, btn, "⧉");
912
+ });
913
+ return btn;
914
+ }
915
+
916
+ function renderCode(code, lang, title) {
917
+ const pre = document.createElement("pre");
918
+ pre.className = "hl";
919
+ const c = document.createElement("code");
920
+ c.innerHTML = highlightCode(code, lang);
921
+ pre.appendChild(c);
922
+ if (!title) {
923
+ const wrap = document.createElement("div");
924
+ wrap.className = "snippet";
925
+ wrap.appendChild(pre);
926
+ wrap.appendChild(copySnippetBtn(code));
927
+ return wrap;
928
+ }
929
+ const det = document.createElement("details");
930
+ det.className = "code-accordion";
931
+ det.dataset.resUrl = `trackio-script://${title}`;
932
+ const sum = document.createElement("summary");
933
+ sum.innerHTML =
934
+ `<span class="code-ico">&lt;/&gt;</span>` +
935
+ `<span class="code-name">${esc(title)}</span>`;
936
+ sum
937
+ .querySelector(".code-name")
938
+ .addEventListener("click", (e) => e.preventDefault());
939
+ det.appendChild(sum);
940
+ const wrap = document.createElement("div");
941
+ wrap.className = "snippet";
942
+ wrap.appendChild(pre);
943
+ wrap.appendChild(copySnippetBtn(code));
944
+ det.appendChild(wrap);
945
+ return det;
946
+ }
947
+
948
+ const IMG_PATH = /^[^\s]+\.(png|jpe?g|gif|svg|webp)$/i;
949
+
950
+ function renderList(items, container) {
951
+ let ul = null;
952
+ items.forEach((item) => {
953
+ if (URL_ONLY.test(item) || IMG_PATH.test(item)) {
954
+ const el = renderStandaloneUrl(item);
955
+ if (el) {
956
+ ul = null;
957
+ container.appendChild(el);
958
+ }
959
+ } else if (item.indexOf("πŸ“¦ Artifact") !== -1) {
960
+ ul = null;
961
+ const div = document.createElement("div");
962
+ div.className = "artifact-chip";
963
+ div.innerHTML = inline(item.replace("πŸ“¦", "πŸͺ£"));
964
+ container.appendChild(div);
965
+ } else if (item.indexOf("trackio-local-dashboard://") !== -1) {
966
+ ul = null;
967
+ const uri = item.match(/trackio-local-dashboard:\/\/\S+/)?.[0] || "";
968
+ const div = document.createElement("div");
969
+ div.className = "artifact-chip";
970
+ if (uri) div.dataset.resUrl = uri;
971
+ div.innerHTML =
972
+ "🎯 <strong>Local dashboard</strong> β€” publish the logbook to share it";
973
+ container.appendChild(div);
974
+ } else {
975
+ if (!ul) {
976
+ ul = document.createElement("ul");
977
+ container.appendChild(ul);
978
+ }
979
+ const li = document.createElement("li");
980
+ li.innerHTML = inline(item);
981
+ ul.appendChild(li);
982
+ }
983
+ });
984
+ }
985
+
986
+ /* -------------------- resources rail -------------------- */
987
+
988
+ function fmt(n) {
989
+ if (n == null) return null;
990
+ if (n >= 1e6) return (n / 1e6).toFixed(1) + "M";
991
+ if (n >= 1e3) return (n / 1e3).toFixed(1) + "k";
992
+ return String(n);
993
+ }
994
+
995
+ const RESOURCE_SECTIONS = [
996
+ ["dashboard", "Dashboards", "🎯"],
997
+ ["model", "Models", "πŸ€—"],
998
+ ["dataset", "Datasets", "πŸ“Š"],
999
+ ["space", "Spaces", "πŸš€"],
1000
+ ["artifact", "Artifacts", "πŸͺ£"],
1001
+ ["paper", "Papers", "πŸ“„"],
1002
+ ["repo", "Code", "πŸ™"],
1003
+ ["job", "Jobs", "βš™οΈ"],
1004
+ ["bucket", "Buckets", "πŸͺ£"],
1005
+ ];
1006
+
1007
+ const RESOURCE_ICONS = Object.fromEntries(
1008
+ RESOURCE_SECTIONS.map(([kind, , icon]) => [kind, icon])
1009
+ );
1010
+
1011
+ const ARTIFACT_ICON_IMG = `<img class="art-ico" src="./bucket-icon.svg" alt="" />`;
1012
+ const DASHBOARD_ICON_IMG = `<img class="art-ico" src="./trackio-logo-light.png" alt="" />`;
1013
+
1014
+ const RESOURCE_DESC = {
1015
+ dashboard: "Dashboard",
1016
+ model: "Model",
1017
+ dataset: "Dataset",
1018
+ space: "Space",
1019
+ artifact: "Artifact β€” in Bucket",
1020
+ paper: "Paper",
1021
+ repo: "Repository",
1022
+ job: "Job β€” status & logs",
1023
+ bucket: "Bucket β€” artifacts & data",
1024
+ };
1025
+
1026
+ const HF_NON_MODEL_PREFIX =
1027
+ /^(datasets|spaces|jobs|buckets|papers|blog|docs|api|posts|collections|organizations|settings|new|join|login|pricing|tasks|learn|chat|models)(\/|$)/;
1028
+
1029
+ function hfId(url, marker) {
1030
+ return url.split(marker)[1].split(/[?#]/)[0].replace(/\/$/, "");
1031
+ }
1032
+
1033
+ function classifyResource(url) {
1034
+ if (IMG_URL.test(url)) {
1035
+ return null;
1036
+ }
1037
+ let m;
1038
+ if (url.startsWith("trackio-local-dashboard://")) {
1039
+ return {
1040
+ kind: "dashboard",
1041
+ id: url.slice("trackio-local-dashboard://".length),
1042
+ url,
1043
+ local: true,
1044
+ };
1045
+ }
1046
+ if (url.startsWith("trackio-artifact://")) {
1047
+ return {
1048
+ kind: "artifact",
1049
+ id: url.slice("trackio-artifact://".length),
1050
+ url,
1051
+ local: true,
1052
+ };
1053
+ }
1054
+ if (url.startsWith("trackio-local-path://")) {
1055
+ return {
1056
+ kind: "artifact",
1057
+ id: url.slice("trackio-local-path://".length),
1058
+ url,
1059
+ local: true,
1060
+ };
1061
+ }
1062
+ if ((m = url.match(/huggingface\.co\/buckets\/[^#\s]+#(.+)/))) {
1063
+ return { kind: "artifact", id: decodeURIComponent(m[1]), url };
1064
+ }
1065
+ if (/huggingface\.co\/datasets\/[^/]+\/[^/]+/.test(url)) {
1066
+ return { kind: "dataset", id: hfId(url, "/datasets/"), url };
1067
+ }
1068
+ if (/huggingface\.co\/spaces\/[^/]+\/[^/]+/.test(url)) {
1069
+ return { kind: "space", id: hfId(url, "/spaces/"), url };
1070
+ }
1071
+ if (/huggingface\.co\/jobs\//.test(url)) {
1072
+ const parts = hfId(url, "/jobs/").split("/");
1073
+ const jid = parts[1] || "";
1074
+ return {
1075
+ kind: "job",
1076
+ id: parts[0] + (jid ? ` Β· ${jid.slice(0, 12)}${jid.length > 12 ? "…" : ""}` : ""),
1077
+ url,
1078
+ };
1079
+ }
1080
+ if (/huggingface\.co\/buckets\//.test(url)) {
1081
+ return { kind: "bucket", id: hfId(url, "/buckets/"), url };
1082
+ }
1083
+ if (/huggingface\.co\/papers\//.test(url)) {
1084
+ return { kind: "paper", id: `Paper ${hfId(url, "/papers/")}`, url };
1085
+ }
1086
+ if ((m = url.match(/arxiv\.org\/(?:abs|pdf)\/([^?#\s]+)/))) {
1087
+ return { kind: "paper", id: `arXiv:${m[1].replace(/\.pdf$/, "")}`, url };
1088
+ }
1089
+ if ((m = url.match(/github\.com\/([^/?#]+\/[^/?#]+)/))) {
1090
+ return { kind: "repo", id: m[1], url };
1091
+ }
1092
+ if ((m = url.match(/huggingface\.co\/([^?#]+)/))) {
1093
+ const rest = m[1].replace(/\/$/, "");
1094
+ if (/^[^/]+\/[^/]+$/.test(rest) && !HF_NON_MODEL_PREFIX.test(rest)) {
1095
+ return { kind: "model", id: rest, url };
1096
+ }
1097
+ }
1098
+ return null;
1099
+ }
1100
+
1101
+ async function fillRailMeta(item, el) {
1102
+ if (item.local) return;
1103
+ const meta = el.querySelector(".rail-meta");
1104
+ const set = (parts) => {
1105
+ const text = parts.filter(Boolean).join(" Β· ");
1106
+ if (text) meta.textContent = text;
1107
+ };
1108
+ if (item.kind === "model") {
1109
+ const d = await getJSON(`https://huggingface.co/api/models/${item.id}`);
1110
+ if (d) set([d.pipeline_tag, `↓ ${fmt(d.downloads)}`, `β™₯ ${fmt(d.likes)}`]);
1111
+ } else if (item.kind === "dataset") {
1112
+ const d = await getJSON(`https://huggingface.co/api/datasets/${item.id}`);
1113
+ if (d) set([`↓ ${fmt(d.downloads)}`, `β™₯ ${fmt(d.likes)}`]);
1114
+ } else if (item.kind === "space" || item.kind === "dashboard") {
1115
+ const d = await getJSON(`https://huggingface.co/api/spaces/${item.id}`);
1116
+ if (d) set([d.sdk, `β™₯ ${fmt(d.likes)}`]);
1117
+ } else if (item.kind === "repo") {
1118
+ const d = await getJSON(`https://api.github.com/repos/${item.id}`);
1119
+ if (d) set([`β˜… ${fmt(d.stargazers_count)}`, d.language]);
1120
+ } else if (item.kind === "paper") {
1121
+ const m = item.id.match(/^(?:arXiv:|Paper )(.+)$/);
1122
+ if (!m) return;
1123
+ const arxivId = m[1].replace(/v\d+$/, "");
1124
+ const d = await getJSON(`https://huggingface.co/api/papers/${arxivId}`);
1125
+ if (d && d.id) {
1126
+ if (el.href) el.href = `https://huggingface.co/papers/${d.id}`;
1127
+ const title =
1128
+ d.title && d.title.length > 70 ? `${d.title.slice(0, 69)}…` : d.title;
1129
+ set([title, d.upvotes ? `β–² ${fmt(d.upvotes)}` : null]);
1130
+ }
1131
+ }
1132
+ }
1133
+
1134
+ const BARE_ID_SKIP_DIRS = new Set([
1135
+ "scripts",
1136
+ "configs",
1137
+ "config",
1138
+ "results",
1139
+ "figures",
1140
+ "data",
1141
+ "datasets",
1142
+ "src",
1143
+ "tests",
1144
+ "test",
1145
+ "examples",
1146
+ "pages",
1147
+ "assets",
1148
+ "docs",
1149
+ "outputs",
1150
+ "output",
1151
+ "checkpoints",
1152
+ "models",
1153
+ "utils",
1154
+ "lib",
1155
+ "bin",
1156
+ "tmp",
1157
+ "node_modules",
1158
+ "dist",
1159
+ "build",
1160
+ ]);
1161
+ const FILE_EXT_RE =
1162
+ /\.(py|pyc|js|ts|jsx|tsx|json|jsonl|yaml|yml|csv|tsv|md|txt|sh|bash|html|css|png|jpe?g|svg|gif|webp|ipynb|toml|cfg|ini|lock|pdf|whl|gz|zip|tar|pt|pth|bin|safetensors|db|sqlite)$/i;
1163
+
1164
+ async function detectBareModelIds(text, groups) {
1165
+ const stripped = text.replace(DETECTED_URL, " ");
1166
+ DETECTED_URL.lastIndex = 0;
1167
+ const seen = new Set();
1168
+ const candidates = [];
1169
+ const re = /(^|[\s"'`(=[])([A-Za-z0-9][\w.-]*\/[A-Za-z0-9][\w.-]*)/g;
1170
+ let m;
1171
+ while ((m = re.exec(stripped)) && candidates.length < 15) {
1172
+ const id = m[2].replace(/[.:,]+$/, "");
1173
+ if (seen.has(id)) continue;
1174
+ seen.add(id);
1175
+ if (FILE_EXT_RE.test(id)) continue;
1176
+ if (BARE_ID_SKIP_DIRS.has(id.split("/")[0].toLowerCase())) continue;
1177
+ candidates.push(id);
1178
+ }
1179
+ const results = await Promise.all(
1180
+ candidates.map((id) => getJSON(`https://huggingface.co/api/models/${id}`))
1181
+ );
1182
+ let added = false;
1183
+ const confirmed = [];
1184
+ results.forEach((d, i) => {
1185
+ if (!d || !d.id) return;
1186
+ const id = candidates[i];
1187
+ confirmed.push(id);
1188
+ const url = `https://huggingface.co/${id}`;
1189
+ if (!groups.has("model")) groups.set("model", new Map());
1190
+ if (!groups.get("model").has(url)) {
1191
+ groups.get("model").set(url, { kind: "model", id, url });
1192
+ added = true;
1193
+ }
1194
+ });
1195
+ return { added, confirmed };
1196
+ }
1197
+
1198
+ function chipifyBareIds(ids, container) {
1199
+ if (!ids.length) return;
1200
+ const escaped = ids.map((id) => id.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"));
1201
+ const pattern = new RegExp("(" + escaped.join("|") + ")");
1202
+ const splitter = new RegExp(pattern.source, "g");
1203
+ container
1204
+ .querySelectorAll(".cell.markdown .cell-body")
1205
+ .forEach((body) => {
1206
+ const walker = document.createTreeWalker(body, NodeFilter.SHOW_TEXT, {
1207
+ acceptNode(node) {
1208
+ if (!pattern.test(node.nodeValue)) return NodeFilter.FILTER_REJECT;
1209
+ for (
1210
+ let el = node.parentElement;
1211
+ el && el !== body;
1212
+ el = el.parentElement
1213
+ ) {
1214
+ if (["A", "CODE", "PRE", "BUTTON"].indexOf(el.tagName) !== -1) {
1215
+ return NodeFilter.FILTER_REJECT;
1216
+ }
1217
+ }
1218
+ return NodeFilter.FILTER_ACCEPT;
1219
+ },
1220
+ });
1221
+ const nodes = [];
1222
+ while (walker.nextNode()) nodes.push(walker.currentNode);
1223
+ nodes.forEach((node) => {
1224
+ const frag = document.createDocumentFragment();
1225
+ node.nodeValue.split(splitter).forEach((part) => {
1226
+ if (ids.indexOf(part) !== -1) {
1227
+ const holder = document.createElement("span");
1228
+ holder.innerHTML = resChipHtml({
1229
+ kind: "model",
1230
+ id: part,
1231
+ url: `https://huggingface.co/${part}`,
1232
+ });
1233
+ frag.appendChild(holder.firstChild);
1234
+ } else if (part) {
1235
+ frag.appendChild(document.createTextNode(part));
1236
+ }
1237
+ });
1238
+ node.parentNode.replaceChild(frag, node);
1239
+ });
1240
+ });
1241
+ }
1242
+
1243
+ let RAIL_TOKEN = 0;
1244
+ const RAIL_EXCLUDE_KINDS = new Set(["paper", "repo", "artifact", "dashboard"]);
1245
+
1246
+ function railDashboardItem(it) {
1247
+ return {
1248
+ kind: "dashboard",
1249
+ id: it.id,
1250
+ url: it.local ? it.resUrl : it.url || it.resUrl,
1251
+ local: it.local,
1252
+ railLabel: "Dashboard",
1253
+ };
1254
+ }
1255
+
1256
+ function promoteTrackioSpacesInRail(groups, dashResUrls, body, rail, token) {
1257
+ const spaceGroup = groups.get("space");
1258
+ if (!spaceGroup || !spaceGroup.size) return;
1259
+ spaceGroup.forEach((item, url) => {
1260
+ getJSON(`https://huggingface.co/api/spaces/${item.id}`)
1261
+ .then((d) => {
1262
+ if (rail.dataset.renderToken !== token) return;
1263
+ const tags = (d && d.tags) || [];
1264
+ if (!tags.some((t) => String(t).toLowerCase() === "trackio")) return;
1265
+ if (dashResUrls.has(url)) return;
1266
+ spaceGroup.delete(url);
1267
+ if (!spaceGroup.size) groups.delete("space");
1268
+ if (!groups.has("dashboard")) groups.set("dashboard", new Map());
1269
+ groups.get("dashboard").set(url, {
1270
+ kind: "dashboard",
1271
+ id: item.id,
1272
+ url: item.url,
1273
+ local: false,
1274
+ railLabel: "Dashboard",
1275
+ });
1276
+ dashResUrls.add(url);
1277
+ paintRail(groups, body, rail);
1278
+ })
1279
+ .catch(() => {});
1280
+ });
1281
+ }
1282
+
1283
+ function renderRail(md, body, rail) {
1284
+ const token = String(++RAIL_TOKEN);
1285
+ rail.dataset.renderToken = token;
1286
+ const scanText = md.replace(
1287
+ /(`{3,4}|~{3,4})(html|raw)[^\n]*\n[\s\S]*?\n\1/g,
1288
+ " "
1289
+ );
1290
+ const groups = new Map();
1291
+ const dashMap = new Map();
1292
+ const dashResUrls = new Set();
1293
+ cellDashboardItems(md).forEach((it) => {
1294
+ if (dashMap.has(it.resUrl)) return;
1295
+ dashMap.set(it.resUrl, railDashboardItem(it));
1296
+ dashResUrls.add(it.resUrl);
1297
+ });
1298
+ if (dashMap.size) groups.set("dashboard", dashMap);
1299
+ extractUrls(scanText).forEach((url) => {
1300
+ const item = classifyResource(url);
1301
+ if (!item) return;
1302
+ if (RAIL_EXCLUDE_KINDS.has(item.kind)) return;
1303
+ if (dashResUrls.has(url)) return;
1304
+ if (!groups.has(item.kind)) groups.set(item.kind, new Map());
1305
+ groups.get(item.kind).set(item.url, item);
1306
+ });
1307
+ const artMap = new Map();
1308
+ cellArtifactItems(md).forEach((it) => {
1309
+ if (artMap.has(it.resUrl)) return;
1310
+ const label = it.type
1311
+ ? it.type.charAt(0).toUpperCase() + it.type.slice(1)
1312
+ : "Artifact";
1313
+ artMap.set(it.resUrl, {
1314
+ kind: "artifact",
1315
+ id: it.name,
1316
+ url: it.local ? it.resUrl : it.url || it.resUrl,
1317
+ local: it.local,
1318
+ railLabel: label,
1319
+ size: it.size,
1320
+ });
1321
+ });
1322
+ if (artMap.size) groups.set("artifact", artMap);
1323
+ paintRail(groups, body, rail);
1324
+ promoteTrackioSpacesInRail(groups, dashResUrls, body, rail, token);
1325
+ detectBareModelIds(scanText, groups)
1326
+ .then((result) => {
1327
+ if (rail.dataset.renderToken !== token) return;
1328
+ chipifyBareIds(result.confirmed, body);
1329
+ if (result.added) paintRail(groups, body, rail);
1330
+ })
1331
+ .catch(() => {});
1332
+ }
1333
+
1334
+ function paintRail(groups, body, rail) {
1335
+ rail.innerHTML = "";
1336
+ RESOURCE_SECTIONS.forEach(([kind, label, icon]) => {
1337
+ const group = groups.get(kind);
1338
+ if (!group || !group.size) return;
1339
+ group.forEach((item) => {
1340
+ const el = document.createElement(item.local ? "div" : "a");
1341
+ el.className = item.local ? "rail-item rail-local" : "rail-item";
1342
+ if (!item.local) {
1343
+ el.href = item.url;
1344
+ el.target = "_blank";
1345
+ el.rel = "noopener";
1346
+ }
1347
+ el.dataset.resUrl = item.url;
1348
+ let desc;
1349
+ if (kind === "artifact") {
1350
+ const state = item.local ? "publish to share" : "Open β†—";
1351
+ desc = item.size ? `${item.size} Β· ${state}` : state;
1352
+ } else if (kind === "dashboard") {
1353
+ desc = item.local ? "publish to share" : "Open β†—";
1354
+ } else {
1355
+ desc = item.local ? "publish to share" : RESOURCE_DESC[kind];
1356
+ }
1357
+ const kindLabel = item.railLabel || label.replace(/s$/, "");
1358
+ const iconHtml =
1359
+ kind === "artifact"
1360
+ ? ARTIFACT_ICON_IMG
1361
+ : kind === "dashboard"
1362
+ ? DASHBOARD_ICON_IMG
1363
+ : `<span>${icon}</span>`;
1364
+ el.innerHTML =
1365
+ `<div class="rail-kind">${iconHtml}${esc(kindLabel)}</div>` +
1366
+ `<div class="rail-title">${esc(item.id)}</div>` +
1367
+ `<div class="rail-meta">${esc(desc)}</div>`;
1368
+ rail.appendChild(el);
1369
+ fillRailMeta(item, el)
1370
+ .catch(() => {})
1371
+ .finally(() => scheduleRailPosition(body, rail));
1372
+ });
1373
+ });
1374
+ rail.hidden = !rail.childElementCount;
1375
+ scheduleRailPosition(body, rail);
1376
+ }
1377
+
1378
+ function resourceAnchor(body, url) {
1379
+ return body.querySelector(`[data-res-url="${CSS.escape(url)}"]`);
1380
+ }
1381
+
1382
+ function positionRail(body, rail) {
1383
+ if (rail.hidden || !rail.isConnected) return;
1384
+ const bodyRect = body.getBoundingClientRect();
1385
+ const items = Array.from(rail.querySelectorAll(".rail-item")).map((el, index) => {
1386
+ const anchor = resourceAnchor(body, el.dataset.resUrl);
1387
+ return {
1388
+ el,
1389
+ index,
1390
+ desired: anchor
1391
+ ? Math.max(0, anchor.getBoundingClientRect().top - bodyRect.top)
1392
+ : 0,
1393
+ };
1394
+ });
1395
+ items.sort((a, b) => a.desired - b.desired || a.index - b.index);
1396
+ let cursor = 0;
1397
+ items.forEach(({ el, desired }) => {
1398
+ const top = Math.max(desired, cursor);
1399
+ el.style.top = `${top}px`;
1400
+ cursor = top + el.offsetHeight + 10;
1401
+ });
1402
+ rail.style.minHeight = `${Math.max(body.offsetHeight, cursor)}px`;
1403
+ }
1404
+
1405
+ function scheduleRailPosition(body, rail) {
1406
+ cancelAnimationFrame(Number(rail.dataset.positionFrame || 0));
1407
+ rail.dataset.positionFrame = String(
1408
+ requestAnimationFrame(() => positionRail(body, rail))
1409
+ );
1410
+ }
1411
+
1412
+ function dashboardSubdomainFromUrl(url) {
1413
+ return spaceIdFromUrl(url).toLowerCase().replace(/[^a-z0-9-]/g, "-");
1414
+ }
1415
+
1416
+ function dashboardOpenLink(head, url) {
1417
+ if (!head || !url) return;
1418
+ const meta = head.querySelector(".cell-meta");
1419
+ if (!meta) return;
1420
+ let link = meta.querySelector(".cell-open");
1421
+ if (!link) {
1422
+ link = document.createElement("a");
1423
+ link.className = "cell-open";
1424
+ link.target = "_blank";
1425
+ link.rel = "noopener";
1426
+ meta.insertBefore(link, meta.firstChild);
1427
+ }
1428
+ link.href = url;
1429
+ link.textContent = "Open β†—";
1430
+ }
1431
+
1432
+ function dashboardFrame(src) {
1433
+ const iframe = document.createElement("iframe");
1434
+ iframe.className = "dashboard-frame";
1435
+ iframe.src = src;
1436
+ iframe.loading = "lazy";
1437
+ iframe.allow = "clipboard-read; clipboard-write; fullscreen";
1438
+ return iframe;
1439
+ }
1440
+
1441
+ function renderDashboardCell(meta, body, container, head) {
1442
+ const project = meta.dashboard_project || "";
1443
+ const holder = document.createElement("div");
1444
+ holder.className = "dashboard-shell";
1445
+ container.appendChild(holder);
1446
+ const space = body.match(/https:\/\/huggingface\.co\/spaces\/[^\s<>)"'`]+/);
1447
+ if (space) {
1448
+ const url = space[0];
1449
+ dashboardOpenLink(head, url);
1450
+ holder.appendChild(
1451
+ dashboardFrame(
1452
+ `https://${dashboardSubdomainFromUrl(url)}.hf.space/?sidebar=hidden&hide_empty_tabs=true`
1453
+ )
1454
+ );
1455
+ return;
1456
+ }
1457
+ if (!isLocalPreview()) {
1458
+ holder.className = "artifact-chip";
1459
+ holder.dataset.resUrl = `trackio-local-dashboard://${project}`;
1460
+ holder.innerHTML =
1461
+ "🎯 <strong>Local Trackio dashboard</strong> β€” publish the logbook to share it";
1462
+ return;
1463
+ }
1464
+ const open = "/dashboard/?project=" + encodeURIComponent(project);
1465
+ dashboardOpenLink(head, open);
1466
+ holder.appendChild(
1467
+ dashboardFrame(open + "&sidebar=hidden&hide_empty_tabs=true"),
1468
+ );
1469
+ }
1470
+
1471
+ const CACHE_PREFIX = "trackio-logbook:";
1472
+ const CACHE_TTL_MS = 24 * 60 * 60 * 1000;
1473
+ const CACHE_MISS_TTL_MS = 60 * 60 * 1000;
1474
+
1475
+ function cacheGet(url) {
1476
+ try {
1477
+ const raw = localStorage.getItem(CACHE_PREFIX + url);
1478
+ if (!raw) return undefined;
1479
+ const entry = JSON.parse(raw);
1480
+ const ttl = entry.d === null ? CACHE_MISS_TTL_MS : CACHE_TTL_MS;
1481
+ if (Date.now() - entry.t > ttl) {
1482
+ localStorage.removeItem(CACHE_PREFIX + url);
1483
+ return undefined;
1484
+ }
1485
+ return entry.d;
1486
+ } catch (e) {
1487
+ return undefined;
1488
+ }
1489
+ }
1490
+
1491
+ function cacheSet(url, data) {
1492
+ try {
1493
+ localStorage.setItem(
1494
+ CACHE_PREFIX + url,
1495
+ JSON.stringify({ t: Date.now(), d: data })
1496
+ );
1497
+ } catch (e) {}
1498
+ }
1499
+
1500
+ async function getJSON(url) {
1501
+ if (UNFURL_CACHE[url] !== undefined) return UNFURL_CACHE[url];
1502
+ const cached = cacheGet(url);
1503
+ if (cached !== undefined) {
1504
+ UNFURL_CACHE[url] = cached;
1505
+ return cached;
1506
+ }
1507
+ try {
1508
+ const r = await fetch(url);
1509
+ if (!r.ok) throw new Error(r.status);
1510
+ const j = await r.json();
1511
+ UNFURL_CACHE[url] = j;
1512
+ cacheSet(url, j);
1513
+ return j;
1514
+ } catch (e) {
1515
+ UNFURL_CACHE[url] = null;
1516
+ cacheSet(url, null);
1517
+ return null;
1518
+ }
1519
+ }
1520
+
1521
+ /* -------------------- routing / render -------------------- */
1522
+
1523
+ function buildTree() {
1524
+ const tree = document.getElementById("tree");
1525
+ tree.innerHTML = "";
1526
+ const nodes = [];
1527
+ (MANIFEST.root.children || []).forEach((c) => flattenTree(c, 0, nodes));
1528
+ nodes.forEach(({ node, depth }) => {
1529
+ const a = document.createElement("a");
1530
+ a.href = "#/" + node.slug;
1531
+ a.className = "depth-" + depth;
1532
+ a.dataset.slug = node.slug;
1533
+ const mark = document.createElement("span");
1534
+ mark.className = "tree-mark";
1535
+ mark.textContent = "Β§";
1536
+ a.appendChild(mark);
1537
+ a.appendChild(document.createTextNode(" " + node.title));
1538
+ tree.appendChild(a);
1539
+ });
1540
+ }
1541
+
1542
+ function highlight(slug) {
1543
+ document
1544
+ .querySelectorAll("#tree a")
1545
+ .forEach((a) => a.classList.toggle("active", a.dataset.slug === slug));
1546
+ document
1547
+ .getElementById("book-head")
1548
+ .classList.toggle("active", slug === MANIFEST.root.slug);
1549
+ }
1550
+
1551
+ function clearPageCache() {
1552
+ Object.keys(PAGE_CACHE).forEach((key) => {
1553
+ delete PAGE_CACHE[key];
1554
+ });
1555
+ }
1556
+
1557
+ function isLocalPreview() {
1558
+ return ["localhost", "127.0.0.1", "::1"].includes(location.hostname);
1559
+ }
1560
+
1561
+ async function fetchManifest() {
1562
+ const suffix = isLocalPreview() ? `?t=${Date.now()}` : "";
1563
+ return await (await fetch("./logbook.json" + suffix, { cache: "no-store" })).json();
1564
+ }
1565
+
1566
+ async function fetchPage(node) {
1567
+ if (PAGE_CACHE[node.file]) return PAGE_CACHE[node.file];
1568
+ try {
1569
+ const suffix = isLocalPreview()
1570
+ ? `?rev=${encodeURIComponent(MANIFEST.revision || "")}`
1571
+ : "";
1572
+ const r = await fetch("./" + node.file + suffix, { cache: "no-store" });
1573
+ PAGE_CACHE[node.file] = await r.text();
1574
+ } catch (e) {
1575
+ PAGE_CACHE[node.file] = "# " + node.title + "\n\n_Could not load section._";
1576
+ }
1577
+ return PAGE_CACHE[node.file];
1578
+ }
1579
+
1580
+ function allNodes() {
1581
+ const nodes = [];
1582
+ flattenTree(MANIFEST.root, 0, nodes);
1583
+ return nodes.map(({ node }) => node);
1584
+ }
1585
+
1586
+ function collectPinnedCells(markdown, nodes) {
1587
+ const cells = [];
1588
+ markdown.forEach((text, index) => {
1589
+ const cellRe = /(^|\n)---\n<!-- trackio-cell\n([\s\S]*?)\n-->\n([\s\S]*?)(?=\n---\n<!-- trackio-cell\n|\s*$)/g;
1590
+ let match;
1591
+ let cellIndex = 0;
1592
+ while ((match = cellRe.exec(text))) {
1593
+ const meta = parseCellMeta(match[2]);
1594
+ if (isPinned(meta)) {
1595
+ cells.push({
1596
+ meta,
1597
+ body: match[3],
1598
+ node: nodes[index],
1599
+ index: cells.length,
1600
+ order: meta.pinned_at || meta.created_at || "",
1601
+ cellIndex,
1602
+ });
1603
+ }
1604
+ cellIndex++;
1605
+ }
1606
+ });
1607
+ return cells.sort(
1608
+ (a, b) =>
1609
+ a.order.localeCompare(b.order) ||
1610
+ a.index - b.index ||
1611
+ a.cellIndex - b.cellIndex
1612
+ );
1613
+ }
1614
+
1615
+ function renderPinnedNotes(cells, container) {
1616
+ if (!cells.length) return;
1617
+ const deck = document.createElement("section");
1618
+ deck.className = "pinned-notes";
1619
+ const list = document.createElement("div");
1620
+ list.className = "pinned-notes-list";
1621
+ cells.forEach(({ meta, body }) => {
1622
+ const cell = renderCell(meta, body, list);
1623
+ cell.classList.add("pinned-copy");
1624
+ });
1625
+ deck.appendChild(list);
1626
+ const anchor =
1627
+ container.querySelector(".logbook-stats") ||
1628
+ container.querySelector(".agent-hint");
1629
+ container.insertBefore(deck, anchor ? anchor.nextSibling : container.firstChild);
1630
+ container.closest(".book-intro").classList.add("has-pinned-notes");
1631
+ }
1632
+
1633
+ function removeIndexProse(body) {
1634
+ const h1 = Array.from(body.children).find((el) => el.tagName === "H1");
1635
+ if (!h1) return;
1636
+ let current = h1.nextElementSibling;
1637
+ while (current && current.tagName !== "H2") {
1638
+ const next = current.nextElementSibling;
1639
+ current.remove();
1640
+ current = next;
1641
+ }
1642
+ }
1643
+
1644
+ function removePageDirectory(body) {
1645
+ const heading = Array.from(body.children).find(
1646
+ (el) => el.tagName === "H2" && el.textContent.trim().toLowerCase() === "pages"
1647
+ );
1648
+ if (!heading) return;
1649
+ let current = heading;
1650
+ while (current) {
1651
+ const next = current.nextElementSibling;
1652
+ current.remove();
1653
+ if (next && ["H1", "H2"].includes(next.tagName)) break;
1654
+ current = next;
1655
+ }
1656
+ }
1657
+
1658
+ const RAIL_OBSERVERS = [];
1659
+
1660
+ async function renderLogbook(opts = {}) {
1661
+ const scrollY = window.scrollY;
1662
+ const page = document.getElementById("page");
1663
+ RAIL_OBSERVERS.splice(0).forEach((observer) => observer.disconnect());
1664
+ page.innerHTML = "";
1665
+ const nodes = allNodes();
1666
+ const markdown = await Promise.all(nodes.map(fetchPage));
1667
+ const pinnedCells = collectPinnedCells(markdown, nodes);
1668
+ let bookIntroBody = null;
1669
+ nodes.forEach((node, index) => {
1670
+ const section = document.createElement("section");
1671
+ section.className = "page-section";
1672
+ section.id = "/" + node.slug;
1673
+ section.dataset.slug = node.slug;
1674
+
1675
+ const layout = document.createElement("div");
1676
+ layout.className = "page-layout";
1677
+ const body = document.createElement("div");
1678
+ body.className = "page-body";
1679
+ const rail = document.createElement("aside");
1680
+ rail.className = "context-rail";
1681
+ rail.setAttribute("aria-label", `Resources for ${node.title}`);
1682
+
1683
+ renderMarkdown(markdown[index], body);
1684
+ if (node.slug === MANIFEST.root.slug) {
1685
+ section.classList.add("book-intro");
1686
+ removeIndexProse(body);
1687
+ removePageDirectory(body);
1688
+ const hint = buildAgentHint();
1689
+ const h1 = body.querySelector("h1");
1690
+ if (h1 && h1.parentNode === body) {
1691
+ body.insertBefore(hint, h1.nextSibling);
1692
+ } else {
1693
+ body.prepend(hint);
1694
+ }
1695
+ hint.after(buildLogbookStats(markdown));
1696
+ bookIntroBody = body;
1697
+ }
1698
+ layout.appendChild(body);
1699
+ layout.appendChild(rail);
1700
+ section.appendChild(layout);
1701
+ page.appendChild(section);
1702
+ renderRail(markdown[index], body, rail);
1703
+ if (window.ResizeObserver) {
1704
+ const observer = new ResizeObserver(() => scheduleRailPosition(body, rail));
1705
+ observer.observe(body);
1706
+ observer.observe(rail);
1707
+ RAIL_OBSERVERS.push(observer);
1708
+ }
1709
+ });
1710
+ if (bookIntroBody) renderPinnedNotes(pinnedCells, bookIntroBody);
1711
+ if (bookIntroBody) {
1712
+ const section = bookIntroBody.closest(".book-intro");
1713
+ const hasExtra = Array.from(bookIntroBody.children).some(
1714
+ (el) =>
1715
+ el.tagName !== "H1" &&
1716
+ !el.classList.contains("agent-hint") &&
1717
+ !el.classList.contains("logbook-stats") &&
1718
+ !el.classList.contains("pinned-notes")
1719
+ );
1720
+ if (section && !section.classList.contains("has-pinned-notes") && !hasExtra) {
1721
+ section.classList.add("book-intro-tight");
1722
+ }
1723
+ }
1724
+ requestAnimationFrame(() => {
1725
+ if (opts.preserveScroll) {
1726
+ window.scrollTo(0, scrollY);
1727
+ } else {
1728
+ scrollToHash({ behavior: "auto" });
1729
+ }
1730
+ updateActiveSection();
1731
+ });
1732
+ }
1733
+
1734
+ function setupResourceHover() {
1735
+ document.addEventListener("mouseover", (e) => {
1736
+ const el = e.target.closest && e.target.closest("[data-res-url]");
1737
+ if (!el || el.classList.contains("rail-item")) return;
1738
+ const url = el.getAttribute("data-res-url");
1739
+ const section = el.closest(".page-section");
1740
+ const scope = section || document;
1741
+ scope.querySelectorAll(".context-rail [data-res-url]").forEach((n) => {
1742
+ n.classList.toggle("res-hl", n.getAttribute("data-res-url") === url);
1743
+ });
1744
+ });
1745
+ document.addEventListener("mouseout", (e) => {
1746
+ const el = e.target.closest && e.target.closest("[data-res-url]");
1747
+ if (!el || el.classList.contains("rail-item")) return;
1748
+ document.querySelectorAll(".context-rail .res-hl").forEach((n) => {
1749
+ n.classList.remove("res-hl");
1750
+ });
1751
+ });
1752
+ }
1753
+
1754
+ let STATS_TOKEN = 0;
1755
+ let STATS_LISTENERS = false;
1756
+
1757
+ function fmtBytes(n) {
1758
+ if (n == null || isNaN(n)) return null;
1759
+ if (n < 1000) return `${n} B`;
1760
+ const units = ["kB", "MB", "GB", "TB"];
1761
+ let v = n;
1762
+ let i = -1;
1763
+ do {
1764
+ v /= 1000;
1765
+ i++;
1766
+ } while (v >= 1000 && i < units.length - 1);
1767
+ return `${v.toFixed(v < 10 ? 1 : 0)} ${units[i]}`;
1768
+ }
1769
+
1770
+ function spaceIdFromUrl(url) {
1771
+ return url.split("/spaces/")[1].split(/[?#]/)[0].replace(/\/$/, "");
1772
+ }
1773
+
1774
+ const LB_CELL_RE = /(^|\n)---\n<!-- trackio-cell\n([\s\S]*?)\n-->\n([\s\S]*?)(?=\n---\n<!-- trackio-cell\n|\s*$)/g;
1775
+
1776
+ function cellDashboardItems(md) {
1777
+ const re = new RegExp(LB_CELL_RE.source, "g");
1778
+ const items = [];
1779
+ let m;
1780
+ while ((m = re.exec(md))) {
1781
+ const meta = parseCellMeta(m[2]);
1782
+ if (meta.type !== "dashboard") continue;
1783
+ const body = m[3];
1784
+ const project = meta.dashboard_project || "";
1785
+ const sp = body.match(/https:\/\/huggingface\.co\/spaces\/[^\s<>)"'`]+/);
1786
+ const local = !sp;
1787
+ const url = sp ? sp[0] : "";
1788
+ const resUrl = local ? `trackio-local-dashboard://${project}` : url;
1789
+ items.push({
1790
+ id: local ? project : spaceIdFromUrl(url),
1791
+ local,
1792
+ url,
1793
+ resUrl,
1794
+ });
1795
+ }
1796
+ return items;
1797
+ }
1798
+
1799
+ function artifactInfoFromCell(meta, body) {
1800
+ const name = meta.artifact || meta.path || "";
1801
+ let size = null;
1802
+ const sm = body.match(/Β·\s*([\d.]+\s*[kMGT]?B)\b/);
1803
+ if (sm) size = sm[1].trim();
1804
+ if (!size && meta.size != null) size = fmtBytes(meta.size);
1805
+ const bucket = body.match(/https:\/\/huggingface\.co\/buckets\/[^\s<>)"'`]+/);
1806
+ const artUri = body.match(/trackio-artifact:\/\/\S+/);
1807
+ const pathUri = body.match(/trackio-local-path:\/\/\S+/);
1808
+ const url = bucket ? bucket[0] : "";
1809
+ const local = !bucket;
1810
+ const resUrl =
1811
+ url || (artUri ? artUri[0] : pathUri ? pathUri[0] : `trackio-artifact://${name}`);
1812
+ return {
1813
+ name,
1814
+ type: meta.artifact_type || "",
1815
+ size,
1816
+ local,
1817
+ isPathRef: !!meta.path,
1818
+ url,
1819
+ resUrl,
1820
+ };
1821
+ }
1822
+
1823
+ function cellArtifactItems(md) {
1824
+ const re = new RegExp(LB_CELL_RE.source, "g");
1825
+ const items = [];
1826
+ let m;
1827
+ while ((m = re.exec(md))) {
1828
+ const meta = parseCellMeta(m[2]);
1829
+ const body = m[3];
1830
+ const order = meta.created_at || "";
1831
+ if (meta.type === "artifact") {
1832
+ const info = artifactInfoFromCell(meta, body);
1833
+ if (info.name) items.push({ ...info, order });
1834
+ }
1835
+ }
1836
+ return items;
1837
+ }
1838
+
1839
+ function collectLogbookResources(markdownList) {
1840
+ const re = new RegExp(LB_CELL_RE.source, "g");
1841
+ const dashboards = new Map();
1842
+ markdownList.forEach((md) => {
1843
+ let m;
1844
+ while ((m = re.exec(md))) {
1845
+ const meta = parseCellMeta(m[2]);
1846
+ const body = m[3];
1847
+ if (meta.type !== "dashboard") continue;
1848
+ const project = meta.dashboard_project || "";
1849
+ const space = body.match(/https:\/\/huggingface\.co\/spaces\/[^\s<>)"'`]+/);
1850
+ const local = !space;
1851
+ const url = space ? space[0] : "";
1852
+ const key = local ? `local:${project}` : `space:${spaceIdFromUrl(url)}`;
1853
+ const resUrl = local ? `trackio-local-dashboard://${project}` : url;
1854
+ if (!dashboards.has(key))
1855
+ dashboards.set(key, { project, local, url, resUrl });
1856
+ }
1857
+ });
1858
+ const artifacts = new Map();
1859
+ markdownList.forEach((md) => {
1860
+ cellArtifactItems(md).forEach((it) => {
1861
+ const key = `${it.type}:${it.name}`;
1862
+ const prev = artifacts.get(key);
1863
+ if (!prev || it.order >= prev.order) artifacts.set(key, it);
1864
+ });
1865
+ });
1866
+ return {
1867
+ dashboards: Array.from(dashboards.values()).sort((a, b) =>
1868
+ a.project.localeCompare(b.project)
1869
+ ),
1870
+ artifacts: Array.from(artifacts.values()).sort((a, b) =>
1871
+ a.name.localeCompare(b.name)
1872
+ ),
1873
+ };
1874
+ }
1875
+
1876
+ function closeStatPopovers() {
1877
+ document
1878
+ .querySelectorAll(".stat-popover")
1879
+ .forEach((p) => (p.hidden = true));
1880
+ document
1881
+ .querySelectorAll(".stat-tile.open")
1882
+ .forEach((t) => t.classList.remove("open"));
1883
+ }
1884
+
1885
+ function ensureStatListeners() {
1886
+ if (STATS_LISTENERS) return;
1887
+ STATS_LISTENERS = true;
1888
+ document.addEventListener("click", closeStatPopovers);
1889
+ document.addEventListener("keydown", (e) => {
1890
+ if (e.key === "Escape") closeStatPopovers();
1891
+ });
1892
+ }
1893
+
1894
+ function stateHtml(remote, url) {
1895
+ return remote
1896
+ ? `<a class="stat-row-state open" href="${esc(url)}" target="_blank" rel="noopener" title="Open in a new tab">Open β†—</a>`
1897
+ : `<span class="stat-row-state">publish to share</span>`;
1898
+ }
1899
+
1900
+ function scrollToResource(resUrl) {
1901
+ closeStatPopovers();
1902
+ if (!resUrl) return;
1903
+ const el = document.querySelector(
1904
+ `#page .page-body [data-res-url="${CSS.escape(resUrl)}"]:not(.stat-row)`
1905
+ );
1906
+ if (!el) return;
1907
+ el.scrollIntoView({ behavior: "smooth", block: "center" });
1908
+ el.classList.add("res-flash");
1909
+ setTimeout(() => el.classList.remove("res-flash"), 1500);
1910
+ }
1911
+
1912
+ function dashRowHtml(d) {
1913
+ const inner =
1914
+ `<span class="stat-row-ico">${DASHBOARD_ICON_IMG}</span>` +
1915
+ `<div class="stat-row-main"><div class="stat-row-title">${esc(d.project)}</div>` +
1916
+ `<div class="stat-row-meta">${stateHtml(!d.local, d.url)}</div></div>`;
1917
+ return `<div class="stat-row" data-res-url="${esc(d.resUrl)}" title="Jump to it in the logbook">${inner}</div>`;
1918
+ }
1919
+
1920
+ function artRowHtml(a) {
1921
+ const remote = !a.local && !!a.url;
1922
+ const parts = [a.type, a.size].filter(Boolean).map(esc);
1923
+ const meta = parts.length
1924
+ ? `${parts.join(" Β· ")} Β· ${stateHtml(remote, a.url)}`
1925
+ : stateHtml(remote, a.url);
1926
+ const inner =
1927
+ `<span class="stat-row-ico">${ARTIFACT_ICON_IMG}</span>` +
1928
+ `<div class="stat-row-main"><div class="stat-row-title">${esc(a.name)}</div>` +
1929
+ `<div class="stat-row-meta">${meta}</div></div>`;
1930
+ return `<div class="stat-row" data-res-url="${esc(a.resUrl)}" title="Jump to it in the logbook">${inner}</div>`;
1931
+ }
1932
+
1933
+ function statTile(icon, alt, singular, plural, head, rowFn) {
1934
+ const tile = document.createElement("button");
1935
+ tile.type = "button";
1936
+ tile.className = "stat-tile";
1937
+ const render = (items) => {
1938
+ const count = items.length;
1939
+ const label = count === 1 ? singular : plural;
1940
+ const caret = count > 0 ? `<span class="stat-caret">β–Ύ</span>` : "";
1941
+ tile.innerHTML =
1942
+ `<img class="stat-icon" src="${icon}" alt="${esc(alt)}" />` +
1943
+ `<div class="stat-text"><div class="stat-num">${count}</div>` +
1944
+ `<div class="stat-label">${esc(label)}</div></div>` +
1945
+ caret;
1946
+ tile.disabled = count === 0;
1947
+ if (count > 0) {
1948
+ const pop = document.createElement("div");
1949
+ pop.className = "stat-popover";
1950
+ pop.hidden = true;
1951
+ pop.innerHTML =
1952
+ `<div class="stat-pop-head">${esc(head)}</div>` +
1953
+ items.map(rowFn).join("");
1954
+ pop.addEventListener("click", (e) => {
1955
+ if (e.target.closest("a.stat-row-state")) {
1956
+ e.stopPropagation();
1957
+ return;
1958
+ }
1959
+ e.stopPropagation();
1960
+ const row = e.target.closest(".stat-row");
1961
+ if (row) scrollToResource(row.dataset.resUrl);
1962
+ });
1963
+ tile.appendChild(pop);
1964
+ }
1965
+ };
1966
+ tile.addEventListener("click", (e) => {
1967
+ if (tile.disabled) return;
1968
+ e.stopPropagation();
1969
+ const pop = tile.querySelector(".stat-popover");
1970
+ if (!pop) return;
1971
+ const isOpen = !pop.hidden;
1972
+ closeStatPopovers();
1973
+ if (!isOpen) {
1974
+ pop.hidden = false;
1975
+ tile.classList.add("open");
1976
+ }
1977
+ });
1978
+ return { tile, render };
1979
+ }
1980
+
1981
+ function buildLogbookStats(markdownList) {
1982
+ const token = ++STATS_TOKEN;
1983
+ ensureStatListeners();
1984
+ const { dashboards, artifacts } = collectLogbookResources(markdownList);
1985
+
1986
+ const el = document.createElement("div");
1987
+ el.className = "logbook-stats";
1988
+ const dash = statTile(
1989
+ "./trackio-logo-light.png",
1990
+ "Trackio",
1991
+ "Trackio Dashboard",
1992
+ "Trackio Dashboards",
1993
+ "Dashboards created in this logbook",
1994
+ dashRowHtml
1995
+ );
1996
+ const art = statTile(
1997
+ "./bucket-icon.svg",
1998
+ "Bucket",
1999
+ "Artifact",
2000
+ "Artifacts",
2001
+ "Artifacts created in this logbook",
2002
+ artRowHtml
2003
+ );
2004
+ dash.render(dashboards);
2005
+ art.render(artifacts);
2006
+ el.appendChild(dash.tile);
2007
+ el.appendChild(art.tile);
2008
+
2009
+ const scanText = markdownList
2010
+ .map((md) =>
2011
+ md.replace(/(`{3,4}|~{3,4})(html|raw)[^\n]*\n[\s\S]*?\n\1/g, " ")
2012
+ )
2013
+ .join("\n");
2014
+ const seen = new Set(
2015
+ dashboards.map((d) =>
2016
+ d.local ? `local:${d.project}` : `space:${spaceIdFromUrl(d.url)}`
2017
+ )
2018
+ );
2019
+ const remoteSpaces = new Map();
2020
+ extractUrls(scanText).forEach((url) => {
2021
+ const item = classifyResource(url);
2022
+ if (item && item.kind === "space" && !item.local) {
2023
+ remoteSpaces.set(item.url, item);
2024
+ }
2025
+ });
2026
+ remoteSpaces.forEach((s) => {
2027
+ const key = `space:${s.id}`;
2028
+ if (seen.has(key)) return;
2029
+ getJSON(`https://huggingface.co/api/spaces/${s.id}`)
2030
+ .then((d) => {
2031
+ if (STATS_TOKEN !== token) return;
2032
+ const tags = (d && d.tags) || [];
2033
+ if (
2034
+ !seen.has(key) &&
2035
+ tags.some((t) => String(t).toLowerCase() === "trackio")
2036
+ ) {
2037
+ seen.add(key);
2038
+ dashboards.push({
2039
+ project: s.id,
2040
+ local: false,
2041
+ url: s.url,
2042
+ resUrl: s.url,
2043
+ });
2044
+ dashboards.sort((a, b) => a.project.localeCompare(b.project));
2045
+ dash.render(dashboards);
2046
+ }
2047
+ })
2048
+ .catch(() => {});
2049
+ });
2050
+ return el;
2051
+ }
2052
+
2053
+ function buildAgentHint() {
2054
+ const onSpaces =
2055
+ /\.hf\.space$/.test(location.hostname) ||
2056
+ /(^|\.)huggingface\.co$/.test(location.hostname);
2057
+ let source = "";
2058
+ if (onSpaces && MANIFEST.space_id) {
2059
+ source = ` ${MANIFEST.space_id}`;
2060
+ } else if (/^https?:$/.test(location.protocol)) {
2061
+ source = ` ${location.origin}/`;
2062
+ }
2063
+ const command = `trackio logbook read${source}`;
2064
+ const tokens = MANIFEST.agent_view_tokens;
2065
+ const div = document.createElement("div");
2066
+ div.className = "agent-hint";
2067
+ const label = document.createElement("span");
2068
+ label.className = "agent-hint-label";
2069
+ label.textContent = "Read from the CLI:";
2070
+ const code = document.createElement("code");
2071
+ code.textContent = command;
2072
+ const copy = document.createElement("button");
2073
+ copy.className = "copy";
2074
+ copy.type = "button";
2075
+ copy.title = "Copy";
2076
+ copy.textContent = "⧉";
2077
+ copy.addEventListener("click", () => copyText(command, copy, "⧉"));
2078
+ const note = document.createElement("span");
2079
+ note.className = "agent-hint-note";
2080
+ note.textContent =
2081
+ "compact view for agents" + (tokens ? ` Β· ~${fmt(tokens)} tokens` : "");
2082
+ div.appendChild(label);
2083
+ div.appendChild(code);
2084
+ div.appendChild(copy);
2085
+ div.appendChild(note);
2086
+ return div;
2087
+ }
2088
+
2089
+ function currentSlug() {
2090
+ const slug = (location.hash || "").replace(/^#\//, "") || MANIFEST.root.slug;
2091
+ return findNode(MANIFEST.root, slug) ? slug : MANIFEST.root.slug;
2092
+ }
2093
+
2094
+ function scrollToHash(opts = {}) {
2095
+ const slug = currentSlug();
2096
+ if (!location.hash) {
2097
+ window.scrollTo({ top: 0, behavior: opts.behavior || "auto" });
2098
+ highlight(slug);
2099
+ return;
2100
+ }
2101
+ const section = document.getElementById("/" + slug);
2102
+ if (section) section.scrollIntoView({ behavior: opts.behavior || "smooth" });
2103
+ highlight(slug);
2104
+ }
2105
+
2106
+ function navigateToLogbookSlug(target) {
2107
+ const slug = String(target || "").replace(/^#?\//, "").trim();
2108
+ if (!slug || !findNode(MANIFEST.root, slug)) return;
2109
+ const hash = "#/" + slug;
2110
+ if (location.hash === hash) {
2111
+ scrollToHash({ behavior: "smooth" });
2112
+ } else {
2113
+ location.hash = hash;
2114
+ }
2115
+ }
2116
+
2117
+ function setupFigureNavigation() {
2118
+ window.addEventListener("message", (event) => {
2119
+ const data = event.data;
2120
+ if (!data || data.type !== "trackio-logbook:navigate") return;
2121
+ // Only accept messages from one of this logbook's sandboxed figure
2122
+ // iframes, rather than from an arbitrary same-origin page.
2123
+ const isFigureFrame = Array.from(
2124
+ document.querySelectorAll("iframe.figure-frame")
2125
+ ).some((frame) => frame.contentWindow === event.source);
2126
+ if (!isFigureFrame) return;
2127
+ navigateToLogbookSlug(data.target);
2128
+ });
2129
+ }
2130
+
2131
+ let SCROLL_FRAME = 0;
2132
+ function updateActiveSection() {
2133
+ cancelAnimationFrame(SCROLL_FRAME);
2134
+ SCROLL_FRAME = requestAnimationFrame(() => {
2135
+ const sections = Array.from(document.querySelectorAll(".page-section"));
2136
+ if (!sections.length) return;
2137
+ const marker = Math.min(window.innerHeight * 0.28, 180);
2138
+ let active = sections[0];
2139
+ sections.forEach((section) => {
2140
+ if (section.getBoundingClientRect().top <= marker) active = section;
2141
+ });
2142
+ if (
2143
+ window.innerHeight + window.scrollY >=
2144
+ document.documentElement.scrollHeight - 2
2145
+ ) {
2146
+ active = sections[sections.length - 1];
2147
+ }
2148
+ highlight(active.dataset.slug);
2149
+ });
2150
+ }
2151
+
2152
+ function startLiveReload() {
2153
+ if (!isLocalPreview()) return;
2154
+ setInterval(async () => {
2155
+ try {
2156
+ const next = await fetchManifest();
2157
+ if (!next || next.revision === MANIFEST.revision) return;
2158
+ MANIFEST = next;
2159
+ clearPageCache();
2160
+ document.title = MANIFEST.title + " Β· Trackio Logbook";
2161
+ document.getElementById("book-title").textContent = MANIFEST.title;
2162
+ document.getElementById("book-head").setAttribute("aria-label", MANIFEST.title);
2163
+ buildTree();
2164
+ renderLogbook({ preserveScroll: true });
2165
+ } catch (e) {}
2166
+ }, LIVE_RELOAD_MS);
2167
+ }
2168
+
2169
+ function setupConnect() {
2170
+ const space = MANIFEST.space_id;
2171
+ if (!space) return;
2172
+ const steps = [
2173
+ { t: "Install Trackio, if you don't have it yet.", c: "uv tool install trackio" },
2174
+ { t: "Add the Trackio skill for your agent, then reload it.", c: "trackio skills add" },
2175
+ { t: "Connect to this logbook.", c: `trackio logbook open ${space}` },
2176
+ ];
2177
+ const ol = document.getElementById("connect-steps");
2178
+ steps.forEach((s, i) => {
2179
+ const li = document.createElement("li");
2180
+ const title = document.createElement("div");
2181
+ title.className = "step-title";
2182
+ title.textContent = `${i + 1}. ${s.t}`;
2183
+ const block = document.createElement("div");
2184
+ block.className = "codeblock";
2185
+ const code = document.createElement("code");
2186
+ code.textContent = s.c;
2187
+ const copy = document.createElement("button");
2188
+ copy.className = "copy";
2189
+ copy.type = "button";
2190
+ copy.title = "Copy";
2191
+ copy.textContent = "⧉";
2192
+ copy.addEventListener("click", () => copyText(s.c, copy, "⧉"));
2193
+ block.appendChild(code);
2194
+ block.appendChild(copy);
2195
+ li.appendChild(title);
2196
+ li.appendChild(block);
2197
+ ol.appendChild(li);
2198
+ });
2199
+
2200
+ const agentPrompt =
2201
+ `Read and help maintain this Trackio experiment logbook ("${MANIFEST.title}").\n\n` +
2202
+ "1. If you don't have Trackio, install it: uv tool install trackio\n" +
2203
+ "2. Add the Trackio skill for your agent: trackio skills add (then reload)\n" +
2204
+ `3. Connect to this logbook: trackio logbook open ${space}\n\n` +
2205
+ "Start with `trackio logbook read`; use `trackio logbook read page \"...\"` " +
2206
+ "for a page-level view, then fetch relevant details with " +
2207
+ "`trackio logbook read cell cell_<id>`. If I've given you " +
2208
+ 'write access to the Space, add findings with `trackio logbook cell markdown "..." ' +
2209
+ '--page "..."` and they will sync back automatically.';
2210
+
2211
+ const foot = document.getElementById("sidebar-foot");
2212
+ foot.hidden = false;
2213
+ const modal = document.getElementById("modal");
2214
+ const open = () => (modal.hidden = false);
2215
+ const close = () => (modal.hidden = true);
2216
+ document.getElementById("connect-btn").addEventListener("click", open);
2217
+ document.getElementById("modal-close").addEventListener("click", close);
2218
+ modal.querySelector(".modal-backdrop").addEventListener("click", close);
2219
+ document.addEventListener("keydown", (e) => {
2220
+ if (e.key === "Escape") close();
2221
+ });
2222
+ const agentBtn = document.getElementById("copy-agent");
2223
+ agentBtn.addEventListener("click", () =>
2224
+ copyText(agentPrompt, agentBtn, "Copy for agent")
2225
+ );
2226
+ }
2227
+
2228
+ function copyText(text, btn, restore) {
2229
+ const done = () => {
2230
+ const prev = btn.textContent;
2231
+ btn.textContent = restore === "⧉" ? "βœ“" : "Copied!";
2232
+ btn.classList.add("copied");
2233
+ setTimeout(() => {
2234
+ btn.textContent = restore;
2235
+ btn.classList.remove("copied");
2236
+ }, 1400);
2237
+ void prev;
2238
+ };
2239
+ if (navigator.clipboard && navigator.clipboard.writeText) {
2240
+ navigator.clipboard.writeText(text).then(done, done);
2241
+ } else {
2242
+ const ta = document.createElement("textarea");
2243
+ ta.value = text;
2244
+ document.body.appendChild(ta);
2245
+ ta.select();
2246
+ try {
2247
+ document.execCommand("copy");
2248
+ } catch (e) {}
2249
+ document.body.removeChild(ta);
2250
+ done();
2251
+ }
2252
+ }
2253
+
2254
+ async function init() {
2255
+ MANIFEST = await fetchManifest();
2256
+ document.title = MANIFEST.title + " Β· Trackio Logbook";
2257
+ document.getElementById("book-title").textContent = MANIFEST.title;
2258
+ document.getElementById("book-head").setAttribute("aria-label", MANIFEST.title);
2259
+ document.getElementById("book-head").addEventListener("click", () => {
2260
+ const target = "#/" + MANIFEST.root.slug;
2261
+ if (location.hash === target) scrollToHash();
2262
+ else location.hash = target;
2263
+ });
2264
+ buildTree();
2265
+ setupConnect();
2266
+ setupResourceHover();
2267
+ setupFigureNavigation();
2268
+ window.addEventListener("hashchange", () => scrollToHash());
2269
+ window.addEventListener("scroll", updateActiveSection, { passive: true });
2270
+ await renderLogbook();
2271
+ startLiveReload();
2272
+ }
2273
+
2274
+ init();
2275
+ })();
logbook.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "title": "Repro - Benchmarking at the Edge of Comprehension",
4
+ "emoji": "\ud83d\udd2c",
5
+ "space_id": "bertfil/BhahZSDowo",
6
+ "paper": {
7
+ "arxiv_id": "2602.14307",
8
+ "openreview_id": "BhahZSDowo",
9
+ "title": "Benchmarking at the Edge of Comprehension",
10
+ "url": "https://arxiv.org/abs/2602.14307"
11
+ },
12
+ "tags": [
13
+ "icml2026-repro",
14
+ "paper-BhahZSDowo"
15
+ ],
16
+ "updated_at": "2026-07-18T16:28:34+00:00",
17
+ "root": {
18
+ "slug": "index",
19
+ "title": "Repro - Benchmarking at the Edge of Comprehension",
20
+ "file": "pages/index.md",
21
+ "children": [
22
+ {
23
+ "slug": "00-scored-evidence-summary",
24
+ "title": "00 - Scored evidence summary",
25
+ "file": "pages/00-scored-evidence-summary/page.md",
26
+ "children": []
27
+ },
28
+ {
29
+ "slug": "claim-1-critique-resilient-benchmarking-framework-produces",
30
+ "title": "Claim 1 - Critique-Resilient Benchmarking framework produces stable sc\u2026",
31
+ "file": "pages/claim-1-critique-resilient-benchmarking-framework-produces/page.md",
32
+ "children": []
33
+ },
34
+ {
35
+ "slug": "claim-2-reformulates-benchmarking-as-adversarial-generatio",
36
+ "title": "Claim 2 - Reformulates benchmarking as adversarial generation-evaluati\u2026",
37
+ "file": "pages/claim-2-reformulates-benchmarking-as-adversarial-generatio/page.md",
38
+ "children": []
39
+ },
40
+ {
41
+ "slug": "reproduction-protocol-and-provenance",
42
+ "title": "Reproduction protocol and provenance",
43
+ "file": "pages/reproduction-protocol-and-provenance/page.md",
44
+ "children": []
45
+ },
46
+ {
47
+ "slug": "conclusion",
48
+ "title": "Conclusion",
49
+ "file": "pages/conclusion/page.md",
50
+ "children": []
51
+ }
52
+ ]
53
+ },
54
+ "agent_view_tokens": 15724,
55
+ "revision": "1784392114542018200"
56
+ }
pages/00-scored-evidence-summary/page.md ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 00 - Scored evidence summary
2
+
3
+
4
+ ---
5
+ <!-- trackio-cell
6
+ {"type": "markdown", "id": "cell_decad5c995eb", "created_at": "2026-07-18T16:28:34+00:00", "title": "Scorecard: 1 verified / 0 falsified / 1 toy / 0 inconclusive of 2 board claims", "pinned": true, "pinned_at": "2026-07-18T16:28:34+00:00"}
7
+ -->
8
+ # Scored evidence β€” 1 verified / 0 falsified / 1 toy / 0 inconclusive of 2 board claims
9
+
10
+ **Paper:** Benchmarking at the Edge of Comprehension
11
+ **IDs:** OpenReview `BhahZSDowo` Β· arXiv `2602.14307`
12
+ **Required tags:** `icml2026-repro`, `paper-BhahZSDowo`
13
+ **Method:** [veritas](https://arxiv.org/abs/2607.02931) replication agent, full mode (paper + authors' code); see the provenance page.
14
+
15
+ | # | Board claim | Verdict | Basis (per-claim replication statuses) |
16
+ |---:|---|---|---|
17
+ | 1 | Critique-Resilient Benchmarking framework produces stable scores that correlate with external capability measures | **TOY** | 3 match / 4 partial / 0 no_match / 1 not-executed (not_attempted); C1, C2, C3, C14, C16, C17, C20, C21 |
18
+ | 2 | Reformulates benchmarking as adversarial generation-evaluation game enabling evaluation beyond full human comprehension | **VERIFIED** | 8 match / 3 partial / 0 no_match / 3 not-executed (missing, not_attempted); C3, C4, C5, C6, C7, C8, C9, C10, C11, C12, C13, C15, C16, C19 |
pages/claim-1-critique-resilient-benchmarking-framework-produces/page.md ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Claim 1: Critique-Resilient Benchmarking framework produces stable scores that correlate …
2
+
3
+
4
+ ---
5
+ <!-- trackio-cell
6
+ {"type": "markdown", "id": "cell_8f8f8ac08ff5", "created_at": "2026-07-18T16:28:34+00:00", "title": "Claim 1 - verdict and summary"}
7
+ -->
8
+ **Board claim 1:** Critique-Resilient Benchmarking framework produces stable scores that correlate with external capability measures
9
+
10
+ **Stated verdict: TOY** (3 match / 4 partial / 0 no_match / 1 not-executed (not_attempted))
11
+
12
+ Every tested component of this claim reproduced in the same direction: answerer Elo scores correlated strongly and monotonically with external AIME/BRUMO/HMMT benchmarks (C1, C17, C20), the fitted Bradley-Terry model beat a baserate predictor with well-calibrated held-out probabilities (C2, C21), and rankings stayed stable whether adjudicated by humans or even weak LLM adjudicators (C3, C16). However, most of these results are marked partial rather than full matches, and the additional stability check across prompt variants (C14) could not be run from the shipped repository.
13
+
14
+
15
+ ---
16
+ <!-- trackio-cell
17
+ {"type": "markdown", "id": "cell_bb7c9abe1a92", "created_at": "2026-07-18T16:28:34+00:00", "title": "Evidence from the replication run"}
18
+ -->
19
+ **C1** [headline/table] β€” status **partial**
20
+
21
+ *Claim:* Answerer Elo scores from the itemized bipartite Bradley-Terry fit correlate strongly and monotonically with published scores on three external human-designed mathematics benchmarks (AIME 2025, BRUMO 2025, HMMT Feb 2025).
22
+
23
+ *Verifier rationale:* [deterministic grade] 18 cell(s), 0 outside tolerance β†’ partial. Comparator notes: Step 8 of the replication produced results/external_validity.json (driver run_external_validity.py, stdout mirrored in results/external_validity_stdout.txt), which correlates the step-5 bootstrapped answerer Elo point estimates against the three external benchmarks over n=8 models. All six coefficients land within the claim's stated tolerances: Spearman rho = 0.881 / 0.833 / 0.826 vs the paper's 0.851 / 0.830 / 0.819 (max deviation 0.030, well inside the +/-0.10 band), and Kendall tau = 0.714 / 0.714 / 0.691 vs 0.706 / 0.673 / 0.662 (max deviation 0.041, inside the +/-0.12 band). The decisive qualitative requirements all hold: every rho > 0.7, every tau > 0.55, and all six bootstrap 95% CIs are strictly positive and exclude zero (the lowest bound is 0.473 for HMMT Kendall tau). The reported CIs also track the paper's closely -- AIME's Spearman CI [0.738, 0.952] is identical and HMMT's [0.707, 0.935] matches to rounding. The method inspected in run_external_validity.py is sound: scipy spearmanr/kendalltau on the point Elo vector, with percentile CIs taken over 1000 step-5 cluster-bootstrap replicate Elo vectors while treating benchmark scores as fixed, which propagates Elo uncertainty coherently. One important caveat, recorded rather than hidden: assess/fix_severity.json rates Fix 9 (external_benchmarks.json) as CRITICAL because the repository ships no AIME/BRUMO/HMMT scores and no code to fetch them, so the agent transcribed the eight models' scores verbatim from the paper's own Table 7 with per-entry provenance. The y-variable is therefore taken from the paper rather than independently recomputed with Matharena; the correlation step is reproduced conditional on the paper's reported inputs, and the Elo side (the x-variable) is genuinely independently fit. This limits how much the run independently corroborates external validity, but it does not undermine the numeric agreement, since these values are inputs to the correlation and not targets of it. Verdict: match.
24
+
25
+ ---
26
+
27
+ **C2** [headline/table] β€” status **partial**
28
+
29
+ *Claim:* In a 5-fold predictive-validity experiment, the fitted Bradley-Terry model predicts held-out interaction outcomes substantially better than a baserate predictor, achieving higher accuracy and significantly lower log-loss and Brier score on the test folds.
30
+
31
+ *Verifier rationale:* [deterministic grade] 9 cell(s), 0 outside tolerance β†’ partial. Comparator notes: The replication produced the full Table 2 grid in results/predictive_validity.json (confirmed identically in results/predictive_validity_stdout.txt and in step 6 of replication_log.json), from a 5-fold CV over the same 1868 eligible binary outcomes the shipped collector resolves, with seed=0. The claim's qualitative core is fully reproduced: the fitted Bradley-Terry model beats the baserate predictor decisively and in the same direction on all three metrics, and every one of the claim's own stated acceptance thresholds is met (model accuracy 0.927 is within 0.5 pp of the paper's 0.922; log-loss 0.185 is well below the ~0.30 bar; Brier 0.054 is well below the ~0.10 bar). Seven of the nine cells fall within 5% relative error, including all three baserate cells and all three deltas. Two cells miss the 5% per-cell band: model log-loss 0.185 vs the paper's 0.208 (11% low) and model Brier 0.054 vs 0.061 (12% low) β€” both errors are in the direction of the replicated model performing slightly better than the paper reports, so they strengthen rather than undercut the claim's substance, but they exceed the table tolerance and drive the partial verdict rather than match. One methodological deviation is worth flagging: the claim's verification instructions state that hyperparameters should be re-estimated on each training split, whereas run_predictive_validity.py:63 defaults to --sigmas-from results/itemized_bt.json and reuses the sigmas selected by empirical Bayes on the full data (sigma_alpha=4.963, sigma_beta=4.237, sigma_delta=1.0). That is a mild hyperparameter leak from the full dataset into every fold and is a plausible contributor to the model's optimistically low log-loss and Brier β€” the two cells that miss. Context from fix_severity.json matters here: this driver is new code (the repo implements no CV at all β€” zero hits for kfold/brier/log_loss/baserate), and it sits downstream of Fix 5, the critical repair to analysis/bt_utils.py, which held the paper's Eq. (5) model as unrunnable, silently-divergent dead code. That fix's verification is strong (the repaired collector reproduces the shipped collect_games exactly β€” an identical 1868-game multiset over (answerer, questioner, outcome) plus an identical skip Counter), so the fit rests on the authors' own resolution rules rather than the agent's reconstruction of them, and I do not treat it as undermining these numbers. The log also records an honest limitation that does not affect the verdict: unseen-item test edges were 0 in every fold, so the documented prior-mean fallback was implemented but never exercised.
32
+
33
+ ---
34
+
35
+ **C3** [headline/table] β€” status **partial**
36
+
37
+ *Claim:* Answerer and benchmarker rankings computed from LLM-generated adjudications are nearly identical to rankings computed from human adjudications (Spearman rho in [0.98, 1], Kendall tau in [0.93, 1]), including when weak models such as GPT-3.5 and GPT-4o serve as adjudicators.
38
+
39
+ *Verifier rationale:* [deterministic grade] 32 cell(s), 0 outside tolerance β†’ partial. Comparator notes: Step 7 of the replication log ran a newly written driver, run_adjudicator_agreement.py, that refits the itemized Bradley-Terry model with each of the eight models installed as the sole adjudicator on the 105 claims its synthetic_adjudications/ file covers (panel retained elsewhere), then correlates the resulting benchmarker (alpha) and answerer (beta) Elo vectors against a baseline. All values are read from replication/codebase/results/adjudicator_agreement.json::ranking_consistency, corroborated by results/adjudicator_stdout.txt. Qualitatively the claim's headline reproduces cleanly: every one of the eight adjudicators yields rho >= 0.976 and tau >= 0.929 for BOTH roles, and the weak adjudicators are not outliers -- GPT-3.5 Turbo (rho 0.976/0.976) and GPT-4o (benchmarker 0.976, answerer 1.0) sit exactly in the same band as GPT-5.2 and Claude 4.5 Opus, far above the ~0.95 floor the verification instructions set as the falsification threshold. Per-cell, however, the replication does not match the paper's table pattern. The paper reports benchmarker = 1.0/1.0 for all eight models and answerer = 1.0/1.0 for three of them; the replication instead finds benchmarker = 0.976/0.929 uniformly for all eight, and answerer = 1.0/1.0 for five (GPT-4o, Gemini, Llama, Phi-4, DeepSeek) and 0.976/0.929 for the other three -- i.e. the one-rank-swap is located in the benchmarker ranking rather than the answerer ranking, roughly transposing the paper's pattern. All 16 Spearman cells fall within 5% relative tolerance (1.0 vs 0.976 is 2.4% error); all 16 Kendall cells fall outside it, but only because a single adjacent-rank swap among 8 items moves tau by 0.071 (7.1%), which mechanically exceeds a 5% band on a coarsely discrete statistic -- this is a tolerance artifact, not evidence of a substantively different result. One methodological deviation is material and I could not eliminate it: the claim specifies correlating each model's Elos against Elos fit from HUMAN adjudications, but run_adjudicator_agreement.py:210 uses --baseline results/itemized_bt.json, the shipped judge-panel Elo fit, as the reference. The agent's own step-7 note and the script's header document why: collect_itemized_games resolves every game outcome from --automated-dir via resolve_automated_victory, while --human-evals-dir feeds only the self-answer gate (bt_utils.py:112,119) and never overrides a critique outcome, so a human-adjudicated Elo fit is not constructible from the shipped code path. The agent caught its own first attempt at this (substituting into --human-evals-dir produced a no-op rho=1.0000 for all eight adjudicators) and rewrote it, which is a credit to the run but leaves the reported quantity as model-adjudicator-vs-panel rather than model-adjudicator-vs-human. Separately, human-vs-model agreement IS computed and rests on only n=71 commonly adjudicated claims. assess/fix_severity.json lists no critical fixes; the fixes touching this path (missing collect_invalid_self_answer_questions in utils.py, pipeline_coverage.py call sites) are rated 'major' and reflect a published snapshot that was never smoke-tested, but none were flagged as invalidating the BT fit itself. Verdict: partial -- the scientific conclusion (adjudication is robust to a weak adjudicator) reproduces, while half the individual cells miss the declared tolerance and the baseline is the panel rather than humans.
40
+
41
+ ---
42
+
43
+ **C14** [supporting/scalar] β€” status **not_attempted**
44
+
45
+ *Claim:* Answerer rankings are highly stable across three prompt variants (standard, no guidance document, no self-check knowledge), with Kendall W = 0.952, mean pairwise Spearman rho = 0.929, and mean pairwise Kendall tau = 0.873.
46
+
47
+ *Verifier rationale:* [deterministic grade] comparator reported no replicated value was produced. Comparator notes: This run produced no prompt-ablation output, so there is no Kendall W, mean pairwise Spearman rho, or mean pairwise Kendall tau to compare against the paper's [0.952, 0.929, 0.873], and no Table 12 pairwise breakdown. Plan step 9 (replication/replication_log.json) did not attempt the ablation; it instead documented C14 as NOT_REPRODUCIBLE from the shipped repository, and I independently confirmed each load-bearing fact rather than relying on that report. (1) No variant axis exists in the generation path: prompt_library.py:5-30 reads fixed guidance/*.md paths (load_question_guidance, load_answer_guidance, load_critique_guidance, load_self_critique_guidance, ...) with no way to drop the guidance document or the self-check knowledge, and generate_benchmark.py's argparse block (lines 238-257) exposes no guidance/variant/ablation flag. (2) No variant dimension exists in the stored data: benchmarks/ entries carry only ['generation_rounds', 'outer_attempt', 'run_id', 'status', 'topic_name', 'topic_slug'], one file per question-author model, so the standard/no-guidance/no-self-check conditions cannot be reconstructed post hoc. (3) No concordance statistic is computed anywhere for prompt variants: grep for 'kendall|spearman' over *.py returns hits only in run_adjudicator_agreement.py, run_external_validity.py, and make_gap_report.py -- all replication drivers computing adjudicator agreement and external-benchmark correlation, which are unrelated analyses; no Kendall's W implementation exists at all, and results/ contains no ablation artifact. Reproducing C14 would require re-running benchmark generation against paid providers under three modified prompt conditions plus writing the ablation statistics from scratch, which this run correctly declined to do. The fixes in assess/fix_severity.json (the major utils.py collect_invalid_self_answer_questions restoration, the pydantic/scipy dependency installs) touch the Bradley-Terry analysis path and do not bear on this claim. Status is not_attempted rather than not_applicable because the ablation is checkable in principle -- it is a re-run cost barrier, not a claim that a replication cannot address.
48
+
49
+ ---
50
+
51
+ **C16** [supporting/figure] β€” status **match**
52
+
53
+ *Claim:* All models, including weak ones such as GPT-3.5 Turbo and GPT-4o, show consistently high agreement with the human adjudicator when adjudicating outcomes (human-model agreement rates of 86.9-94.6% using outcome categories).
54
+
55
+ *Verifier rationale:* The produced figure figures/figure3_agreement.pdf is the Figure 3 equivalent: a heatmap of pairwise adjudicator agreement on outcome categories among the eight model adjudicators plus a Human row/column, color-coded by agreement percentage with per-cell n values. The underlying data (results/adjudicator_agreement.json, human_model_agreement) confirms the Human row is uniformly high and every human-model rate exceeds 80%: GPT-5.2 90.1%, Claude Opus 4.5 94.4%, Gemini 3 Pro 93.0%, GPT-4o 93.0%, GPT-3.5 Turbo 93.0%, Llama 4 Maverick 93.0%, Phi-4 93.0%, DeepSeek V3.2 Speciale 88.7% (range 88.7-94.4%, vs paper's stated 86.9-94.6%). Critically, the weak adjudicators GPT-3.5 Turbo and GPT-4o both agree with humans at 93.0%, comparable to frontier models and not systematic outliers β€” exactly the claimed qualitative pattern. Individual per-model values differ by at most ~2 pts from the paper's reported figures (e.g. DeepSeek 88.7 vs 86.9, Phi-4 93.0 vs 94.6), well within replication noise given n=71 shared human-adjudicated claims. Figure 9 (suboutcome granularity) was also produced. The only structural difference from the described figure is that the matrix is rendered full-symmetric rather than strictly lower-triangular, which does not affect the content. The claim is supported.
56
+
57
+ ---
58
+
59
+ **C17** [supporting/table] β€” status **match**
60
+
61
+ *Claim:* External benchmark scores recomputed with Matharena (with confidence intervals) for the eight evaluated models on AIME 2025, BRUMO 2025, and HMMT Feb 2025.
62
+
63
+ *Verifier rationale:* [deterministic grade] 24 cell(s), 0 outside tolerance β†’ match. Comparator notes: IMPORTANT CAVEAT ON PROVENANCE: this table was NOT independently recomputed by the replication. The repository ships no external benchmark scores in any form and no code to fetch them; a search for 'aime|brumo|hmmt|matharena' matches only substring collisions and an incidental prompt (verified exhaustively; fix_severity.json Fix 9 rates this a critical completeness gap). Recomputing with Matharena would require fresh paid API calls to all eight models. The agent therefore transcribed the eight-model x three-benchmark scores verbatim from the paper's own Table 7 into external_benchmarks.json (with a per-entry `source` provenance field), and they are echoed back in results/external_validity.json. Consequently every cell of the replicated table equals the paper value exactly (all 24 score cells and their +/- CI half-widths are identical copies), so per-cell tolerance is trivially satisfied and the qualitative separation the verification asks for holds by construction: the four frontier models (GPT-5.2 94.9/90.8/98.3, Claude 4.5 Opus 89.2/92.5/81.7, Gemini 3 Pro 95.0/96.7/97.5, DeepSeek v3.2 Speciale 95.8/95.0/97.5) all exceed ~80% on all three benchmarks, while GPT-3.5 Turbo (5.8/16.7/1.7), GPT-4o (7.5/15.8/0.8), Phi-4 (20.0/23.3/3.3) and Llama 4 Maverick (16.7/27.5/10.0) all score below ~30%. I mark this `match` because the extracted table equals the paper table cell-for-cell, but the match is tautological rather than an independent corroboration of the Matharena scores; it confirms the scores are correctly carried as inputs into the downstream C1 correlation (which the pipeline reproduced to rho 0.826-0.881, close to the paper's), and it means a C17 mismatch could not be the cause of any C1 discrepancy. Evidence: external_benchmarks.json, results/external_validity.json, replication_log.json step 8, assess/fix_severity.json Fix 9.
64
+
65
+ ---
66
+
67
+ **C20** [supporting/figure] β€” status **partial**
68
+
69
+ *Claim:* Scatter plots of answerer Elo against external benchmark scores show an approximately monotonic relationship for all three benchmarks.
70
+
71
+ *Verifier rationale:* The exact expected file figures/scatter_external.pdf was not produced, but the claim's constituent panels were emitted as three separate per-benchmark PDFs (scatter_elo_vs_aime2025.pdf, scatter_elo_vs_brumo2025.pdf, scatter_elo_vs_hmmt_feb2025.pdf), which counts as produced. Reading all three (multimodal Read) confirms the central scientific content of the claim: each panel plots external benchmark Score (%) on the y-axis against Answerer Elo (itemized BT, beta) on the x-axis, with one point per model for all eight models, and the relationship is clearly approximately monotone increasing in every panel -- the four frontier models cluster in the upper right (high Elo ~1500-1700, high score ~80-100%), the weak models cluster in the lower left (low Elo, low score), with a distinct gap between the groups and no strong non-monotonicity. This matches the described visual pattern well. However, three described structural sub-features are absent: (1) the combined figure was never assembled (only the panels), (2) points are not color-coded per model (all rendered a uniform blue, though a name legend is present), and (3) the 95% CI error bars on both axes are not drawn. fix_severity.json shows this figure comes from a new agent-authored driver (Fix 10, minor) fed by paper-transcribed external scores (Fix 9, critical -- the artifact ships no external benchmark values), so the y-values are conditional on the paper's own Table 7 rather than independently recomputed; this does not change the visual monotonicity assessment. Because the substantive monotonic pattern reproduces cleanly but the combined assembly, per-model color coding, and error bars are missing, this is a partial match.
72
+
73
+ ---
74
+
75
+ **C21** [supporting/figure] β€” status **match**
76
+
77
+ *Claim:* Calibration curves for the predictive-validity experiment show that predicted win probabilities track empirical frequencies on both train and test splits.
78
+
79
+ *Verifier rationale:* The pipeline produced figures/calibration_curves.pdf (the exact expected file figures/calibration.pdf does not exist, but this is the same figure under a related name). The figure plots Empirical frequency (y, 0-1) against Predicted P(answerer wins) (x, 0-1) with a dashed y=x 'perfect calibration' diagonal, exactly as the claim describes. It shows both a train curve and a test (pooled held-out) curve, both of which track the diagonal reasonably closely across the full probability range. The test curve shows more deviation than the train curve, including a visible dip around predicted p ~0.5 (empirical ~0.46), consistent with the paper's qualitative description of a test-panel dip. The qualitative requirement - that the model is approximately calibrated rather than systematically over- or under-confident on both splits - is clearly satisfied. The one structural deviation is layout: the paper describes two side-by-side panels (Train and Test) while the reproduction overlays both curves in a single panel; this is a cosmetic difference that preserves all information content and the claim's message. The fix_severity.json fixes concern the compute_bt_elo entry point rather than the calibration plotting path and do not undermine this figure.
pages/claim-2-reformulates-benchmarking-as-adversarial-generatio/page.md ADDED
@@ -0,0 +1,137 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Claim 2: Reformulates benchmarking as adversarial generation-evaluation game enabling eva…
2
+
3
+
4
+ ---
5
+ <!-- trackio-cell
6
+ {"type": "markdown", "id": "cell_d7be8035c882", "created_at": "2026-07-18T16:28:34+00:00", "title": "Claim 2 - verdict and summary"}
7
+ -->
8
+ **Board claim 2:** Reformulates benchmarking as adversarial generation-evaluation game enabling evaluation beyond full human comprehension
9
+
10
+ **Stated verdict: VERIFIED** (8 match / 3 partial / 0 no_match / 3 not-executed (missing, not_attempted))
11
+
12
+ The adversarial generation-evaluation game mechanics themselves reproduced well: the question-generation and feasibility-gating pipeline counts matched (C7, C9, C10, C11), answerer and benchmarker Elo outputs largely matched published tiers and rankings (C4, C5, C6, C19), and automated or even weak-model adjudicators agreed closely with human adjudicators (C3, C16), though the adjudication-volume breakdown was only partially reproduced (C8). The specific evidence for evaluation 'beyond full human comprehension'β€”the critique-vs-answer-correctness gap (C12) and the mutual-incompatibility check on ambiguous answers (C13)β€”could not be attempted from the released repository, and the benchmarker/answerer skill-correlation result (C15) is missing entirely.
13
+
14
+
15
+ ---
16
+ <!-- trackio-cell
17
+ {"type": "markdown", "id": "cell_3f89ab5f9ab5", "created_at": "2026-07-18T16:28:34+00:00", "title": "Evidence from the replication run"}
18
+ -->
19
+ **C3** [headline/table] β€” status **partial**
20
+
21
+ *Claim:* Answerer and benchmarker rankings computed from LLM-generated adjudications are nearly identical to rankings computed from human adjudications (Spearman rho in [0.98, 1], Kendall tau in [0.93, 1]), including when weak models such as GPT-3.5 and GPT-4o serve as adjudicators.
22
+
23
+ *Verifier rationale:* [deterministic grade] 32 cell(s), 0 outside tolerance β†’ partial. Comparator notes: Step 7 of the replication log ran a newly written driver, run_adjudicator_agreement.py, that refits the itemized Bradley-Terry model with each of the eight models installed as the sole adjudicator on the 105 claims its synthetic_adjudications/ file covers (panel retained elsewhere), then correlates the resulting benchmarker (alpha) and answerer (beta) Elo vectors against a baseline. All values are read from replication/codebase/results/adjudicator_agreement.json::ranking_consistency, corroborated by results/adjudicator_stdout.txt. Qualitatively the claim's headline reproduces cleanly: every one of the eight adjudicators yields rho >= 0.976 and tau >= 0.929 for BOTH roles, and the weak adjudicators are not outliers -- GPT-3.5 Turbo (rho 0.976/0.976) and GPT-4o (benchmarker 0.976, answerer 1.0) sit exactly in the same band as GPT-5.2 and Claude 4.5 Opus, far above the ~0.95 floor the verification instructions set as the falsification threshold. Per-cell, however, the replication does not match the paper's table pattern. The paper reports benchmarker = 1.0/1.0 for all eight models and answerer = 1.0/1.0 for three of them; the replication instead finds benchmarker = 0.976/0.929 uniformly for all eight, and answerer = 1.0/1.0 for five (GPT-4o, Gemini, Llama, Phi-4, DeepSeek) and 0.976/0.929 for the other three -- i.e. the one-rank-swap is located in the benchmarker ranking rather than the answerer ranking, roughly transposing the paper's pattern. All 16 Spearman cells fall within 5% relative tolerance (1.0 vs 0.976 is 2.4% error); all 16 Kendall cells fall outside it, but only because a single adjacent-rank swap among 8 items moves tau by 0.071 (7.1%), which mechanically exceeds a 5% band on a coarsely discrete statistic -- this is a tolerance artifact, not evidence of a substantively different result. One methodological deviation is material and I could not eliminate it: the claim specifies correlating each model's Elos against Elos fit from HUMAN adjudications, but run_adjudicator_agreement.py:210 uses --baseline results/itemized_bt.json, the shipped judge-panel Elo fit, as the reference. The agent's own step-7 note and the script's header document why: collect_itemized_games resolves every game outcome from --automated-dir via resolve_automated_victory, while --human-evals-dir feeds only the self-answer gate (bt_utils.py:112,119) and never overrides a critique outcome, so a human-adjudicated Elo fit is not constructible from the shipped code path. The agent caught its own first attempt at this (substituting into --human-evals-dir produced a no-op rho=1.0000 for all eight adjudicators) and rewrote it, which is a credit to the run but leaves the reported quantity as model-adjudicator-vs-panel rather than model-adjudicator-vs-human. Separately, human-vs-model agreement IS computed and rests on only n=71 commonly adjudicated claims. assess/fix_severity.json lists no critical fixes; the fixes touching this path (missing collect_invalid_self_answer_questions in utils.py, pipeline_coverage.py call sites) are rated 'major' and reflect a published snapshot that was never smoke-tested, but none were flagged as invalidating the BT fit itself. Verdict: partial -- the scientific conclusion (adjudication is robust to a weak adjudicator) reproduces, while half the individual cells miss the declared tolerance and the baseline is the panel rather than humans.
24
+
25
+ ---
26
+
27
+ **C4** [supporting/qualitative] β€” status **match**
28
+
29
+ *Claim:* Answerer Elos separate models into a high-performing frontier cluster (DeepSeek v3.2 Speciale, Claude 4.5 Opus, GPT-5.2, Gemini 3 Pro Preview), a middle group (Llama 4 Maverick, Phi-4, GPT-4o), and GPT-3.5 Turbo as the weakest model.
30
+
31
+ *Verifier rationale:* Every component of the claim is unambiguously demonstrated by the produced evidence. results/itemized_bt.json and results/itemized_bt_stdout.txt give the answerer Elos from the paper's Eq. (5) model: the four frontier models take the top four slots in exactly the paper's order, Llama/Phi-4/GPT-4o form a distinctly lower middle band, and GPT-3.5 Turbo is last. The frontier-to-middle gap is 789 Elo on point estimates and 841 on bootstrap means, matching the paper's 'roughly 800+'. results/bootstrap_elo.json (1000 cluster-bootstrap replicates resampled by question) shows the frontier band's CIs are mutually overlapping but strictly disjoint from Llama's, and GPT-3.5's CI is strictly below Phi-4's, so both the frontier clustering and GPT-3.5's last place are statistically clean rather than orderings of noise. The bootstrap means reproduce the paper's approximate Figure 2 values to within ~40 Elo across all eight models, which is far tighter agreement than this claim requires. The only ordering difference from the paper's quoted values is a Phi-4/GPT-4o swap inside the middle band (436.3 vs 434.3 point estimates, a 2-Elo separation that is a tie, and the bootstrap means reverse it back to the paper's order); the claim explicitly treats the middle three as an unordered band, so this is immaterial. Two robustness notes from assess/fix_severity.json: the code path underlying this claim was the subject of Fix 5, rated critical, because analysis/bt_utils.py shipped the paper's Eq. (5) as unrunnable, silently-divergent dead code whose outcome-resolution logic predated the utils.question_key refactor. That fix does not undermine the verdict, because the repaired collector was verified to reproduce the shipped collect_games exactly, on an identical 1868-game multiset over (answerer, questioner, outcome) plus an identical skip Counter, which establishes that the replicated Elos rest on the authors' resolution rules rather than reconstructed guesses. Fix 7 (minor) notes genuine cluster-bootstrap bias for thin models (Phi-4 and GPT-3.5 authored few questions), but that affects benchmarker Elos and CI width, not the answerer rank structure this claim asserts. The regenerated figures/figure2_answerer_elo.pdf visually confirms the three-tier structure.
32
+
33
+ ---
34
+
35
+ **C5** [supporting/figure] β€” status **match**
36
+
37
+ *Claim:* Figure 2 plots answerer Elos for all eight models with bootstrap 95% confidence intervals.
38
+
39
+ *Verifier rationale:* The guessed path figures/answerer_elos.pdf does not exist, but the run produced figures/figure2_answerer_elo.pdf, which is unambiguously the claimed figure. Reading it directly confirms every structural feature the claim describes: a point-and-errorbar plot with exactly eight models on the x-axis, Elo on the y-axis spanning roughly -200 to 2050, vertical 95% CI bars, models sorted by descending Elo with deepseek-v3.2-speciale leftmost/highest (~1740) and gpt-3.5-turbo-0125 rightmost/lowest (~-25). The four frontier models' CIs overlap each other (all spanning roughly 1420-2020) while being clearly separated from the middle band (llama-4-maverick ~730, then gpt-4o and phi-4 near ~400). The title confirms the intervals come from a cluster bootstrap over questions (1000 replicates, resampled by question), matching the claim's stated CI construction. One deviation worth noting for interpretation rather than verdict: fix_severity.json Fix 7 records that the bootstrap driver was written by the agent (the repo ships no bootstrap code, only the unused per-edge-weight hook fit_itemized_bt_weighted at bt_utils.py:399), and that the plot centers error bars on the bootstrap mean while overlaying the full-data point estimate as a separate orange marker. This two-series presentation is a small addition relative to the paper's single-series figure, but it was a deliberate choice to keep genuine cluster-bootstrap bias for thin models visible rather than clipping it away, and it does not remove or contradict any claimed feature. Fix 5 (critical, analysis/bt_utils.py) affects the upstream game-collection path, but its verification -- an exact 1868-game multiset match against the shipped collect_games -- supports that the Elos underlying this figure rest on the authors' resolution rules.
40
+
41
+ ---
42
+
43
+ **C6** [supporting/figure] β€” status **partial**
44
+
45
+ *Claim:* Benchmarker (questioner) Elos rank GPT-5.2 highest, followed by Gemini 3 Pro Preview, DeepSeek v3.2 Speciale, and Claude 4.5 Opus, with GPT-4o, GPT-3.5 Turbo, Llama 4 Maverick, and Phi-4 clustered near or below zero.
46
+
47
+ *Verifier rationale:* The underlying data for this figure was fully produced and reproduces the claim closely, but the figure itself was never rendered. results/bootstrap_elo.json and results/bootstrap_stdout.txt contain benchmarker (questioner) Elos for all eight models with percentile 95% CIs from 1000 cluster-bootstrap replicates over 271 question clusters (1868 games). GPT-5.2 is highest by a clear, non-overlapping margin (point 1195.6 vs paper ~1220), and the top-four ordering matches the claim exactly: Gemini 3 Pro Preview 934.7 (~920), DeepSeek v3.2 Speciale 820.4 (~800), Claude Opus 4.5 646.0 (~630). The bottom four all sit near or below zero β€” GPT-4o 69.9 (~0), GPT-3.5 Turbo -7.8 (~-90), Phi-4 -112.2 (~-220), Llama 4 Maverick -128.3 (~-200) β€” with heavily overlapping CIs, so their internal ordering is not decisive, as the claim itself notes. The bootstrap means track the paper even more tightly (1205.4 / 925.6 / 806.7 / 620.6 / 4.6 / -77.5 / -214.0 / -214.2). What is missing is only the drawing step: run_bootstrap.py:124 computes the benchmarker summary and writes it to JSON, but its sole plotting call at line 161 is _figure2(answerer, ...), so the run emitted figures/figure2_answerer_elo.pdf (the answerer counterpart, C-claim for Figure 2) and no questioner figure. A search of the whole codebase excluding venv found no PDF/PNG/SVG plotting benchmarker Elos, and the replication log references no such path. This is the 'content reproduced, assembly missing' case, so partial rather than not_attempted (evidence_found is true) and rather than match (no figure exists to inspect for structural features). fix_severity.json rates the relevant driver work as minor and not a repo defect in the resolution path: Fix 7 created run_bootstrap.py using the shipped-but-unused fit_itemized_bt_weighted hook, and Fix 5 (critical) realigned analysis/bt_utils.py's stale collect_itemized_games with the shipped collect_games, verified by an exact 1868-game multiset match over (answerer, questioner, outcome) plus an identical skip Counter β€” so the Elo values these figures would plot rest on the authors' own resolution rules. Fix 7 also notes real cluster-bootstrap bias for thin models (Phi-4 authored 15 questions, GPT-3.5 only 8), which is why those two models' point estimates fall outside their bootstrap CIs and why their wide intervals should not be over-read.
48
+
49
+ ---
50
+
51
+ **C7** [supporting/scalar] β€” status **match**
52
+
53
+ *Claim:* Of 2464 candidate question-answer pairs, 1897 (76.99%) passed feasibility gating; in 1819 cases (73.82%) the answerer provided an answer, which the benchmarker declared correct in 1201 (48.74%) and incorrect in 618 (25.08%) cases.
54
+
55
+ *Replicated:* `[2464, 1905, 77.31, 1831, 74.31, 1212, 49.19, 619, 25.12]` *Paper:* `[2464, 1897, 76.99, 1819, 73.82, 1201, 48.74, 618, 25.08]`
56
+ *Grading rule:* [0] rel err 0.00% ≀ 5% β†’ match; [1] rel err 0.42% ≀ 5% β†’ match; [2] rel err 0.42% ≀ 5% β†’ match; [3] rel err 0.66% ≀ 5% β†’ match; [4] rel err 0.66% ≀ 5% β†’ match; [5] rel err 0.92% ≀ 5% β†’ match; [6] rel err 0.92% ≀ 5% β†’ match; [7] rel err 0.16% ≀ 5% β†’ match; [8] rel err 0.16% ≀ 5% β†’ match
57
+
58
+ *Verifier rationale:* [deterministic grade] [0] rel err 0.00% ≀ 5% β†’ match; [1] rel err 0.42% ≀ 5% β†’ match; [2] rel err 0.42% ≀ 5% β†’ match; [3] rel err 0.66% ≀ 5% β†’ match; [4] rel err 0.66% ≀ 5% β†’ match; [5] rel err 0.92% ≀ 5% β†’ match; [6] rel err 0.92% ≀ 5% β†’ match; [7] rel err 0.16% ≀ 5% β†’ match; [8] rel err 0.16% ≀ 5% β†’ match. Comparator notes: Step 2 of the replication ran the shipped statistics entry point unmodified (`python stats.py --disable-disagreement-logs`, exit 0; git status confirms stats.py is not among the patched files) and its 'Protocol flow' block in results/stats_stdout.txt reproduces the paper's Appendix D funnel closely. The total candidate pair count matches exactly at 2464. The feasibility-gating stage corresponds to stats.py's Step 3 (lines 595-646), where Bob critiques Alice's self-answer: of 1925 candidate games, the only drop is 'Bob's critique no victory' (20), since 'critique missing', 'critique unknown' and 'critique upheld' are all 0, so 1889 'critique says correct' + 16 'critique rejected' = 1905 pairs pass gating (77.31% of 2464, vs the paper's 1897 / 76.99%). That 1905 is confirmed by the answer stage partitioning it exactly: 70 'Bob fails to answer' + 4 'Bob claims ill-posed' + 1831 'Bob answers successfully' = 1905. Bob provided an answer in 1831 cases (74.31% vs the paper's 1819 / 73.82%); Alice declared the answer correct in 1212 (49.19% vs 1201 / 48.74%) and incorrect in 11 + 576 + 32 = 619 (25.12% vs 618 / 25.08%). The internal-consistency checks the claim asks for both hold: 1212 + 619 = 1831 exactly mirrors the paper's 1201 + 618 = 1819, and answered + unanswered reconciles against the gated total. Every count is within ~1% relative error of the paper, the feasibility pass rate of 77.31% sits inside the accepted 65-88% band, and the correct/incorrect split among answered cases is 1.96:1, matching the expected ~2:1. Note the replication's own percentages after the candidate-games line are denominated in candidate games, not the 2464 total (stats.py:770-794), so I recomputed all rates against 2464 to match the paper's convention. Of the two critical fixes in assess/fix_severity.json, neither is in this code path: Fix 5 touches analysis/bt_utils.py (itemized BT) and Fix 9 adds external_benchmarks.json (external validity); no fix alters stats.py or the funnel it computes.
59
+
60
+ ---
61
+
62
+ **C8** [supporting/scalar] β€” status **partial**
63
+
64
+ *Claim:* Adjudication volume: automated labellers upheld the incorrectness claim in 566 cases (22.97%), rejected it in 10 (0.41%), and escalated 44 cases (1.79%) to human adjudicators; humans additionally adjudicated 16 self-answer-incorrect claims and 1 ill-posedness claim (61 total), plus 125 claims from preliminary experiments (186 overall).
65
+
66
+ *Replicated:* `[576, 23.38, 11, 0.45, 32, 1.3, 26, 2, 73, 103, 176]` *Paper:* `[566, 22.97, 10, 0.41, 44, 1.79, 16, 1, 61, 125, 186]`
67
+ *Grading rule:* [0] rel err 1.77% ≀ 5% β†’ match; [1] rel err 1.78% ≀ 5% β†’ match; [2] rel err 10.00% ≀ 30% β†’ partial; [3] rel err 9.76% ≀ 30% β†’ partial; [4] rel err 27.27% ≀ 30% β†’ partial; [5] rel err 27.37% ≀ 30% β†’ partial; [6] rel err 62.50% > 30% β†’ no_match; [7] rel err 100.00% > 30% β†’ no_match; [8] rel err 19.67% ≀ 30% β†’ partial; [9] rel err 17.60% ≀ 30% β†’ partial; [10] rel err 5.38% ≀ 30% β†’ partial
68
+
69
+ *Verifier rationale:* [deterministic grade] [0] rel err 1.77% ≀ 5% β†’ match; [1] rel err 1.78% ≀ 5% β†’ match; [2] rel err 10.00% ≀ 30% β†’ partial; [3] rel err 9.76% ≀ 30% β†’ partial; [4] rel err 27.27% ≀ 30% β†’ partial; [5] rel err 27.37% ≀ 30% β†’ partial; [6] rel err 62.50% > 30% β†’ no_match; [7] rel err 100.00% > 30% β†’ no_match; [8] rel err 19.67% ≀ 30% β†’ partial; [9] rel err 17.60% ≀ 30% β†’ partial; [10] rel err 5.38% ≀ 30% β†’ partial. Comparator notes: The shipped stats.py ran clean at its documented defaults and reproduced the adjudication funnel over an identical base of 2464 total pairs considered (results/stats_stdout.txt:57-83). The three automated-panel dispositions of the answer-incorrectness claim replicate closely: upheld ('Alice says incorrect, Alice wins') 576 = 23.38% of 2464 vs the paper's 566/22.97% (1.8% relative error); rejected ('Alice says incorrect, Bob wins') 11 = 0.45% vs 10/0.41% (off by one case); and escalated ('Alice says incorrect, no victory', broken down at stats_stdout.txt:90-91 as 'no unanimity among judges: 32', i.e. exactly the non-unanimous automated panel that the paper's protocol sends to humans) 32 = 1.30% vs 44/1.79%. The contested totals agree almost exactly: 576+11+32 = 619 here vs 566+10+44 = 620 in the paper. Both decisive qualitative checks pass: the human-escalation rate is low single-digit (1.30%, paper 1.79%), so 94.8% of contested claims (587/619) are resolved by unanimous automated vote, and the rejected count is far smaller than the upheld count (11 vs 576, a 52x ratio, matching the paper's 57x). One important caveat on provenance for the escalation element: replication_log.json step 2 and results/reproducibility_gaps.json record that 'escalat' has zero hits in shipped Python, so the repo has no explicit escalation code path and the agent reported the count UNAVAILABLE rather than infer it; I derived 32 from stats.py's own no-unanimity drop bucket, which is the direct structural analogue. An alternative reading -- counting answer-critique claims that actually received a human label (evaluator mode) -- gives 45 = 1.83%, an even closer match to 44/1.79%; the two readings differ because some human labels were collected for the RQ2 adjudicator-agreement study on claims the panel had already decided unanimously. Either reading satisfies the claim's qualitative test, so the verdict is unaffected. The human-adjudication decomposition, which I computed with the repo's own stats.count_human_labels over evaluations/ (8 labeller files, 179 decisions, 176 unique claim keys), is weaker: 73 contested claims in the current experiment carry a human label (paper 61), splitting as 45 evaluator-mode / 26 self-answer-incorrect contradictor-mode (paper 16, the clearest miss at 63% relative error) / 2 ill-posedness (paper 1); 176 unique human-adjudicated claim keys overall (paper 186) minus the 73 current ones leaves 103 superseded/preliminary claims (paper 125). Note that stats.py's funnel independently prints 'Bob's critique rejected: 16', which coincidentally equals the paper's self-answer figure but is a different quantity (an automated outcome, not a human adjudication count). fix_severity.json records no fix touching stats.py -- it ran unpatched -- and the one fix in this step (pipeline_coverage.py, rated major) is a diagnostic script that does not feed these counts, so no known fix explains the residual gaps. Verdict is partial: the headline automated-panel counts and the protocol's claimed tractability replicate well, while the human-adjudication volume decomposition runs 15-60% high.
70
+
71
+ ---
72
+
73
+ **C9** [supporting/scalar] β€” status **match**
74
+
75
+ *Claim:* Of the 1897 traces used in the Bradley-Terry fit, 671 (35.37%) were benchmarker (Alice) victories and 1226 (64.23%) were answerer (Bob) victories.
76
+
77
+ *Replicated:* `[1868, 656, 35.12, 1212, 64.88]` *Paper:* `[1897, 671, 35.37, 1226, 64.23]`
78
+ *Grading rule:* [0] rel err 1.53% ≀ 5% β†’ match; [1] rel err 2.24% ≀ 5% β†’ match; [2] rel err 0.71% ≀ 5% β†’ match; [3] rel err 1.14% ≀ 5% β†’ match; [4] rel err 1.01% ≀ 5% β†’ match
79
+
80
+ *Verifier rationale:* [deterministic grade] [0] rel err 1.53% ≀ 5% β†’ match; [1] rel err 2.24% ≀ 5% β†’ match; [2] rel err 0.71% ≀ 5% β†’ match; [3] rel err 1.14% ≀ 5% β†’ match; [4] rel err 1.01% ≀ 5% β†’ match. Comparator notes: Two independent paths in this run produced the same trace split, and both land within ~2% of every paper number. Step 3 ran the shipped entry point (compute_bt_elo.py) at documented defaults and reported 'Games: 1868' with the per-role win/loss records summing to answerer 1212 + questioner 656 = 1868 (replication_log.json step 3 stdout and notes; results/compute_bt_elo_stdout.txt). Step 4 fit the paper's actual itemized Eq. (5) model and independently recorded n_games=1868, n_answerer_victories=1212, n_benchmarker_victories=656 in results/itemized_bt.json, giving 656/1868 = 35.12% Alice victories and 1212/1868 = 64.88% Bob victories. Element-wise relative error against the paper's [1897, 671, 35.37, 1226, 64.23] is -1.53%, -2.24%, -0.71%, -1.14% and +1.01% β€” all comfortably inside the 5% match band β€” and the qualitative check the claim asks for is satisfied: the answerer-win share is the clear majority at 64.88%, squarely inside the accepted ~55-72% window, and unlike the paper's own figures the replicated percentages sum to 100.00% exactly. The ~29-trace shortfall (1868 vs 1897, 1.5%) most plausibly reflects small differences in the feasibility gate; the step-3 skip Counter accounts for the exclusions (question_invalid_self_answer: 588, critique_no_majority: 28, question_missing_or_failed: 14, illposed_no_majority: 1). Regarding fix_severity.json: Fix 5 is rated critical and sits directly in this code path β€” analysis/bt_utils.py held the paper's Eq. (5) model as unrunnable dead code forked from an older generation of the outcome-resolution logic, and a signature-only patch would have silently fit a different game set. That risk did not materialize here: the repaired collector was verified to reproduce compute_bt_elo.collect_games exactly (identical 1868-game multiset over (answerer, questioner, outcome) plus an identical skip Counter), which is why steps 3 and 4 agree on the split and why the count rests on the authors' shipped resolution rules rather than the agent's reconstruction of them. The fix therefore explains nothing about the residual 1.5% gap and does not undercut the verdict.
81
+
82
+ ---
83
+
84
+ **C10** [supporting/scalar] β€” status **match**
85
+
86
+ *Claim:* The experiment generates exactly one question package per model per MSC category, yielding 352 questions total (44 categories x 8 models) and 2464 candidate question-answer pairs (each question answered by the 7 other models).
87
+
88
+ *Replicated:* `[44, 8, 352, 2464]` *Paper:* `[44, 8, 352, 2464]`
89
+ *Grading rule:* [0] rel err 0.00% ≀ 5% β†’ match; [1] rel err 0.00% ≀ 5% β†’ match; [2] rel err 0.00% ≀ 5% β†’ match; [3] rel err 0.00% ≀ 5% β†’ match
90
+
91
+ *Verifier rationale:* [deterministic grade] [0] rel err 0.00% ≀ 5% β†’ match; [1] rel err 0.00% ≀ 5% β†’ match; [2] rel err 0.00% ≀ 5% β†’ match; [3] rel err 0.00% ≀ 5% β†’ match. Comparator notes: All four structural counts reproduce exactly. (1) 44 categories: configs/ reports 'run_ids: 44' (results/config_counts.txt), and every one of the 8 files in benchmarks/ spans exactly 44 distinct topic_slug values with a union of exactly 44 across the corpus. (2) 8 models: configs/models.json and results/config_counts.txt list 8 models, and benchmarks/ holds exactly 8 per-model files; pipeline_coverage.py independently reports 8 benchmark/answer/critique/judge models. (3) 352 questions, one per (model, category) cell: the shipped pipeline_coverage.py computes its expectation as len(run_ids) * len(benchmark_specs) (pipeline_coverage.py:337) and prints 'Questions: Generated: 352/352 (100.0%)' (results/pipeline_coverage_stdout.txt). I confirmed the one-per-cell distribution independently: each of the 8 benchmark files covers all 44 slugs with at least one succeeded entry, summing to 8 x 44 = 352. The raw entry count is higher (772) purely because of outer_attempt retries; under the repo's latest-outer-attempt semantics the grid is exactly one question per cell. One question (of 352) has status 'failed', so 351 succeeded -- this does not affect the authored-question count of 352. (4) 2464 candidate pairs: the shipped stats.py prints 'Total pairs considered: 2464' (results/stats_stdout.txt:58). This is computed, not hardcoded -- pairs_total is incremented by enumeration (stats.py:575) over latest-outer-attempt benchmark entries crossed with eligible answerers, and the b != a rule is implemented explicitly in the parallel path at pipeline_coverage.py:352-356 ('allow_self_answering or spec.slug != q_slug'), giving 352 x 7 = 2464. The stats.py funnel then partitions it exactly: 1925 candidate games + 7 dropped for missing/failed question + 532 dropped for missing answer = 2464. Note on a quantity that could be mistaken for this one: pipeline_coverage.py reports only 2261 answers *generated* against an expectation of 2457 (= 351 x 7, its denominator excluding the single failed question), because microsoft-phi-4 as answerer covers only 16 of 44 categories against every questioner -- a 196-answer shortfall in the shipped artifact that reconciles exactly with my independent count (2268 distinct (answerer, questioner, slug) cells, minus the 7 belonging to the failed question, equals the tool's 2261). That is a real incompleteness in the released answer data, but it measures answers actually produced, not the claim's quantity: the paper's 2464 is the count of *candidate* pairs where an answer is attempted, which is the pre-drop enumeration that stats.py labels 'Total pairs considered' and then decomposes, with the missing phi-4 answers falling into its 'Dropped before candidate (answer missing): 532' bucket. No fix in assess/fix_severity.json touches this code path: the critical fixes concern analysis/bt_utils.py (dead Eq. (5) code) and the absent external_benchmarks.json, and the pipeline_coverage.py repair (Fix 3, major) preserved per-question counting semantics, which the agreement between that tool, stats.py, and my independent enumeration of the raw JSON corroborates.
92
+
93
+ ---
94
+
95
+ **C11** [supporting/table] β€” status **match**
96
+
97
+ *Claim:* Question-generation success rates vary sharply by model: frontier models (GPT-5.2, Gemini 3 Pro, DeepSeek v3.2 Speciale, Claude 4.5 Opus) produced a valid question for all 44 topics, while GPT-3.5 Turbo succeeded for only 8 of 44 topics.
98
+
99
+ *Verifier rationale:* [deterministic grade] 48 cell(s), 0 outside tolerance β†’ match. Comparator notes: Reproduced Table 6 exactly, all 48 cells. The shipped stats.py run (replication/codebase/results/stats_stdout.txt) does not print the Valid/Invalid/Failed breakdown directly, so I recomputed it from the checked-in artifacts using the repo's own resolution logic -- utils.collect_invalid_self_answer_questions plus the latest-outer-attempt filtering and STATUS_SUCCEEDED / final_question checks that compute_bt_elo.collect_games applies (compute_bt_elo.py:218-308) -- rather than reimplementing the rules. Per questioner model over 44 topics: Claude 4.5 Opus, DeepSeek v3.2 Speciale, Gemini 3 Pro Preview and GPT-5.2 each 44/0/0; GPT-4o 37/7/0; Llama 4 Maverick 35/9/0; Phi-4 15/29/0; GPT-3.5 Turbo 8/35/1. Every row sums to 44 and every cell equals the paper value. Three independent cross-checks in the produced evidence corroborate this: the total Valid count of 271 equals the paper Valid column sum and matches the "Question items: 271" reported by the itemized Eq. (5) fit (results/itemized_bt_stdout.txt); the per-questioner game counts in results/compute_bt_elo_stdout.txt track 7x the Valid counts (GPT-3.5 55 games ~ 8 questions, Phi-4 105 = 7x15); and assess/fix_severity.json (Fix 7) independently records "phi-4 at n=15, gpt-3.5 at n=8 authored questions". The fixes recorded in fix_severity.json (notably the critical Fix 5 to analysis/bt_utils.py and Fix 9 external_benchmarks.json) sit in the Bradley-Terry/external-validity paths downstream of question generation and do not touch this counting path, which reads the shipped benchmarks/, critiques/, automated_evaluations/ and evaluations/ artifacts through unmodified utils.py helpers. The qualitative pattern the claim asserts holds: the four frontier models are at 100% valid while GPT-3.5 Turbo reaches only 18.2%.
100
+
101
+ ---
102
+
103
+ **C12** [supporting/table] β€” status **not_attempted**
104
+
105
+ *Claim:* Critique agreement rates are uniformly high across all eight models (0.88-0.96) while answer correctness rates vary widely (0.16-0.99), showing that catching errors is substantially easier than avoiding them.
106
+
107
+ *Verifier rationale:* [deterministic grade] comparator reported no replicated table was produced. Comparator notes: This run did not produce a Table 10 equivalent, so there is no replicated value to compare. The claim's verification instructions require a critic-reliability artifact that uses GPT-5.4-high as a ground-truth proxy to compute per-model answer correctness and critique agreement rates, but the replication established that no such artifact or code path exists in the repository. The plan step that targeted C12 (replication_plan.json step 2, replication_log.json::step_outcomes[1]) ran the shipped stats.py, and its own notes record C12 as 'only partially addressed': 'gpt-5.4' has zero hits repo-wide, and configs/models.json lists exactly the eight evaluated models with no GPT-5.4-high entry. The independent gap report (codebase/results/reproducibility_gaps.json, caveats entry for C12) reaches the same conclusion with grep evidence, noting that stats.py's agreement rates are victory rates resolved over the full judge panel (victory.py:183 resolve_automated_victory), a different quantity from agreement against a ground-truth proxy. I confirmed this directly: a grep across the codebase for critique_agreement / agreement_rate / answer_correctness / correctness_rate / critique_f1 returns no hits in any shipped or agent-written Python or JSON, and no results file contains per-model values for these three columns. The per-model blocks that stats.py does print (codebase/results/stats_stdout.txt: 'Self-answer correctness', 'Defender success', 'Self-answers declared correct by critics') are panel-resolved debated victory rates over small and unequal critique counts, not the paper's per-model correctness and agreement rates, so they cannot be substituted cell-for-cell -- and in particular no critique-agreement or critique-F1 quantity is computed anywhere, meaning the decisive asymmetry check (uniformly high agreement with spread under ~0.10 versus correctness spanning over ~0.7, and specifically GPT-3.5 Turbo at ~0.16 correctness yet ~0.95 agreement) cannot be evaluated from this run in either direction. Producing it would require new code plus fresh paid API calls to a model the artifact never references. None of the ten entries in assess/fix_severity.json touch this code path, so no applied fix explains the absence; it is a genuine gap in the published artifact rather than a replication failure. Reported as not_attempted rather than no_match: the run produced no value, and the claim is neither supported nor refuted by the available evidence.
108
+
109
+ ---
110
+
111
+ **C13** [supporting/scalar_range] β€” status **not_attempted**
112
+
113
+ *Claim:* Among pairs of non-proven-wrong answers to the same question, only 4.23% (160/3779, spanning 271 protocol-ready questions) are mutually incompatible, and incompatibility is concentrated in pairs involving weak models (GPT-3.5 pairs 14.5-18.3%) versus near-zero among frontier-model pairs (0.0-0.4%).
114
+
115
+ *Verifier rationale:* [deterministic grade] comparator reported no replicated value was produced. Comparator notes: This run produced no Table 11 equivalent, so there is no pairwise mutual-incompatibility rate to compare against the paper's 4.23% pooled / 14.5-18.3% GPT-3.5 / 0.0-0.4% frontier stratification. Plan step 9 (replication/replication_log.json, step_id 9) explicitly classified C13 as NOT_REPRODUCIBLE, and I independently confirmed both halves of that finding rather than relying on the report's self-description. Code side: `grep -rIn -i -E 'incompat|mutually' --include='*.py'` over the patched codebase returns exactly one hit, labeller_app_ordered.py:802 `parser.add_mutually_exclusive_group()` β€” an argparse substring collision, not incompatibility logic; there are also zero hits for `protocol.?ready|proven.?wrong`. Data side: I loaded all 8 files in automated_evaluations/ (13,665 decisions) and the union of decision fields is ['answer_model','confidence','critic_model','error','judge_model','mode','outer_attempt','question_model','raw_response','reasoning','run_id','status','topic_slug','type','verdict'] β€” each decision carries exactly one `answer_model`, i.e. verdicts are per-answer, never a pairwise judgement of two answers to the same question, so the 3779 pairs cannot be reconstructed from shipped artifacts. The paper's stated classifier GPT-5.4-high appears nowhere in the repo (grep -rIn -i 'gpt-5\.4' hits only this run's own output files); configs/models.json lists exactly the 8 evaluated models, none of them GPT-5.4-high. Reproducing C13 would therefore require both writing the pairwise-classification code from scratch and making new paid API calls to a model absent from the repo β€” and fix_severity.json Fix 2 records that provider API keys were deliberately unset to guarantee no paid calls. No fix in assess/fix_severity.json touches this path (Fixes 1/3/5 concern compute_bt_elo.py's import-time breakage and utils.question_key drift), so the fixes neither explain nor alter this verdict: the claim was not attempted rather than attempted-and-failed, and no evidence here either supports or refutes the paper's stratification.
116
+
117
+ ---
118
+
119
+ **C15** [supporting/scalar] β€” status **missing**
120
+
121
+ *Claim:* Benchmarker strength (alpha) and answerer strength (beta) are moderately but not fully correlated (Spearman rho = 0.69, Kendall tau = 0.47), with Llama 4 Maverick ranking notably higher as an answerer than as a benchmarker and GPT-4o showing the opposite pattern.
122
+
123
+ ---
124
+
125
+ **C16** [supporting/figure] β€” status **match**
126
+
127
+ *Claim:* All models, including weak ones such as GPT-3.5 Turbo and GPT-4o, show consistently high agreement with the human adjudicator when adjudicating outcomes (human-model agreement rates of 86.9-94.6% using outcome categories).
128
+
129
+ *Verifier rationale:* The produced figure figures/figure3_agreement.pdf is the Figure 3 equivalent: a heatmap of pairwise adjudicator agreement on outcome categories among the eight model adjudicators plus a Human row/column, color-coded by agreement percentage with per-cell n values. The underlying data (results/adjudicator_agreement.json, human_model_agreement) confirms the Human row is uniformly high and every human-model rate exceeds 80%: GPT-5.2 90.1%, Claude Opus 4.5 94.4%, Gemini 3 Pro 93.0%, GPT-4o 93.0%, GPT-3.5 Turbo 93.0%, Llama 4 Maverick 93.0%, Phi-4 93.0%, DeepSeek V3.2 Speciale 88.7% (range 88.7-94.4%, vs paper's stated 86.9-94.6%). Critically, the weak adjudicators GPT-3.5 Turbo and GPT-4o both agree with humans at 93.0%, comparable to frontier models and not systematic outliers β€” exactly the claimed qualitative pattern. Individual per-model values differ by at most ~2 pts from the paper's reported figures (e.g. DeepSeek 88.7 vs 86.9, Phi-4 93.0 vs 94.6), well within replication noise given n=71 shared human-adjudicated claims. Figure 9 (suboutcome granularity) was also produced. The only structural difference from the described figure is that the matrix is rendered full-symmetric rather than strictly lower-triangular, which does not affect the content. The claim is supported.
130
+
131
+ ---
132
+
133
+ **C19** [supporting/figure] β€” status **match**
134
+
135
+ *Claim:* Victory rates among valid traces show near-total dominance of frontier answerers over weak benchmarkers: every question authored by GPT-4o, Llama 4 Maverick, Phi-4, or GPT-3.5 Turbo was answered successfully (100% answerer victory) by each of the four frontier models.
136
+
137
+ *Verifier rationale:* The produced figure figures/victory_rate_matrix.pdf and its backing data results/victory_matrix.json reproduce Figure 5's structure exactly. It is an 8x8 heatmap with question authors (Alice) as rows and answerers (Bob) as columns, diagonal excluded, each cell giving Bob's victory rate as a percentage with n. The core claim holds: every one of the 16 cells for weak authors (GPT-4o, GPT-3.5 Turbo, Llama-4-Maverick, Phi-4) crossed with the four frontier answerers (GPT-5.2, Claude Opus 4.5, Gemini 3 Pro, DeepSeek-v3.2) is 100% (answerer_victory_rate=1.0). The paper's specifically cited cells also match: GPT-5.2 author vs GPT-4o/Llama/Phi-4/GPT-3.5 answerers = 0.0/9.3/2.3/0.0 (paper 0.0/9.1/2.3/0.0) and GPT-3.5 Turbo author vs GPT-5.2/Gemini/DeepSeek/Claude answerers = 100/100/100/100 (paper identical). The strong block structure (lower-left ~100%, upper-right ~0%) is clearly present. The expected filename figures/victory_rates.pdf was not the exact name used (figure emitted as victory_rate_matrix.pdf), but this is the correct figure content.
pages/conclusion/page.md ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Conclusion
2
+
3
+
4
+ ---
5
+ <!-- trackio-cell
6
+ {"type": "markdown", "id": "cell_4876c79726ac", "created_at": "2026-07-18T16:28:34+00:00", "title": "Summary of reproduction", "pinned": true, "pinned_at": "2026-07-18T16:28:34+00:00"}
7
+ -->
8
+ **Scorecard: 1 verified / 0 falsified / 1 toy / 0 inconclusive of 2 board claims.**
9
+
10
+ The paper reproduces well. Every central result comes back in the same direction and close to the reported magnitude, including the external-validity correlations, the predictive-validity gap over baserate, the adjudicator-robustness finding, and the three-tier answerer Elo ranking. The main caveat is about the artifact, not the science: the shipped code did not run as-is, the paper's own Eq. (5) model was present only as unrunnable dead code, and three claims (a mutual-incompatibility rate, a prompt-ablation, and a GPT-5.4-high critic-reliability table) cannot be reproduced from the repository at all without new paid API calls to a model the repo never defines.
11
+
12
+ This is a strong replication of an offline-analysis paper. Of 21 claims, the substantive verdict is that the science holds: the structural counts (C7, C9, C10, C11) match the paper to within about 1%, the Elo cluster structure and victory-rate block structure (C4, C5, C6, C19) reproduce, the two RQ headlines (C1, C2, C3, C16) reproduce, and the supporting figures (C20, C21) and hyperparameters (C18) land in the paper's regime. Three claims were correctly not attempted (C12, C13, C14) because they require resources the environment and repository lack. The headline replication score is deterministic and driven by per-cell tolerance grading, which pushes several genuinely-successful claims to 'partial' for reasons that do not reflect a real disagreement: C1 misses a 5% band by 0.03 on a correlation coefficient, C2 misses on two cells where the replicated model is slightly BETTER calibrated than the paper (log-loss 0.185 vs 0.208), and C3's Kendall tau cells miss because a single adjacent-rank swap among 8 items moves tau by 0.071, a mechanical artifact of a coarse discrete statistic. What the score does not capture, in either direction: it does not reward that the reproduced numbers were computed independently and still landed on the paper (a copier would score identically), and it does not penalize that the code needed substantial repair to run at all. Attribution overall: the numeric agreement is a faithful reproduction; the not-attempted claims are a mix of missing paid compute and repository incompleteness; and the paper's own model shipped as silently-divergent dead code, which is an artifact defect rather than a methodology defect.
pages/index.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Repro - Benchmarking at the Edge of Comprehension
2
+
3
+ ## Pages
4
+
5
+ | Page |
6
+ | --- |
7
+ | [00 - Scored evidence summary](#/00-scored-evidence-summary) |
8
+ | [Claim 1 - Critique-Resilient Benchmarking framework produces stable sc…](#/claim-1-critique-resilient-benchmarking-framework-produces) |
9
+ | [Claim 2 - Reformulates benchmarking as adversarial generation-evaluati…](#/claim-2-reformulates-benchmarking-as-adversarial-generatio) |
10
+ | [Reproduction protocol and provenance](#/reproduction-protocol-and-provenance) |
11
+ | [Conclusion](#/conclusion) |
pages/reproduction-protocol-and-provenance/page.md ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ # Reproduction protocol and provenance
2
+
3
+
4
+ ---
5
+ <!-- trackio-cell
6
+ {"type": "markdown", "id": "cell_66bd852376ba", "created_at": "2026-07-18T16:28:34+00:00", "title": "Protocol and provenance"}
7
+ -->
8
+ **Provenance.** This logbook backfills a replication executed 2026-07-17 to 2026-07-17 by [veritas](https://arxiv.org/abs/2607.02931), an open-source replication agent; verdicts and evidence quote the run's per-claim verification records verbatim.
trackio-logo-light.png ADDED
trackio-logo.png ADDED
trackio-wordmark-dark.png ADDED