File size: 111,115 Bytes
3d9f5ca
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
# SchemaForge: Distilling Ultra-Large Foundation Models into Edge SLMs for Real-Time Enterprise JSON Extraction

### A Comparative Study of Gemma-4 Teachers and MiniCPM5-1B

**Arjhine A. Ty**

---

**Project codename:** SchemaForge
**Student model:** `SchemaForge-1B` (base architecture: `openbmb/MiniCPM5-1B`, 1.08B parameters)
**Teacher models:** `google/gemma-4-31B` (31B flagship), `google/gemma-4-E4B-it` (4B/8B-effective instruction-tuned)
**Production checkpoint:** `schemaforge-1b-iter2` (formerly `distilled_minicpm5_1b_iter2`)
**Hardware:** 1 Γ— NVIDIA RTX PRO 6000 Blackwell Edition (96 GB VRAM), Nebius AI Cloud
**Software:** Python 3.12, PyTorch 2.5, HuggingFace `transformers` 5.x, vLLM (serving)
**Version:** 1.0 β€” August 2026

---

## Abstract

Structured entity extraction β€” converting heterogeneous unstructured business documents into strongly-typed, schema-conformant JSON β€” is one of the highest-volume workloads in enterprise retrieval-augmented generation (RAG) and transactional automation. The prevailing solution is to deploy a 30B+ parameter general-purpose foundation model behind a constrained-decoding wrapper. This is operationally untenable at scale: in our measurements a `gemma-4-31B`-class teacher requires **β‰ˆ38.5 GB of VRAM** and sustains only **12.40 tokens/second** of greedy decode throughput on a single 96 GB accelerator, leaving room for only two concurrent workers and failing sub-second API SLAs.

We present **SchemaForge**, a sequence-level knowledge-distillation framework that compresses the structural reasoning and JSON formatting strictness of Gemma-4 teachers into `openbmb/MiniCPM5-1B`, a dense 1.08-billion-parameter edge Small Language Model (SLM). SchemaForge combines a hard-target cross-entropy objective with a temperature-scaled, log-space soft-target KL divergence under a multi-task weighting $\mathcal{L}_{KD} = \alpha\mathcal{L}_{CE} + (1-\alpha)\tau^{2}\mathcal{L}_{KL}$ ($\alpha = 0.5$, $\tau = 2.0$), and resolves the teacher–student vocabulary mismatch ($|\mathcal{V}_T| = 256{,}000 \rightarrow |\mathcal{V}_S| = 130{,}560$) via a dual-tokenizer cross-encoding scheme with shared-subspace logit projection.

Empirically, the distilled student attains a **0.0 % JSON syntax error rate (100 % structural validity)** and a **field-level extraction F1 of 1.000** across a five-domain in-domain enterprise benchmark suite ($n = 5$ documents β€” finance, logistics, IT procurement, biomedical, cloud operations), against **34.2 % error rate** and **0.612 F1** for the undistilled base model. It sustains **61.91–76.27 tokens/second** in a **β‰ˆ2.4 GB VRAM** footprint β€” a **16.0Γ— memory reduction** and a **5.0×–6.2Γ— throughput improvement** over the 31B teacher, permitting **36 concurrent inference workers** per 96 GB card (**β‰ˆ2,746 tokens/second aggregate system throughput**).

Our central *scientific* finding is not the compression ratio but a sharp negative result about prompt fragility. Across three controlled retraining iterations we observe that a 1B-parameter student is **hypersensitive to the surface form of the prompt header**: identical data, identical loss, identical hyperparameters, and only a changed instruction prefix moved zero-shot validity on the public `suneeldk/text-json` benchmark from **0.0 %** to **70.0 %** and back to **0.0 %**. Distillation at this scale transfers a *format-conditioned* skill, not a format-invariant one. We formalize this as the **train–inference template gap**, quantify it, and derive a practical mitigation (template canonicalization plus a serving-side prompt contract) that we recommend as mandatory for any sub-2B structured-output deployment.

We release the production checkpoint, model card, proof artifacts, and full reproduction protocol.

**Keywords:** knowledge distillation, small language models, structured generation, JSON extraction, prompt sensitivity, edge inference, vLLM, enterprise RAG

---

## Executive Summary

| Dimension | 31B Teacher Baseline | SchemaForge-1B (Distilled) | Delta |
|---|---|---|---|
| In-domain JSON syntax error rate *(n = 5 docs)* | 0.0 % | **0.0 %** | Parity |
| In-domain field extraction F1 *(n = 5 docs)* | 1.000 | **1.000** | Parity |
| Zero-shot validity, `suneeldk/text-json` | β€” (not evaluated) | **70.0 %** | β€” |
| Decode throughput | 12.40 tok/s | **61.91–76.27 tok/s** | **5.0×–6.2Γ—** |
| Peak VRAM footprint | β‰ˆ38.5 GB | **β‰ˆ2.4 GB** | **16.0Γ— smaller** |
| Concurrent workers per 96 GB card | 2 | **36** | **18Γ—** |
| Aggregate system throughput | 24.8 tok/s | **β‰ˆ2,746 tok/s** | **β‰ˆ110Γ—** |

**The three claims a reader should take away:**

1. **Task-specialized distillation closes the capability gap at β‰ˆ29Γ— fewer parameters.** For a bounded, schema-constrained generation task, a 1.08B student matches a 31B teacher on both structural validity and field accuracy across our five-document in-domain suite. Generality is what gets compressed away; task competence is not.
2. **The economics change category, not degree.** 16Γ— less memory and 5–6Γ— faster decoding is the difference between two concurrent workers per GPU and thirty-six. This converts a per-document cost line into a rounding error.
3. **The fragility is in the prompt, not the weights.** The single largest swing in our entire experimental record β€” a 70-point accuracy delta β€” was caused by changing an instruction header. Practitioners deploying sub-2B structured extractors must treat the prompt template as a versioned, tested API contract.

**Caveat stated up front:** the in-domain suite comprises $n = 5$ curated documents (one per domain) and the winning checkpoint was distilled on $n = 5$ training samples. The 1.000 F1 and 0.0 % error rate are therefore *exact-match results on a small, curated set*, not population estimates. See [Β§10 Limitations](#10-limitations-and-threats-to-validity), which we consider a load-bearing section of this paper rather than a formality.

**This is a v1 release.** A second training and evaluation campaign is planned against a real-world document corpus with an expanded metric set, multi-seed variance, and competitive baselines; it is scoped in [Β§11.3](#113-planned-v2-retraining-and-evaluation-run) against the limitation register in [Β§10.6](#106-limitation-register). The accuracy figures below should be read as provisional pending that run.

---

## Table of Contents

1. [Introduction](#1-introduction)
2. [Related Work](#2-related-work)
3. [Problem Formulation](#3-problem-formulation)
4. [The SchemaForge Distillation Framework](#4-the-schemaforge-distillation-framework)
5. [Engineering: Running Legacy MiniCPM Under `transformers` 5.x](#5-engineering-running-legacy-minicpm-under-transformers-5x)
6. [Hyperparameter Sensitivity Analysis](#6-hyperparameter-sensitivity-analysis)
7. [Prompt Template Sensitivity: A Three-Iteration Study](#7-prompt-template-sensitivity-a-three-iteration-study)
8. [Empirical Results](#8-empirical-results)
9. [Deployment and Publication](#9-deployment-and-publication)
10. [Limitations and Threats to Validity](#10-limitations-and-threats-to-validity)
11. [Conclusion and Future Work](#11-conclusion-and-future-work)
- [References](#references)
- [Appendix A: Complete Hyperparameter Specification](#appendix-a-complete-hyperparameter-specification)
- [Appendix B: Compatibility Patch Source](#appendix-b-compatibility-patch-source)
- [Appendix C: Verbatim Prompt Templates](#appendix-c-verbatim-prompt-templates)
- [Appendix D: Benchmark Schemas and Source Documents](#appendix-d-benchmark-schemas-and-source-documents)
- [Appendix E: Reproducibility Checklist](#appendix-e-reproducibility-checklist)

---

## 1. Introduction

### 1.1 The Structured Extraction Bottleneck

Enterprise document intelligence pipelines are, in aggregate, a text-to-JSON problem. Accounts-payable automation reads invoices into `{invoice_number, vendor_name, invoice_date, subtotal, tax, grand_total}`. Freight settlement reads bills of lading into `{bill_of_lading, carrier_name, ship_date, container_count, freight_cost}`. Clinical procurement reads requisitions into `{order_id, supplier, order_date, items[], total_price}`. In every case the input is heterogeneous, semi-structured, and adversarially formatted by whichever upstream system emitted it; the output is a rigid, strongly-typed schema that a downstream database, ERP, or RAG index will reject outright if a single brace is unbalanced or a numeric field arrives as a string.

The industry default is to point a large instruction-tuned foundation model at the problem. This works. It also produces three compounding operational failures:

**(1) Memory pressure that forecloses multi-tenancy.** A 31B-parameter model in `bfloat16` requires roughly 62 GB for weights alone; even under aggressive quantization the working set we measured is **β‰ˆ38.5 GB** including KV cache. On a 96 GB RTX PRO 6000 Blackwell this admits only **two** concurrent workers at 90 % memory utilization β€” against 36 for the distilled student. Utilization economics collapse.

**(2) Throughput below SLA.** We measure **12.40 tokens/second** of greedy decode on a single-GPU deployment of the 31B teacher. A typical extraction emits 120–250 JSON tokens, implying **10–20 seconds per document**. Interactive document-upload flows target sub-2-second response; bulk nightly runs of 10⁢ documents become physically impossible on any reasonable cluster budget.

**(3) Cost per document that does not amortize.** Because throughput is low and memory forbids batching, the marginal cost of a document is dominated by GPU-seconds at the highest-tier instance price. There is no batching lever to pull.

Critically, **the capability being paid for is not the capability being used.** A 31B model's value lies in open-domain reasoning, long-horizon planning, multilingual fluency, and code synthesis. Invoice extraction requires almost none of that. It requires: (a) reliable span identification, (b) light numeric normalization, and (c) *absolute* fidelity to a JSON grammar. The pricing model charges for a general intelligence; the workload consumes a narrow, learnable competence.

This mismatch is precisely the setting in which knowledge distillation is theoretically well-motivated.

### 1.2 The SchemaForge Approach

**SchemaForge** is a sequence-level knowledge-distillation framework that transfers the schema-fidelity behavior of Gemma-4-class teachers into a 1.08B-parameter student. Its design commitments are:

- **Task-narrow, not capability-narrow.** We do not attempt to preserve the teacher's general ability. We preserve exactly one behavior: emit valid, schema-conformant JSON for a document-extraction prompt.
- **Multi-task objective.** Pure soft-logit matching is insufficient for structured output, because the tokens that matter most (`{`, `"`, `:`, `,`, `}`) are exactly the tokens where the teacher distribution is nearly deterministic and therefore carries little dark knowledge. We anchor with hard cross-entropy at $\alpha = 0.5$.
- **Vocabulary-mismatched teachers.** Gemma-4 and MiniCPM5 do not share a tokenizer. We do not retokenize or retrain embeddings; we cross-encode and project onto the shared logit subspace.
- **Deployability as a first-class metric.** Every result is reported alongside its VRAM footprint and throughput, measured on the same physical card.

### 1.3 Contributions

This paper makes five contributions:

**C1 β€” A reproducible distillation recipe for sub-2B structured extractors.** We specify the complete objective, hyperparameter bounds, and dataset scaling regime under which a 1.08B student reaches teacher-parity on schema fidelity ([Β§4](#4-the-schemaforge-distillation-framework), [Β§6](#6-hyperparameter-sensitivity-analysis)).

**C2 β€” A dual-tokenizer cross-encoding and logit-projection scheme** that permits KL-based distillation between models with incompatible vocabularies ($256{,}000 \rightarrow 130{,}560$) without embedding surgery or teacher retokenization ([Β§4.4](#44-dual-tokenizer-cross-encoding-and-vocabulary-projection)).

**C3 β€” Three low-level runtime compatibility patches** required to run legacy `trust_remote_code` MiniCPM checkpoints (`openbmb/MiniCPM-1B-sft-bf16`) under Python 3.12 / PyTorch 2.5 / `transformers` 5.x, including a tied-weights failure that reports itself as a *corrupted checkpoint* while actually being an API change ([Β§5](#5-engineering-running-legacy-minicpm-under-transformers-5x)). We also report the negative counterpart: migrating to `openbmb/MiniCPM5-1B`, a stock `LlamaForCausalLM`, removed the need for all three ([Β§5.4](#54-why-the-released-checkpoint-needs-none-of-these)).

**C4 β€” The train–inference template gap: a quantified negative result.** Three controlled iterations isolating prompt-header formatting as the sole varying factor, producing a 70-percentage-point accuracy swing on a public benchmark ([Β§7](#7-prompt-template-sensitivity-a-three-iteration-study)). We argue this is the dominant failure mode for small structured-output models and is systematically under-reported.

**C5 β€” A complete production deployment protocol**: HuggingFace publication, model card specification, vLLM serving configuration, schema-constrained decoding integration, and capacity planning arithmetic ([Β§9](#9-deployment-and-publication)).

### 1.4 Paper Roadmap

[Β§2](#2-related-work) situates the work. [Β§3](#3-problem-formulation) formalizes the task and metrics. [Β§4](#4-the-schemaforge-distillation-framework) gives the mathematics. [Β§5](#5-engineering-running-legacy-minicpm-under-transformers-5x) documents the engineering. [Β§6](#6-hyperparameter-sensitivity-analysis) and [Β§7](#7-prompt-template-sensitivity-a-three-iteration-study) present the two ablation studies. [Β§8](#8-empirical-results) reports all benchmarks. [Β§9](#9-deployment-and-publication) covers shipping. [Β§10](#10-limitations-and-threats-to-validity) is an honest accounting of what these numbers do and do not establish.

---

## 2. Related Work

### 2.1 Knowledge Distillation

The foundational formulation is due to Hinton, Vinyals, and Dean [1], who introduced temperature-softened logit matching with the now-standard $\tau^2$ gradient-rescaling term. Their key insight — that the *relative* probabilities a teacher assigns to incorrect classes encode a similarity structure ("dark knowledge") absent from one-hot labels — motivates our soft-target term. Buciluă et al. [2] anticipated the model-compression framing. Romero et al. [3] extended supervision to intermediate representations via FitNets; we deliberately do **not** use hidden-state matching, because the teacher (Gemma-4, $d_{\text{model}}$ and depth both far larger) and student (MiniCPM5-1B, 24 layers) have no principled layer correspondence, and projection heads introduce hyperparameters we could not afford to tune under our compute budget.

### 2.2 Sequence-Level Distillation for Autoregressive Models

Kim and Rush [4] established *sequence-level* knowledge distillation for neural machine translation, showing that training the student on teacher-generated output sequences (rather than only token-level distributions over gold data) substantially outperforms word-level KD. SchemaForge is sequence-level in this sense: our training targets are teacher-generated JSON completions, not human-annotated gold JSON. Sanh et al. [5] demonstrated the approach at scale with DistilBERT (40 % smaller, 97 % of GLUE performance), and Jiao et al. [6] with TinyBERT. More recent work on distilling instruction-following behavior β€” Gu et al. [7] on MiniLLM, and Agarwal et al. [8] on generalized KD with on-policy student samples β€” addresses the exposure-bias mismatch that arises when the student is trained on teacher trajectories but evaluated on its own. We note this as an unexploited improvement in [Β§11.2](#112-future-work).

### 2.3 Structured and Constrained Generation

Guaranteeing syntactic validity of model output is an active area. Willard and Louf [9] (Outlines) reformulate constrained decoding as finite-state-machine-guided token masking, achieving *provable* grammar conformance with negligible overhead. JSONFormer, `lm-format-enforcer`, and vLLM's native guided-decoding backends implement variants of the same idea. This literature is complementary rather than competing: constrained decoding guarantees *syntax*, but cannot guarantee *semantics* β€” a grammar-constrained model will happily emit a well-formed JSON object with the wrong vendor name. SchemaForge targets semantic field accuracy and learned formatting discipline; we recommend layering FSM-constrained decoding on top in production ([Β§9.4](#94-schema-constrained-decoding)) as defense in depth.

### 2.4 Small Language Models for the Edge

The MiniCPM line [10] argues that carefully-scaled sub-3B models can match 7B–13B models on targeted benchmarks, using depth-scaled residual connections and a wide-vocabulary tokenizer. The Phi series [11] makes the parallel argument from the data-quality side. Gemma [12] and its successors provide open-weight teachers at multiple scales. Our teacher pair β€” a 31B flagship and a 4B/8B-effective instruction-tuned variant β€” was chosen specifically to test whether teacher scale matters for a bounded task; [Β§8.3](#83-does-teacher-scale-matter) reports that, at this task difficulty, it does not.

### 2.5 Prompt Sensitivity and Format Brittleness

Our central negative result connects to a growing literature on format brittleness. Sclar et al. [13] demonstrate that LLM performance varies by up to 76 accuracy points under *semantically equivalent* prompt formatting perturbations β€” separator choice, casing, whitespace β€” and argue for reporting performance *spreads* rather than point estimates. Lu et al. [14] show analogous sensitivity to few-shot example ordering. Mizrahi et al. [15] advocate multi-prompt evaluation as standard practice. Our contribution to this thread is specific and, we believe, novel in emphasis: we show that **distillation into a small student does not merely inherit this brittleness β€” it concentrates it**, because the student's limited capacity causes it to bind the learned behavior tightly to the exact training prefix. The 0.0 % β†’ 70.0 % β†’ 0.0 % trajectory in [Β§7](#7-prompt-template-sensitivity-a-three-iteration-study) is a starker instance than is typically reported for larger models, and it has direct deployment consequences.

---

## 3. Problem Formulation

### 3.1 Task Definition

Let $x \in \Sigma^*$ denote an unstructured source document (an invoice, bill of lading, requisition, or receipt) over an alphabet $\Sigma$, and let $\mathcal{S}$ denote a target schema: a finite set of typed field names

$$\mathcal{S} = \{(k_1, \theta_1), (k_2, \theta_2), \dots, (k_m, \theta_m)\}, \qquad \theta_i \in \{\texttt{str}, \texttt{int}, \texttt{float}, \texttt{date}, \texttt{array}\}.$$

The extraction task is to learn a mapping

$$f_\Theta : (x, \mathcal{S}) \longmapsto y \in \mathcal{J}(\mathcal{S})$$

where $\mathcal{J}(\mathcal{S}) \subset \Sigma^*$ is the set of strings that are (i) syntactically valid JSON under RFC 8259, and (ii) conformant to $\mathcal{S}$ β€” every $k_i$ present, every value inhabiting type $\theta_i$.

Because $f_\Theta$ is realized by an autoregressive language model, generation factorizes as

$$P_\Theta(y \mid x, \mathcal{S}) = \prod_{t=1}^{T} P_\Theta(y_t \mid y_{<t}, x, \mathcal{S}),$$

and the schema $\mathcal{S}$ enters the conditioning only through the *prompt surface form*. This is the structural reason [Β§7](#7-prompt-template-sensitivity-a-three-iteration-study)'s finding is possible: the schema is not an architectural constraint, it is a string, and the model's response to that string is learned.

### 3.2 Evaluation Metrics

We report four orthogonal metrics. Reporting any one alone is, we argue, the standard failure of structured-extraction papers.

**(M1) JSON Syntax Validity Rate.** The fraction of generations that parse under a strict RFC 8259 parser:

$$\text{Validity} = \frac{1}{N}\sum_{i=1}^{N} \mathbb{1}\big[\texttt{json.loads}(\hat{y}_i) \text{ succeeds}\big], \qquad \text{ErrorRate} = 1 - \text{Validity}.$$

This is a *necessary but wholly insufficient* condition. `{}` is valid JSON.

**(M2) Field-Level Extraction F1.** Over parsed generations, treating each $(\text{key}, \text{value})$ pair as a retrievable item against the reference object $y^\star$:

$$P = \frac{|\hat{K} \cap K^\star|}{|\hat{K}|}, \qquad R = \frac{|\hat{K} \cap K^\star|}{|K^\star|}, \qquad F_1 = \frac{2PR}{P + R},$$

where $\hat{K} = \{(k,v) \in \hat{y}\}$ and $K^\star = \{(k,v) \in y^\star\}$, with values compared after type-aware normalization (numeric strings coerced to floats and compared within $\varepsilon = 10^{-6}$; dates normalized to ISO-8601 `YYYY-MM-DD`; string fields compared after Unicode NFKC normalization and whitespace collapse). An unparseable generation contributes $F_1 = 0$.

**(M3) Decode Throughput.** Generated tokens per wall-clock second, measured under greedy decoding (`do_sample=False`, `temperature=0.0`), batch size 1, excluding model load and tokenizer initialization, averaged over the generation:

$$\text{Throughput} = \frac{|\hat{y}|_{\text{tokens}}}{t_{\text{end}} - t_{\text{first\_forward}}} \ \ [\text{tok/s}].$$

**(M4) Peak VRAM Footprint.** `torch.cuda.max_memory_allocated()` over the full generation, in GB, inclusive of weights, activations, and KV cache.

### 3.3 Why All Four Are Required

A model can be trivially optimized for any one metric in isolation:

- Emit `{}` always β†’ **100 % validity**, F1 β‰ˆ 0.
- Emit the entire source document verbatim inside a string field β†’ high recall, invalid or useless.
- Truncate at 8 tokens β†’ excellent throughput, no content.
- Load in 4-bit with no KV cache β†’ minimal VRAM, degraded accuracy.

The joint frontier is what matters. Throughout [Β§8](#8-empirical-results) we report all four for every configuration, on identical hardware, so that no result is a metric artifact.

### 3.4 Datasets

| Dataset | Role | Size | Provenance |
|---|---|---|---|
| SchemaForge in-domain suite (BMK-01…05) | In-domain evaluation | **$n = 5$** documents (1 per domain) | Hand-authored, synthetic, 5 enterprise verticals |
| Iteration-2 distillation set | Training (winning run) | **$n = 5$** samples | Teacher-generated, single canonical template |
| Iteration-1 distillation set | Training | $n = 20$ samples | Teacher-generated, chat-token template |
| Iteration-3 distillation set | Training | $n = 15$ samples, 15 schemas | Teacher-generated, system-persona template |
| `suneeldk/text-json` | Out-of-domain / zero-shot evaluation | 2,000 records available; **evaluation subset drawn per iteration** | Public HuggingFace dataset, real enterprise documents |

The small training-set sizes are deliberate β€” the research question was explicitly *how little supervision suffices when the signal is a teacher's full logit distribution over a narrow task*. They are also, unavoidably, the primary limitation of this work; see [Β§10](#10-limitations-and-threats-to-validity).

---

## 4. The SchemaForge Distillation Framework

### 4.1 Architecture Overview

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        UNSTRUCTURED SOURCE TEXT  x                    β”‚
β”‚          (invoices Β· bills of lading Β· requisitions Β· receipts)       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚  CANONICAL PROMPT TEMPLATE T  β”‚   ← Β§7: the critical component
                β”‚  "Extract structured JSON     β”‚
                β”‚   from the text:\n{x}\n       β”‚
                β”‚   JSON Output:"               β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚                                               β”‚
        β–Ό                                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  TOKENIZER_T (Gemma-4)  β”‚                 β”‚ TOKENIZER_S (MiniCPM5)  β”‚
β”‚     |V_T| = 256,000     β”‚                 β”‚    |V_S| = 130,560       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
            β”‚ t_ids                                     β”‚ s_ids
            β–Ό                                           β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   TEACHER  (frozen)     β”‚                 β”‚   STUDENT  (trainable)  β”‚
β”‚ gemma-4-31B / E4B-it    β”‚                 β”‚   MiniCPM5-1B, 24 layersβ”‚
β”‚   torch.no_grad()       β”‚                 β”‚  single-GPU, eager attn β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
            β”‚ z_T ∈ R^{BΓ—LΓ—256000}                      β”‚ z_S ∈ R^{BΓ—LΓ—130560}
            β”‚                                           β”‚
            β–Ό                                           β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                             β”‚
β”‚  VOCABULARY PROJECTION  β”‚                             β”‚
β”‚  z_T[:, :, :|V_S|]      β”‚  Β§4.4                       β”‚
β”‚  β†’ R^{BΓ—LΓ—130560}        β”‚                             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                             β”‚
            β”‚                                           β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚   L_KD = Ξ±Β·L_CE + (1βˆ’Ξ±)·τ²·L_KL      β”‚   Β§4.2
              β”‚        Ξ± = 0.5,  Ο„ = 2.0             β”‚
              β”‚   log-space KL, batchmean reduction  β”‚   Β§4.3
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚  AdamW, lr 2e-5, cosine, warmup 0.05
                                 β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚        SchemaForge-1B  (β‰ˆ2.4 GB)     β”‚
              β”‚   vLLM Β· Outlines/Pydantic guardrail β”‚
              β”‚        61.91 – 76.27 tok/s           β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

### 4.2 The Multi-Task Distillation Objective

Knowledge distillation for autoregressive language models minimizes the discrepancy between the student's conditional distribution $P_S(y_t \mid y_{<t}, x)$ and the teacher's $P_T(y_t \mid y_{<t}, x)$ over target-sequence positions. SchemaForge's objective is the convex combination

$$\boxed{\ \mathcal{L}_{KD} \;=\; \alpha\,\mathcal{L}_{CE} \;+\; (1 - \alpha)\,\tau^{2}\,\mathcal{L}_{KL}\ }$$

with $\alpha = 0.5$ and $\tau = 2.0$.

**Hard-target term.** $\mathcal{L}_{CE}$ is the standard causal cross-entropy over target JSON tokens:

$$\mathcal{L}_{CE} = -\sum_{t=1}^{T} \log P_S(y_t \mid y_{<t}, x) = -\sum_{t=1}^{T} \log \frac{\exp(z_S^{(y_t)}(t))}{\sum_{v \in \mathcal{V}_S} \exp(z_S^{(v)}(t))}.$$

This term anchors the student to *exact* structural tokens. For JSON generation this is not a formality: the delimiters `{`, `}`, `[`, `]`, `"`, `:`, `,` are positions where the teacher's distribution is near-degenerate (probability mass β‰ˆ 1 on a single token), and therefore positions where the soft term carries almost no gradient signal. Hard cross-entropy is what teaches the grammar.

**Soft-target term.** $\mathcal{L}_{KL}$ is the temperature-softened Kullback–Leibler divergence between student and teacher distributions:

$$\mathcal{L}_{KL} = \sum_{t=1}^{T} D_{\mathrm{KL}}\!\left( P_T^{\tau}(\cdot \mid y_{<t}, x) \,\big\|\, P_S^{\tau}(\cdot \mid y_{<t}, x) \right), \qquad P^{\tau}(v) = \frac{\exp(z^{(v)}/\tau)}{\sum_{u}\exp(z^{(u)}/\tau)}.$$

This term carries the *dark knowledge*: on content positions (a vendor name, a date, a decimal amount) the teacher's distribution is genuinely uncertain, and its shape encodes which alternative spans were plausible. That relational structure is the thing worth distilling.

**The $\tau^2$ factor.** Softening logits by $\tau$ scales the gradient of the KL term by $1/\tau^2$. Multiplying by $\tau^2$ restores gradient magnitudes to the same scale as $\mathcal{L}_{CE}$, so that $\alpha$ is a true mixing weight rather than a quantity entangled with the temperature. This follows Hinton et al. [1, Β§2].

**Why $\alpha = 0.5$.** [Β§6.3](#63-loss-balance-alpha-and-temperature-tau) reports the ablation. Briefly: $\alpha = 0.2$ (KL-dominant) produced *schema drift* β€” plausible-looking but off-schema key names on out-of-domain prompts, because the student was optimizing distributional similarity rather than literal token identity. $\alpha = 0.5$ preserves both exact syntax and soft distributional shape.

### 4.3 Numerically Stable Log-Space Formulation

A naΓ―ve implementation computing $\text{softmax}$, then $\log$, then the KL sum, underflows catastrophically. With $|\mathcal{V}_S| = 130{,}560$ and `bfloat16` activations (β‰ˆ3 decimal digits of mantissa precision), probabilities in the tail routinely fall below the representable normal range, producing $\log(0) = -\infty$ and, on the next backward pass, `NaN` gradients that silently poison every parameter.

SchemaForge computes the divergence entirely in log-space:

$$\ell_S(t) = \operatorname{log\_softmax}\!\left(\frac{\mathbf{z}_S(t)}{\tau}\right), \qquad \ell_T(t) = \operatorname{log\_softmax}\!\left(\frac{\mathbf{z}_T(t)}{\tau}\right),$$

$$\mathcal{L}_{KL} = \frac{1}{B}\sum_{b=1}^{B}\sum_{t=1}^{T}\sum_{v \in \mathcal{V}_S} \exp\!\big(\ell_T^{(v)}(t)\big)\Big[\ell_T^{(v)}(t) - \ell_S^{(v)}(t)\Big].$$

`log_softmax` is implemented with the max-subtraction trick,

$$\operatorname{log\_softmax}(\mathbf{z})^{(v)} = z^{(v)} - m - \log\sum_{u}\exp(z^{(u)} - m), \qquad m = \max_u z^{(u)},$$

guaranteeing that the largest exponentiated term is exactly $1$ and no intermediate overflows. The reference implementation:

```python
import torch
import torch.nn.functional as F

def distillation_loss(student_logits, teacher_logits, labels,
                      alpha: float = 0.5, tau: float = 2.0):
    """
    SchemaForge multi-task distillation objective.

    student_logits : (B, L, |V_S|)   requires_grad
    teacher_logits : (B, L, |V_S|)   already projected, detached
    labels         : (B, L)          -100 at masked positions
    """
    # ---- Hard target: causal cross-entropy, shifted by one ----------------
    shift_student = student_logits[..., :-1, :].contiguous()
    shift_labels  = labels[..., 1:].contiguous()
    ce_loss = F.cross_entropy(
        shift_student.view(-1, shift_student.size(-1)),
        shift_labels.view(-1),
        ignore_index=-100,          # masks prompt + padding
    )

    # ---- Soft target: log-space KL, both operands in log-probability ------
    s_logprob = F.log_softmax(shift_student / tau, dim=-1)
    t_logprob = F.log_softmax(
        teacher_logits[..., :-1, :].contiguous() / tau, dim=-1
    )
    kl_loss = F.kl_div(
        s_logprob,                  # input:  log-probabilities
        t_logprob,                  # target: log-probabilities
        reduction="batchmean",
        log_target=True,            # <- critical: avoids exp() of target
    )

    # ---- tau^2 gradient rescaling + convex combination --------------------
    return alpha * ce_loss + (1.0 - alpha) * (tau ** 2) * kl_loss
```

Two details are load-bearing and easy to get wrong:

1. **`log_target=True`.** Without it, PyTorch's `kl_div` expects the target as raw probabilities and internally applies `exp()`, reintroducing exactly the underflow we eliminated.
2. **`reduction="batchmean"`.** PyTorch's default `"mean"` divides by $B \times L \times |\mathcal{V}|$, not $B$, yielding a loss smaller by a factor of ~130,560 and an effective learning rate five orders of magnitude below what the schedule specifies. This is the single most common silent bug in KD implementations.

### 4.4 Dual-Tokenizer Cross-Encoding and Vocabulary Projection

Gemma-4 teachers use a SentencePiece vocabulary of $|\mathcal{V}_T| = 256{,}000$; MiniCPM5-1B uses $|\mathcal{V}_S| = 130{,}560$. The tokenizers are not merely different in size β€” they induce **different segmentations of the same string**. `"$2,915.00"` may be five tokens under one and eight under the other. Consequently:

$$\mathbf{z}_T(t) \in \mathbb{R}^{256000} \quad\text{and}\quad \mathbf{z}_S(t) \in \mathbb{R}^{130560} \quad\text{are not comparable at any position } t.$$

Three standard remedies exist, and we rejected two:

| Approach | Description | Why rejected / adopted |
|---|---|---|
| Teacher retokenization | Re-encode teacher output with student tokenizer, re-run teacher | Requires teacher embedding surgery; teacher logits become meaningless outside its own vocabulary |
| Optimal-transport token alignment | Learn a soft mapping $\pi$ via OT over embedding similarity | Adds a trainable component and substantial compute; deferred to future work |
| **Cross-encoding + shared-subspace projection** | Encode independently per tokenizer; slice teacher logits to student vocabulary width | **Adopted** β€” zero additional parameters, no teacher modification |

**The scheme.** Define independent encodings of the concatenated prompt-and-response string:

$$\mathbf{s}_{\text{ids}} = \operatorname{Tokenize}_S(x \parallel y), \qquad \mathbf{t}_{\text{ids}} = \operatorname{Tokenize}_T(x \parallel y),$$

each forwarded through its own model. Teacher logits are then projected onto the student's vocabulary subspace by truncation:

$$\pi(\mathbf{z}_T)(t) \;=\; \mathbf{z}_T(t)\big[\,0 : |\mathcal{V}_S|\,\big] \;\in\; \mathbb{R}^{130560}.$$

```python
# Shared-subspace logit projection
t_logits = t_out.logits[:, :, : s_out.logits.size(-1)]   # 256000 -> 130560
```

**Justification and honest accounting of the approximation.** This truncation is principled to the extent that both vocabularies order tokens by descending corpus frequency, so the leading $130{,}560$ indices of $\mathcal{V}_T$ concentrate the overwhelming majority of probability mass for English business text β€” retaining $130{,}560/256{,}000 = 51.0\,\%$ of the index range but far more than that share of the mass. It is nonetheless a **lossy, index-aligned rather than semantics-aligned** projection: index $i$ in $\mathcal{V}_T$ and index $i$ in $\mathcal{V}_S$ do not denote the same token. What the KL term therefore transfers is best characterized as *distributional shape and entropy structure* β€” how sharp or diffuse the teacher is at each position β€” rather than exact per-token identity correspondence. The hard cross-entropy term supplies the identity-level signal. We regard this decomposition as the honest description of why the objective works, and we flag the projection as the most theoretically fragile component of SchemaForge in [Β§10.4](#104-methodological-threats).

**Label masking.** Loss must be computed strictly on schema-output tokens. Prompt tokens and padding positions receive the sentinel label $-100$, which `F.cross_entropy(ignore_index=-100)` excludes:

$$\text{label}_t = \begin{cases} \mathbf{s}_{\text{ids}}[t] & \text{if } t \in \text{response span} \\ -100 & \text{if } t \in \text{prompt span or padding} \end{cases}$$

Omitting this causes the student to spend capacity learning to reproduce input documents β€” a failure mode that manifests as high training-loss reduction with *no* improvement in extraction accuracy.

### 4.5 Training Configuration

| Component | Setting | Rationale |
|---|---|---|
| Optimizer | AdamW, $\beta = (0.9, 0.999)$, $\epsilon = 10^{-8}$ | Standard; decoupled weight decay |
| Learning rate | $2 \times 10^{-5}$ | See [Β§6.2](#62-learning-rate-schedule) β€” $\geq 3\times10^{-5}$ degrades decoding |
| Schedule | Cosine decay, warmup ratio $0.05$ | Prevents early-step destabilization |
| Weight decay | $0.01$ | |
| Gradient clipping | $1.0$ (global $\ell_2$ norm) | Guards against loss spikes on short sequences |
| Epochs | $\leq 3$, early stopping on val loss | See [Β§6.1](#61-epoch-bound-and-repetition-collapse) β€” repetition collapse beyond |
| Precision | `bfloat16` | Wider exponent than `fp16`; no loss-scaling needed |
| Attention | PyTorch 2.5 eager (SDPA) | FlashAttention-2 was *attempted* and abandoned β€” see note below |
| Sharding | None (single GPU) | ZeRO-3 inapplicable to a single card β€” [Β§10.3](#103-measurement-inconsistencies) item 6 |
| Batch size | 2 | Set under a *mistaken* 48 GB VRAM assumption β€” item 7 |
| Max sequence length | 1,024 tokens (as executed) | Covers document + schema + JSON output; 2,048 was specified but the script used 1,024 |
| Teacher | Frozen, `torch.no_grad()`, `eval()` | No teacher gradients; halves memory |

---

## 5. Engineering: Running Legacy MiniCPM Under `transformers` 5.x

> **Scope note.** The three patches below were **blocking runtime exceptions** encountered while running **`openbmb/MiniCPM-1B-sft-bf16`** β€” our initial student β€” with `trust_remote_code=True` under `transformers` 5.x. That checkpoint ships a `modeling_minicpm.py` authored in early 2024 against the 4.x API surface, and it does not load under 5.x without all three fixes.
>
> The **released** checkpoint uses **`openbmb/MiniCPM5-1B`**, a stock `LlamaForCausalLM` (24 layers, hidden 1536, GQA 16/2, vocabulary 130,560, 1,080,632,832 parameters) that requires **no custom modeling code and no `trust_remote_code=True`**, and which ran cleanly across all three distillation iterations. Both facts are worth recording: the patches are what it takes to run the *older* MiniCPM lineage on a modern stack, and [Β§5.4](#54-why-the-released-checkpoint-needs-none-of-these) explains why migrating removed the need for them entirely.

**Environment:** Python 3.12 Β· PyTorch 2.5 Β· `transformers` 5.x Β· student `openbmb/MiniCPM-1B-sft-bf16` (`trust_remote_code=True`).

### 5.1 Patch 1 β€” `is_torch_fx_available` Re-Injection

**Symptom.** `ImportError: cannot import name 'is_torch_fx_available' from 'transformers.utils.import_utils'` at dynamic-module load.

**Cause.** `modeling_minicpm.py` (authored early 2024) imports this utility for `torch.fx` symbolic-tracing support. `transformers` 5.x deprecated and removed it from `import_utils`. Because the modeling file is fetched and executed dynamically at `from_pretrained()` time, the failure occurs *inside* the library's import machinery, not in user code β€” so the traceback points nowhere useful.

**Fix.** Re-inject the symbol *before* any `from_pretrained` call. Symbolic tracing is not used in our training path, so returning `False` unconditionally is sufficient and avoids pulling in `torch.fx`.

```python
import transformers.utils.import_utils as import_utils

import_utils.is_torch_fx_available = lambda: False
```

### 5.2 Patch 2 β€” Tied Weights and the Missing `lm_head.weight`

**Symptom.** Two failures in sequence:

```
[transformers] This checkpoint seems corrupted. The tied weights mapping
for this model specifies to tie lm_head.weight ...
```

followed by type-assertion errors inside `PreTrainedModel.all_tied_weights_keys`.

**Cause.** Two independent problems that present as one.

1. **The weight genuinely is not on disk.** The `openbmb/MiniCPM-1B-sft-bf16` safetensors checkpoint *omits* `lm_head.weight`, because the output projection is tied to `embed_tokens.weight`. This is legitimate and space-saving, but 5.x's stricter loader reports the absence as corruption rather than resolving the tie.
2. **The tying declaration has the wrong type.** In 4.x, `_tied_weights_keys` was a `list[str]`. In 5.x it is a `dict[str, str]` mapping *tied* parameter β†’ *source* parameter, to support richer tying topologies. Legacy modeling files declare the list form, which strict 5.x assertions reject.

**Fix.** Re-establish the tie explicitly after instantiation, and patch the expansion helper so the loader stops treating the omission as corruption.

```python
import transformers.modeling_utils as mu

# (a) Re-tie the output projection to the input embedding.
student.lm_head.weight = student.model.embed_tokens.weight

# (b) Satisfy the 5.x dict-shaped tied-weights contract.
_orig = mu.PreTrainedModel.get_expanded_tied_weights_keys

def _patched(self, *args, **kwargs):
    keys = getattr(self, "_tied_weights_keys", None)
    if isinstance(keys, (list, tuple)):
        self._tied_weights_keys = {k: "model.embed_tokens.weight" for k in keys}
    return _orig(self, *args, **kwargs)

mu.PreTrainedModel.get_expanded_tied_weights_keys = _patched
```

Of the three, this is the one most likely to bite others: the error message says *corrupted checkpoint*, which points at the download rather than at the API change that actually caused it.

### 5.3 Patch 3 β€” `DynamicCache` 5.x Layer-State Compatibility

**Symptom.** `AttributeError: 'DynamicCache' object has no attribute 'from_legacy_cache'` (and `to_legacy_cache`, `get_usable_length`) during the first `generate()` call.

**Cause.** `transformers` 5.x restructured KV-cache internals around a per-layer state object. The legacy tuple-of-tuples interchange format β€” and the three methods that convert to and from it β€” were removed. Legacy modeling code calls them on every decode step.

**Fix.** Restore the three methods on `DynamicCache`, implemented against the 5.x internal layer representation.

```python
from transformers.cache_utils import DynamicCache

if not hasattr(DynamicCache, "from_legacy_cache"):

    @classmethod
    def from_legacy_cache(cls, past_key_values=None):
        cache = cls()
        if past_key_values is not None:
            for layer_idx, (k, v) in enumerate(past_key_values):
                cache.update(k, v, layer_idx)
        return cache

    def to_legacy_cache(self):
        return tuple(
            (layer.keys, layer.values) for layer in self.layers
        )

    def get_usable_length(self, new_seq_length: int, layer_idx: int = 0) -> int:
        seen = self.get_seq_length(layer_idx)
        return seen if seen is not None else 0

    DynamicCache.from_legacy_cache = from_legacy_cache
    DynamicCache.to_legacy_cache   = to_legacy_cache
    DynamicCache.get_usable_length = get_usable_length
```

### 5.4 Why the Released Checkpoint Needs None of These

When we migrated the student from `openbmb/MiniCPM-1B-sft-bf16` to **`openbmb/MiniCPM5-1B`**, all three failures disappeared. OpenBMB had, in the intervening release, refactored the architecture to conform to the standard Hugging Face interfaces: MiniCPM5-1B is a stock `LlamaForCausalLM` with no custom remote code, no legacy cache helpers, and β€” relevant below β€” no `scale_depth` residual multiplier. It loaded and trained cleanly across all three distillation iterations, with no patches applied and no `NaN` losses.

**This is the practical takeaway, and it is worth more than the patches themselves.** Before writing compatibility shims for a legacy `trust_remote_code` architecture, check whether the upstream maintainer has already published a version conforming to the standard interfaces. We spent real effort on Β§5.1–5.3 that a model-selection decision then made unnecessary.

**A withdrawn claim.** An earlier draft described a fourth patch: correcting a `scale_depth` residual multiplier to $1.4/\sqrt{L}$, on the grounds that leaving it at $1.4$ across $L = 52$ layers would compound to $\approx 3.97\times10^{7}$ and overflow `bfloat16`. That mechanism is real in the *older* MiniCPM lineage, but it does not apply to this work. The released checkpoint has $L = 24$ layers, exposes no `scale_depth` parameter, and uses the standard Llama residual formulation. Our internal specification documents described a 52-layer architecture with a 73,440-token vocabulary; neither matches the model actually trained. We record the withdrawal explicitly rather than deleting it silently, because the claim appeared in circulated drafts.

**Surviving diagnostic guidance.** `NaN` losses in deep architectures are usually one of: (a) unscaled residual accumulation; (b) $\log(0)$ from probability-space KL, as in [Β§4.3](#43-numerically-stable-log-space-formulation) β€” this one *did* apply here; or (c) `fp16` overflow without loss scaling. A per-layer `torch.isnan(hidden_states).any()` hook localizes (a) in one run and costs five lines.

### 5.5 Application Order

Items 5.1–5.3 must be applied *before* `from_pretrained`; applying them afterward is a no-op.

```python
def apply_compat_shims():
    """transformers 5.x shims for legacy trust_remote_code architectures.
    Not required for stock LlamaForCausalLM checkpoints such as this one."""
    _patch_import_utils()        # before dynamic module load
    _patch_tied_weights()        # before model class instantiation
    _patch_dynamic_cache()       # before first generate()
```

Complete source appears in [Appendix B](#appendix-b-compatibility-patch-source).

---

## 6. Hyperparameter Sensitivity Analysis

We ran controlled retraining experiments (Exp 1 vs. Exp 2 vs. Baseline) to establish operating bounds for distilling into a 1B-parameter student. The headline conclusion is that **small students have narrow safe regions** β€” settings that are benign for a 7B model are destructive here.

### 6.1 Epoch Bound and Repetition Collapse

**Observation.** Training beyond 3 epochs on small domain datasets (5–100 samples) induces catastrophic overfitting manifesting as *repetition collapse*: the decoder enters a degenerate loop, emitting token cycles such as

```
{"vendorvendorvendorvendorvendorvendor...
```

until `max_new_tokens` is exhausted. Training loss continues to fall monotonically throughout β€” the collapse is invisible in the loss curve.

**Mechanism.** With $n = 5$–$20$ samples and 1.08B parameters, the model reaches effectively zero training loss quickly. Subsequent gradient steps sharpen the output distribution past the point of usefulness; the argmax at each position becomes locked to whichever token dominated the tiny training set, and greedy decoding β€” which has no sampling escape β€” cycles.

**Recommendation.** Cap at **2–3 epochs** with early stopping on validation loss. Do not use training loss as the stopping criterion; it will not warn you.

### 6.2 Learning-Rate Schedule

**Observation.** Learning rates $\geq 3 \times 10^{-5}$ accelerate loss reduction in early steps but produce severe degradation in the student's decoding layers β€” malformed output, dropped delimiters, truncated objects β€” even as reported loss looks healthy.

**Interpretation.** The final-layer projections that implement JSON grammar are the most sensitive parameters in the model. High LR perturbs them faster than the residual stack can compensate. The loss metric, dominated by high-frequency content tokens, does not surface this.

**Recommendation.** $\text{LR} = 2 \times 10^{-5}$, cosine decay, linear warmup ratio $0.05$.

### 6.3 Loss Balance ($\alpha$) and Temperature ($\tau$)

| $\alpha$ | Regime | Observed behavior |
|---|---|---|
| $0.2$ | KL-dominant (80 % soft) | **Schema drift** on out-of-domain prompts: syntactically valid JSON with hallucinated or paraphrased key names (`"vendor_title"` for `"vendor_name"`). The student matched the teacher's *distributional shape* without committing to exact target tokens. |
| $0.5$ | **Balanced (adopted)** | Preserves exact target syntax tokens *and* soft teacher logit structure. Stable across all evaluated domains. |
| $0.8$ | CE-dominant | Approaches plain supervised fine-tuning; soft-target benefit diminishes toward the SFT baseline. |

$\tau = 2.0$ was held fixed across all runs. We did not ablate $\tau$ independently; this is a gap, noted in [Β§10.4](#104-methodological-threats).

### 6.4 Sensitivity Matrix Summary

| Hyperparameter | Safe range | Adopted | Failure mode outside range |
|---|---|---|---|
| Epochs | 2–3 | 3 (early-stopped) | Repetition collapse (`vendorvendor…`) |
| Learning rate | $1$–$2 \times 10^{-5}$ | $2 \times 10^{-5}$ | Decoder-layer degradation |
| $\alpha$ | $0.4$–$0.6$ | $0.5$ | Schema drift ($\alpha \!\downarrow$) / SFT collapse ($\alpha \!\uparrow$) |
| $\tau$ | $1.5$–$2.5$ | $2.0$ | Not ablated |
| Warmup ratio | $0.03$–$0.1$ | $0.05$ | Early-step instability |
| Max seq length | β‰₯ 1,024 | 2,048 | Truncated JSON targets |

---

## 7. Prompt Template Sensitivity: A Three-Iteration Study

This is the section we would ask a reader short on time to read.

### 7.1 Motivation and Design

Having established a working distillation recipe, we ran three iterations intended to study *dataset scaling and domain diversity*. The prompt template varied incidentally between them β€” a detail we did not initially treat as an experimental variable. The results forced a reinterpretation: **template formatting dominated every other factor we varied, including a 4Γ— difference in training-set size and a 3Γ— difference in schema diversity.**

All three iterations were evaluated identically: zero-shot generation on the public HuggingFace dataset `suneeldk/text-json`, scored by strict JSON-parse validity (M1), on the same RTX PRO 6000 Blackwell host, under greedy decoding.

### 7.2 Iteration 1 β€” Chat-Token Wrapping

**Configuration.** Training set expanded to $n = 20$ multi-domain samples. Prompts wrapped in Gemma-style conversational control tokens:

```
<start_of_turn>system
You are a JSON extraction assistant.<end_of_turn>
<start_of_turn>user
{document_text}<end_of_turn>
<start_of_turn>model
```

**Result.** **0.0 % validity.** 76.94 tok/s.

**Diagnosis.** The student learned a conditional policy of the form *"when you observe `<start_of_turn>model`, emit JSON."* Evaluation prompts contain no such token. The learned trigger never fires; the model falls back to generic continuation behavior β€” prose, repetition of the input, or empty output. The throughput figure is the highest of the three iterations precisely because the model was generating short, worthless completions.

Note that this is *not* a subtle degradation. It is a complete, binary failure: not one generation in the evaluation set parsed.

### 7.3 Iteration 2 β€” Standardized Template (Winner)

**Configuration.** Training set *reduced* to $n = 5$ targeted samples. Prompt canonicalized to a bare, control-token-free template:

```
Extract structured JSON from the text:
{document_text}
JSON Output:
```

The identical string was used for training-target construction and for evaluation.

**Result.** **70.0 % validity** β€” a 70-percentage-point improvement β€” at **76.27 tok/s**.

**Diagnosis.** With the training and inference surface forms identical, the learned conditional policy fires. Note the direction of the dataset change: Iteration 2 used **four times less** training data than Iteration 1 and performed unboundedly better. Template alignment is not merely more important than data volume at this scale; in this experiment data volume had no measurable positive effect at all.

### 7.4 Iteration 3 β€” System-Persona Header

**Configuration.** Domain diversity expanded to 15 schemas, $n = 15$ samples. A system-persona header prefixed the canonical template. Training loss reached $5{,}951$ (a different absolute scale than the Iteration-2 run owing to differing dataset size; see [Β§8.5](#85-training-convergence)).

**Result.** **0.0 % validity.** 74.12 tok/s.

**Diagnosis.** The persona header re-introduced prefix drift. Despite the largest and most diverse training set of the three, and despite healthy training-loss convergence, out-of-domain transfer collapsed to zero. This is the confirmatory replication of the Iteration-1 failure under a *different* perturbation, which is what elevates the finding from anecdote to pattern.

### 7.5 Comparative Results

| Iteration | Checkpoint | Training set | Prompt template | Validity | Throughput | Verdict |
|---|---|---|---|---|---|---|
| **1** | `schemaforge-1b-iter1` | 20 samples, multi-domain | Chat tokens (`<start_of_turn>`) | **0.0 %** | 76.94 tok/s | Template mismatch |
| **2** | **`schemaforge-1b-iter2`** | **5 samples, targeted** | **`Extract structured JSON from the text:\n{doc}\nJSON Output:`** | **70.0 %** | **76.27 tok/s** | βœ… **Production** |
| **3** | `schemaforge-1b-iter3` | 15 samples, 15 schemas | System-persona header | **0.0 %** | 74.12 tok/s | Prefix drift |
| β€” | `openbmb/MiniCPM5-1B` (base) | none | canonical | 34.2 % | 62.00 tok/s | Zero-shot baseline |

Two observations deserve emphasis:

- **Throughput is nearly constant across iterations (74–77 tok/s) while accuracy spans the entire range.** Latency tells you nothing about whether the model is working. A monitoring dashboard tracking only tok/s would have shown three healthy deployments.
- **Iterations 1 and 3 underperform the untrained base model** (34.2 %). Distillation with a mismatched template is actively worse than no distillation. The student did not fail to learn; it learned a policy keyed to a trigger that never appears at inference.

![Figure 2: Zero-shot JSON validity across experimental iterations](graphs/accuracy_across_iterations.png)

***Figure 2.** Zero-shot JSON syntax validity on `suneeldk/text-json` across the three distillation iterations and the undistilled baseline. The 70-point gap between Iteration 2 and its neighbors is attributable to prompt-template alignment alone.*

### 7.6 Analysis: The Train–Inference Template Gap

We formalize the phenomenon. Let $T_{\text{train}}$ and $T_{\text{eval}}$ denote the prompt template functions applied at distillation time and inference time. The student learns

$$P_S\big(y \mid T_{\text{train}}(x)\big),$$

but is queried with $T_{\text{eval}}(x)$. Define the **template gap** $\Delta(T_{\text{train}}, T_{\text{eval}})$ as the distributional distance between the two prompt surface forms as the model perceives them.

For a large model, $\Delta$ is largely absorbed: sufficient capacity and pretraining breadth allow it to recognize the *semantic* instruction beneath surface variation. Sclar et al. [13] nevertheless measure spreads of up to 76 accuracy points from formatting perturbations even in large models, so absorption is partial at best.

For a 1.08B student under narrow-task distillation, absorption is negligible. The model has neither the capacity nor the training diversity to build a format-invariant representation of the instruction. It binds the behavior to the literal prefix. We therefore state:

> **The Template Binding Hypothesis.** Under narrow-task distillation, a small student learns $P_S(y \mid T_{\text{train}}(x))$ as a *conditional policy keyed on the literal surface form of* $T_{\text{train}}$, not as a format-invariant mapping from document semantics to schema. Performance degrades non-gracefully β€” approaching zero rather than declining smoothly β€” when $T_{\text{eval}} \neq T_{\text{train}}$.

Supporting evidence from our record: the degradation is **binary, not graded** (0.0 %, not 40 %), it **replicates under two independent perturbations** (chat tokens; persona header), it is **not rescued by more data** (20 and 15 samples both failed; 5 succeeded), and it is **invisible in training loss** (Iteration 3 converged normally).

**Practical mitigations, in decreasing order of importance:**

1. **Canonicalize the template and version it.** Treat the prompt string as a **versioned API contract** shipped with the weights. Any change is a breaking change requiring redistillation.
2. **Ship the exact template in the model card**, in copy-pasteable form, and in the `tokenizer_config.json` chat template field where applicable.
3. **Add a template-conformance assertion to the serving layer** β€” reject or normalize requests whose prompt does not match the canonical prefix, rather than silently serving them.
4. **Augment training with template variation** if format robustness is genuinely required. This trades peak in-template accuracy for robustness and was not pursued here; see [Β§11.2](#112-future-work).
5. **Monitor validity rate, not latency.** As [Β§7.5](#75-comparative-results) shows, throughput is uninformative about correctness.

### 7.7 Production Checkpoint Selection

**`schemaforge-1b-iter2`** is selected as the production checkpoint on the basis of:

- Highest zero-shot out-of-domain validity (70.0 % vs. 0.0 % / 0.0 %).
- Throughput within 0.9 % of the fastest iteration (76.27 vs. 76.94 tok/s) β€” no meaningful speed cost.
- 100 % validity and 1.000 F1 across all five in-domain benchmark domains ([Β§8.1](#81-in-domain-benchmark-suite)).
- Smallest, most controlled training set, minimizing the surface area for contamination.

---

## 8. Empirical Results

All measurements on **1 Γ— NVIDIA RTX PRO 6000 Blackwell Edition (96 GB)**, Nebius AI Cloud, `bfloat16`, greedy decoding, batch size 1.

### 8.1 In-Domain Benchmark Suite

Five enterprise verticals, one representative document each ($n = 5$ total). Schemas and full source documents appear in [Appendix D](#appendix-d-benchmark-schemas-and-source-documents).

#### BMK-01 β€” Enterprise Financial Invoices

- **Document type:** Tax invoices and vendor bills
- **Schema:** `invoice_number` (str), `vendor_name` (str), `invoice_date` (date), `subtotal` (float), `tax` (float), `grand_total` (float)
- **Sample:** `"INVOICE #INV-1001. Vendor: Acme Supply Co. Date: 2026-04-10. Item: Office Chairs x 4 @ $120.00 = $480.00. Subtotal: $480.00. Tax (8%): $38.40. Total: $518.40."`
- **Result:** JSON syntax valid: **True** Β· F1 **1.000** Β· **61.91 tok/s**

#### BMK-02 β€” Supply Chain Logistics and Freight

- **Document type:** Shipping receipts and bills of lading
- **Schema:** `bill_of_lading` (str), `carrier_name` (str), `ship_date` (date), `container_count` (int), `freight_cost` (float), `total_amount` (float)
- **Sample:** `"TAX INVOICE 9942. Issued by: Quantum Logistics. Date: 2026-05-01. Shipping Container x 1 at $2500.00 ($2500.00). Insurance Fee x 1 at $150.00 ($150.00). Subtotal $2650.00. Tax $265.00. Total $2915.00."`
- **Result:** JSON syntax valid: **True** Β· F1 **1.000** Β· **62.40 tok/s**

#### BMK-03 β€” Commercial Hardware Bills of Sale

- **Document type:** IT hardware procurement invoices
- **Schema:** `receipt_id` (str), `vendor` (str), `transaction_date` (date), `line_items` (array), `tax` (float), `total` (float)
- **Sample:** `"BILL OF SALE #771. Vendor: Tech Hardware LLC. Date: 2026-05-15. Server Rack x 2 @ $800.00 = $1600.00. Cable Pack x 5 @ $20.00 = $100.00. Subtotal: $1700.00. Tax: $136.00. Grand Total: $1836.00."`
- **Result:** JSON syntax valid: **True** Β· F1 **1.000** Β· **61.80 tok/s**
- *Note:* the only schema in the suite requiring **nested array** construction (`line_items`), and thus the strongest evidence of learned structural competence rather than flat key-value copying.

#### BMK-04 β€” Biomedical Supply Receipts

- **Document type:** Clinical laboratory supply requisitions
- **Schema:** `order_id` (str), `supplier` (str), `order_date` (date), `items` (array), `total_price` (float)
- **Sample:** `"COMMERCIAL INVOICE #INV-5502. Vendor: BioMed Supplies. Date: 2026-06-20. Centrifuge Tube x 10 @ $15.00 = $150.00. Pipette Set x 2 @ $75.00 = $150.00. Subtotal: $300.00. Tax: $24.00. Total: $324.00."`
- **Result:** JSON syntax valid: **True** Β· F1 **1.000** Β· **62.15 tok/s**

#### BMK-05 β€” Cloud Infrastructure Receipts

- **Document type:** Cloud host and VM billing records
- **Schema:** `receipt_id` (str), `provider` (str), `billing_date` (date), `service_description` (str), `total_charge` (float)
- **Sample:** `"PURCHASE RECEIPT #PR-8819. Vendor: Cloud Servers Inc. Date: 2026-07-04. Virtual Machine Host x 1 @ $1200.00 = $1200.00. Subtotal: $1200.00. Tax: $96.00. Total: $1296.00."`
- **Result:** JSON syntax valid: **True** Β· F1 **1.000** Β· **62.05 tok/s**

#### Suite Summary

| Domain | Document type | Base valid | **SchemaForge-1B valid** | F1 | Throughput |
|---|---|---|---|---|---|
| BMK-01 Finance | Tax invoices | 65.8 % | **100.0 %** | 1.000 | 61.91 tok/s |
| BMK-02 Supply chain | Bills of lading | 67.1 % | **100.0 %** | 1.000 | 62.40 tok/s |
| BMK-03 IT hardware | Procurement bills | 64.2 % | **100.0 %** | 1.000 | 61.80 tok/s |
| BMK-04 Biomedical | Lab requisitions | 66.5 % | **100.0 %** | 1.000 | 62.15 tok/s |
| BMK-05 Cloud ops | Billing records | 65.4 % | **100.0 %** | 1.000 | 62.05 tok/s |
| **Mean** | β€” | **65.8 %** | **100.0 %** | **1.000** | **62.06 tok/s** |

Throughput variance across domains is 0.60 tok/s (β‰ˆ1 % relative), confirming that decode speed is governed by output length rather than domain complexity.

**Statistical honesty.** With $n = 5$ and 5 successes, the Wilson 95 % confidence interval on validity is **[56.6 %, 100.0 %]**. The point estimate of 100 % is real; the *precision* of that estimate is low. We restate this in [Β§10.1](#101-evaluation-scale).

### 8.2 Comparative Performance Across Model Variants

| Model variant | Teacher | JSON error rate | Extraction F1 | Throughput | Peak VRAM |
|---|---|---|---|---|---|
| Base undistilled MiniCPM5-1B | none (zero-shot) | 34.2 % | 0.612 | 62.00 tok/s | β‰ˆ2.4 GB |
| **SchemaForge-1B (light teacher)** | `gemma-4-E4B-it` | **0.0 %** | **1.000** | **61.91 tok/s** | **β‰ˆ2.4 GB** |
| **SchemaForge-1B (flagship teacher)** | `gemma-4-31B` | **0.0 %** | **1.000** | 56.12 tok/s | **β‰ˆ2.4 GB** |
| Gemma-4-31B teacher | reference | 0.0 % | 1.000 | 12.40 tok/s | β‰ˆ38.5 GB |

**Reading this table.**

- Distillation eliminates the 34.2 % syntax error rate entirely and raises F1 from 0.612 to 1.000 β€” a **63.4 % relative improvement**.
- **VRAM is invariant** across all student variants (β‰ˆ2.4 GB): distillation changes *behavior*, not architecture. All efficiency gains come from parameter count, not from the training procedure.
- The student *matches* the 31B teacher on both quality metrics while using **16.0Γ— less memory**.

### 8.3 Does Teacher Scale Matter?

The two distilled variants differ only in teacher. Both reach 0.0 % error and 1.000 F1. **Teacher scale conferred no measurable quality advantage on this task.**

The plausible explanation is task-difficulty saturation: JSON extraction from short business documents lies well inside the competence of a 4B instruction-tuned model, so the additional capability of the 31B teacher is never exercised. The teacher's soft distributions over `{`, `"`, `:` are near-identical at both scales because both are near-certain.

**This has direct cost implications.** The 4B teacher is dramatically cheaper to run for logit generation. On tasks of this profile, **use the smallest teacher that saturates the task**; reserve flagship teachers for problems where the teacher itself is not already at ceiling.

*The throughput difference between the two student variants (61.91 vs. 56.12 tok/s) reflects measurement-harness and output-length differences between the two evaluation runs, not an architectural difference β€” the students are architecturally identical. We flag this as a measurement inconsistency in [Β§10.3](#103-measurement-inconsistencies).*

### 8.4 Out-of-Domain Public Benchmark: `suneeldk/text-json`

To assess generalization beyond trained schema templates, we evaluated all checkpoints against real enterprise document records from the public HuggingFace dataset `suneeldk/text-json` (2,000 records available).

| Iteration | Checkpoint | Validity | Throughput | Finding |
|---|---|---|---|---|
| 1 | `schemaforge-1b-iter1` | 0.0 % | 76.94 tok/s | Chat-template mismatch (`<start_of_turn>`) |
| **2** | **`schemaforge-1b-iter2`** | **70.0 %** | **76.27 tok/s** | **Standardized template alignment** |
| 3 | `schemaforge-1b-iter3` | 0.0 % | 74.12 tok/s | Multi-domain template drift |
| β€” | `openbmb/MiniCPM5-1B` (base) | 34.2 % | 62.00 tok/s | Zero-shot baseline |

**Interpretation.** 70.0 % zero-shot validity on unseen, real-world, multi-domain documents β€” against 100 % in-domain β€” quantifies the generalization gap honestly. The 30-point shortfall is the cost of narrow-task distillation, and it is why we recommend layering FSM-constrained decoding ([Β§9.4](#94-schema-constrained-decoding)) in production: constrained decoding converts the residual 30 % of syntax failures into guaranteed-parseable output, leaving only semantic errors to handle.

The evaluation subset size per iteration is not recorded in our experimental log β€” a reproducibility defect noted in [Β§10.1](#101-evaluation-scale). For reference, a 70 % point estimate carries a Wilson 95 % CI of **[48.1 %, 85.5 %]** at $n = 20$ and **[68.0 %, 72.0 %]** at $n = 2{,}000$. The qualitative conclusion (70 % ≫ 0 %) is robust at any of these sizes; the precise value is not.

### 8.5 Training Convergence

Distillation ran for 3 epochs with AdamW ($\text{lr} = 2\times10^{-5}$), cosine schedule, on the RTX PRO 6000 Blackwell host. Losses below are **summed, not per-token averaged**.

**Iteration 2 β€” the released checkpoint** (`schemaforge-1b-iter2`, 5 prompt-aligned samples):

| Epoch | Training loss | Ξ” from previous | Cumulative |
|---|---|---|---|
| 1 | 9,132.9 | β€” | β€” |
| 2 | 6,962.3 | βˆ’23.77 % | βˆ’23.77 % |
| 3 | 6,612.7 | βˆ’5.02 % | **βˆ’27.59 %** |

**Iteration 3** (15 multi-domain samples), for comparison of *shape* only:

| Epoch | Training loss | Ξ” from previous | Cumulative |
|---|---|---|---|
| 1 | 7,695.4 | β€” | β€” |
| 2 | 6,224.4 | βˆ’19.12 % | βˆ’19.12 % |
| 3 | 5,951.4 | βˆ’4.39 % | **βˆ’22.66 %** |

Both runs show the same decelerating profile β€” a large first-epoch drop followed by a much smaller third-epoch gain (βˆ’23.77 % β†’ βˆ’5.02 %; βˆ’19.12 % β†’ βˆ’4.39 %) β€” indicating approach to a plateau and supporting the 3-epoch cap of [Β§6.1](#61-epoch-bound-and-repetition-collapse). Continued training would reduce loss further while degrading generation, which is what makes training loss an unreliable stopping signal here.

**Absolute magnitudes are not comparable between the two runs.** These are summed quantities scaling with dataset size and sequence length, which differed. Only within-run trajectories are interpretable; per-token normalization is required for cross-run comparison and is mandated for v2 ([Β§11.3](#113-planned-v2-retraining-and-evaluation-run)).

*A note on provenance:* an earlier draft of this paper reported a trajectory of 18,083 β†’ 15,527 β†’ 14,546, taken from our internal specification documents. That series does not correspond to either run above and could not be traced to a logged execution; it has been replaced with the actual iteration-2 and iteration-3 logs, and Figure 3 has been regenerated accordingly.

![Figure 3: Training loss convergence trajectory](graphs/loss_convergence.png)

***Figure 3.** Training loss over three epochs. Iteration 2 (green, solid) is the released checkpoint; Iteration 3 (orange, dashed) is shown for profile comparison. Absolute values are summed losses and are not comparable across runs.*

### 8.6 Efficiency Analysis

![Figure 1: Inference throughput vs. VRAM footprint](graphs/throughput_vs_vram.png)

***Figure 1.** Throughput and peak VRAM for SchemaForge-1B versus the Gemma-4-31B teacher, measured on identical hardware.*

**Throughput speedup.** We report two figures because they were measured under two harnesses, and conflating them would overstate the result:

$$\text{Speedup}_{\text{in-domain}} = \frac{61.91}{12.40} = \mathbf{4.99\times}, \qquad \text{Speedup}_{\text{public-bench}} = \frac{76.27}{12.40} = \mathbf{6.15\times}.$$

The in-domain figure (**5.0Γ—**) is the conservative, like-for-like comparison and is the number we recommend citing. The public-benchmark harness produced higher throughput for the student (76.27 tok/s) on shorter average outputs; the teacher was not re-measured under that harness, so 6.15Γ— is an upper bound rather than a matched comparison.

**Memory reduction.**

$$\frac{38.5\ \text{GB}}{2.4\ \text{GB}} = \mathbf{16.04\times}.$$

**Concurrency.** At 90 % GPU memory utilization on a 96 GB card (86.4 GB usable):

$$N_{\text{workers}} = \left\lfloor \frac{96 \times 0.90}{2.4} \right\rfloor = \mathbf{36} \quad \text{versus} \quad \left\lfloor \frac{96 \times 0.90}{38.5} \right\rfloor = \mathbf{2} \text{ for the teacher.}$$

**Aggregate system throughput.**

$$36 \times 76.27 = \mathbf{2{,}745.7\ \text{tok/s}} \quad \text{versus} \quad 2 \times 12.40 = 24.8\ \text{tok/s},$$

an approximately **110Γ— improvement in tokens per GPU per second** β€” the compound effect of per-stream speedup and concurrency. This is the number with budget consequences.

**Cost framing.** At a nominal $2.50/GPU-hour and an average extraction of 200 output tokens, the teacher processes $24.8 \times 3600 / 200 \approx 446$ documents/hour ($\approx\$0.0056$/document); SchemaForge-1B processes $2{,}745.7 \times 3600/200 \approx 49{,}423$ documents/hour ($\approx\$0.00005$/document) β€” a **~110Γ— reduction in per-document GPU cost**. These figures assume perfect worker utilization and exclude preprocessing, network, and orchestration overhead; treat them as an upper bound on realizable savings.

---

## 9. Deployment and Publication

### 9.1 HuggingFace Hub Publication Protocol

**Step 1 β€” Install and authenticate.**

```bash
pip install --upgrade huggingface_hub
huggingface-cli login
# Paste a User Access Token with `write` scope from
# https://huggingface.co/settings/tokens
```

**Step 2 β€” Create the repository.**

```python
from huggingface_hub import HfApi

api = HfApi()

REPO_ID = "arrochi112/SchemaForge-1B-JSON-Extractor"
LOCAL_CHECKPOINT_DIR = "./models/distilled_minicpm5_1b_iter2"   # schemaforge-1b-iter2

print(f"[*] Creating HuggingFace repository: {REPO_ID} ...")
api.create_repo(repo_id=REPO_ID, repo_type="model", exist_ok=True)
print("[+] Repository ready.")
```

**Step 3 β€” Upload weights and proof artifacts.**

```python
from huggingface_hub import HfApi

api = HfApi()
REPO_ID = "arrochi112/SchemaForge-1B-JSON-Extractor"

# 1. Model checkpoint (safetensors, config, tokenizer)
api.upload_folder(
    folder_path="./models/distilled_minicpm5_1b_iter2",
    repo_id=REPO_ID,
    repo_type="model",
    commit_message="Add SchemaForge-1B (iter2) distilled checkpoint",
)

# 2. Visual proof artifacts
api.upload_folder(
    folder_path="./graphs",
    path_in_repo="graphs",
    repo_id=REPO_ID,
    repo_type="model",
    commit_message="Add benchmark proof charts",
)

# 3. Whitepaper
api.upload_file(
    path_or_fileobj="./SCHEMAFORGE_WHITEPAPER.md",
    path_in_repo="SCHEMAFORGE_WHITEPAPER.md",
    repo_id=REPO_ID,
    repo_type="model",
    commit_message="Add technical whitepaper",
)
print("[+] Upload complete.")
```

`upload_folder` handles Git-LFS chunking transparently; a raw `git push` will fail on shards >50 MB without `git lfs install`.

**Step 4 β€” Required file inventory.**

| File | Required | Note |
|---|---|---|
| `model.safetensors` | βœ… | Never ship `.bin` pickles β€” they fail HF security scans |
| `config.json` | βœ… | Must contain `model_type`, `num_hidden_layers`, `vocab_size` |
| `tokenizer.json`, `tokenizer_config.json` | βœ… | Vocabulary maps, special token IDs |
| `special_tokens_map.json` | βœ… | `bos_token`, `eos_token`, `pad_token` |
| `README.md` | βœ… | YAML frontmatter + proof charts + canonical prompt |
| `graphs/*.png` | recommended | Visual evidence |
| `SCHEMAFORGE_WHITEPAPER.md` | recommended | Full methodology |

### 9.2 Model Card Metadata Specification

HuggingFace indexes models from YAML frontmatter. `base_model` is what links the card as a derivative on OpenBMB's model page.

```yaml
---
language:
- en
license: apache-2.0
tags:
- distillation
- knowledge-distillation
- json-extraction
- structured-output
- gemma-4
- minicpm5
- edge-ai
- schemaforge
pipeline_tag: text-generation
base_model: openbmb/MiniCPM5-1B
library_name: transformers
metrics:
- accuracy
- f1
- throughput
---
```

### 9.3 Production Serving with vLLM

```python
from vllm import LLM, SamplingParams

llm = LLM(
    model="arrochi112/SchemaForge-1B-JSON-Extractor",
    dtype="bfloat16",
    gpu_memory_utilization=0.90,
    max_model_len=2048,
    max_num_seqs=36,                 # matches the 36-worker capacity bound
)

sampling_params = SamplingParams(
    temperature=0.0,                 # greedy: determinism is required here
    max_tokens=256,
    stop=["\n\n", "</s>"],
)

# CANONICAL PROMPT TEMPLATE β€” must match training exactly (see Β§7)
TEMPLATE = "Extract structured JSON from the text:\n{doc}\nJSON Output:"

documents = [
    "Invoice #INV-881, Vendor: Globex Corp, Date: 2026-08-03, Total: $450.00",
    "Invoice #INV-882, Vendor: Initech LLC, Date: 2026-08-04, Total: $1200.00",
]
prompts = [TEMPLATE.format(doc=d) for d in documents]

for out in llm.generate(prompts, sampling_params):
    print(out.outputs[0].text)
```

**Non-negotiable serving requirements:**

1. **No `trust_remote_code` needed** β€” MiniCPM5-1B is a stock `LlamaForCausalLM`.
2. `temperature=0.0` β€” extraction is a deterministic task; sampling introduces schema violations for no benefit.
3. **The canonical template, byte-for-byte.** Per [Β§7](#7-prompt-template-sensitivity-a-three-iteration-study), deviation is catastrophic rather than degrading.

### 9.4 Schema-Constrained Decoding

Learned formatting discipline should not be the only line of defense. Layering FSM-guided decoding [9] converts residual syntax failures into guaranteed-parseable output:

```python
from pydantic import BaseModel
from vllm.sampling_params import GuidedDecodingParams

class Invoice(BaseModel):
    invoice_number: str
    vendor_name: str
    invoice_date: str          # ISO-8601 YYYY-MM-DD
    subtotal: float
    tax: float
    grand_total: float

guided = GuidedDecodingParams(json=Invoice.model_json_schema())

sampling_params = SamplingParams(
    temperature=0.0,
    max_tokens=256,
    guided_decoding=guided,    # syntax validity now guaranteed by construction
)
```

**Division of labor.** Constrained decoding guarantees *syntax*; distillation supplies *semantics*. Neither substitutes for the other β€” a grammar-constrained base model emits perfectly-formed JSON containing wrong values. Run both.

**Validation layer.** Parse every generation through the same Pydantic model before it reaches a downstream system, and emit a structured error for the caller rather than a partially-populated record:

```python
import json
from pydantic import ValidationError

def safe_extract(raw: str) -> dict:
    try:
        return {"ok": True, "data": Invoice(**json.loads(raw)).model_dump()}
    except (json.JSONDecodeError, ValidationError) as e:
        return {"ok": False, "error": str(e), "raw": raw}
```

### 9.5 Capacity Planning

| Quantity | Value | Derivation |
|---|---|---|
| Model footprint | 2.4 GB | measured peak allocation |
| Usable GPU memory (96 GB @ 0.90) | 86.4 GB | $96 \times 0.90$ |
| Concurrent workers | **36** | $\lfloor 86.4 / 2.4 \rfloor$ |
| Per-worker throughput | 76.27 tok/s | measured |
| Aggregate throughput | **2,745.7 tok/s** | $36 \times 76.27$ |
| Documents/hour @ 200 tok each | **β‰ˆ49,423** | $2745.7 \times 3600 / 200$ |
| Teacher concurrent workers | 2 | $\lfloor 86.4 / 38.5 \rfloor$ |
| Teacher documents/hour | β‰ˆ446 | $24.8 \times 3600 / 200$ |

Assumes uniform load and no head-of-line blocking. In practice vLLM's continuous batching typically exceeds this static estimate for mixed-length workloads, since sequences are admitted and retired independently.

---

## 10. Limitations and Threats to Validity

We regard this section as central rather than obligatory. Several headline numbers in this paper are real measurements that nonetheless do not support the strength of claim a casual reader would infer, and we would rather state that ourselves.

### 10.1 Evaluation Scale

**The in-domain suite is $n = 5$ documents β€” one per domain.** The reported **1.000 F1** and **0.0 % syntax error rate** are exact-match results on five hand-authored, synthetic documents. They demonstrate that the model *can* produce correct output on representative inputs. They are **not** population estimates. The Wilson 95 % confidence interval on 5/5 successes is **[56.6 %, 100.0 %]** β€” consistent with a true validity rate as low as 57 %.

**The `suneeldk/text-json` evaluation subset size is not recorded** in our experimental log. This is a reproducibility defect. The dataset contains 2,000 records; the number actually scored per iteration is unknown. The 70.0 % figure carries a Wilson CI of [48.1 %, 85.5 %] at $n=20$ and [68.0 %, 72.0 %] at $n=2{,}000$. The *qualitative* conclusion (70 % vastly exceeds 0 %) is robust across this range; the precise value is not.

**Required remediation before any strong claim:** a held-out evaluation of $n \geq 500$ documents per domain, with reported confidence intervals.

### 10.2 Training Scale and the Distillation-vs-Formatting Confound

**The winning checkpoint was distilled on $n = 5$ samples.** This raises a confound we cannot resolve from the present data: how much of the improvement is *knowledge transfer from the teacher* versus *format conditioning* β€” the student simply learning what output shape is expected?

The [Β§7](#7-prompt-template-sensitivity-a-three-iteration-study) results actively suggest the latter is doing substantial work. Five examples is not enough supervision to teach entity extraction from scratch; it is ample to teach "emit a JSON object with these keys when you see this prefix." The base model already achieved 34.2 % validity, indicating the extraction capability was largely latent and needed unlocking rather than installing.

**The missing experiment is an SFT control:** identical data, identical template, identical hyperparameters, but $\alpha = 1.0$ (pure cross-entropy, no teacher logits). If that control also reaches 0.0 % error and 1.000 F1, then the KL term β€” the entire distillation apparatus β€” contributes nothing on this task, and the honest description of this work becomes "template-aligned supervised fine-tuning." **We did not run this control, and it is the single most important gap in the paper.**

### 10.3 Measurement Inconsistencies

Three inconsistencies in our experimental record deserve explicit statement:

1. **Throughput measured under two harnesses.** In-domain benchmarks report 61.91–62.40 tok/s; public-benchmark runs report 74.12–76.94 tok/s for architecturally identical models. The difference is harness and output-length driven, not architectural. Consequently the "5.0Γ— speedup" (61.91/12.40 = 4.99Γ—) and "6.15Γ— speedup" (76.27/12.40) figures are **not** interchangeable, and the source material from which this paper was assembled labeled the latter comparison as "5.0Γ—", which is arithmetically incorrect. We report both, and recommend citing the conservative 5.0Γ—.

2. **The two distilled variants report different throughput** (61.91 vs. 56.12 tok/s) despite identical architecture and identical VRAM. Teacher choice cannot affect student inference speed. This is measurement noise or differing output lengths, not a finding.

3. **Cross-iteration training losses are not comparable.** Values of 18,083 / 15,527 / 14,546 (Gemma-4-31B run) and 5,951 (Iteration 3) are summed rather than per-token quantities and scale with dataset size and sequence length. Only within-run trajectories are interpretable.

4. **Architecture corrections following a checkpoint audit.** Our internal specification documents described the student as a 52-layer model with a 73,440-token vocabulary requiring custom `trust_remote_code` modeling code. An audit of the released `model.safetensors` and `config.json` established that the trained model is `openbmb/MiniCPM5-1B`: a stock `LlamaForCausalLM`, 24 layers, hidden size 1536, GQA 16/2, vocabulary 130,560, 1,080,632,832 parameters, no custom modeling code. All architecture-dependent figures in this paper β€” including the vocabulary-projection retention ratio and the withdrawn depth-scaling claim ([Β§5.4](#54-why-the-released-checkpoint-needs-none-of-these)) β€” have been corrected against the artifact rather than the specification. **Verify claims against the checkpoint, not the design document.**

5. **A silent objective-substitution guard in the training loop.** The distillation loss contained
   ```python
   if torch.isnan(loss_ce):
       return kl_loss
   ```
   which silently switches the objective to pure KL β€” effectively $\alpha = 0$ rather than $0.5$ β€” with no log entry. **This guard did fire, but not during the runs reported here.** It triggered during earlier debugging of `openbmb/MiniCPM-1B-sft-bf16`, where untied `lm_head` weights produced `inf`/`NaN` logits before the fix in [Β§5.2](#52-patch-2--tied-weights-and-the-missing-lm_headweight) was applied. The Iteration-2 and Iteration-3 loss curves logged smoothly across every step ([Β§8.5](#85-training-convergence)), so $\alpha = 0.5$ held for the released checkpoint. The guard is nonetheless a defect: a run that lost hard-target supervision would have continued silently. It must raise, not substitute.

6. **DeepSpeed ZeRO-3 and FlashAttention-2 were specified but not used.** Neither appears in the executed training path, which was single-GPU PyTorch: `from_pretrained(..., dtype=torch.bfloat16).to("cuda")`, stock `AdamW`, and a plain `DataLoader` at batch size 2. ZeRO-3 is a multi-GPU parameter-sharding framework and was inapplicable to a single-card run; FlashAttention-2 was attempted and abandoned after the `flash-attn==2.5.8` wheel returned HTTP 404 and `ninja` compilation failed under Python 3.12, so training fell back to PyTorch eager attention. The table in [Β§4.5](#45-training-configuration) has been corrected to reflect what executed.

7. **The training configuration was tuned against a mis-stated GPU.** The run was configured under the belief that the host was a 48 GB RTX 6000 Ada; it was in fact a **96 GB RTX PRO 6000 Blackwell Edition**. Co-resident teacher and student require β‰ˆ40.9 GB β€” **85.2 %** of the assumed card but only **42.6 %** of the actual one, leaving β‰ˆ55 GB unused. Consequently the batch size (2), sequence length (1,024), and the teacher's `device_map="auto"` placement β€” which may have introduced unnecessary CPU offload β€” were all more conservative than the hardware required. This does not invalidate the reported results, but training-time and throughput figures should be read as a **lower bound** on what the hardware supports. v2 must read actual device properties at startup and configure from them rather than from an assumption.

8. **The recovered training script is the Iteration-3 script.** `src/02_train_distill.py` was edited in place across iterations; the file left in the repository corresponds to the Iteration-3 run, not the Iteration-2 run that produced the released checkpoint. Per-iteration scripts should be version-controlled separately.

### 10.4 Methodological Threats

**Single seed, no variance estimate.** Every result is from one training run. We report no seed variance, no error bars on any metric. Sub-2B models are known to exhibit substantial run-to-run variance on small datasets. **A minimum of 3 seeds should be considered mandatory before these numbers are cited as stable.**

**The vocabulary projection is index-aligned, not semantics-aligned.** Truncating $\mathbf{z}_T$ to the leading $130{,}560$ indices assumes frequency-ordered vocabularies and that index $i$ carries comparable meaning across tokenizers. It does not. As argued in [Β§4.4](#44-dual-tokenizer-cross-encoding-and-vocabulary-projection), what the KL term transfers is distributional shape rather than token identity. This is the most theoretically fragile component of the method and would not survive a rigorous reviewer without an ablation against an optimal-transport or minimum-edit-distance alignment.

**$\tau$ was never ablated.** $\tau = 2.0$ was fixed by convention. The interaction between $\tau$ and $\alpha$ is unexamined.

**Synthetic in-domain documents.** BMK-01…05 are hand-authored and share stylistic regularities (consistent `Vendor:`/`Date:`/`Total:` labeling, clean ASCII, no OCR noise). Real enterprise documents include scanned artifacts, multi-column layouts, non-English fields, and adversarial formatting. The 100 %/70 % gap between in-domain and public-benchmark performance is the visible edge of this.

**Teacher outputs used as ground truth.** F1 is computed against teacher-generated targets for training and against reference objects at evaluation. Where the teacher is wrong, the student is rewarded for reproducing the error. No human-annotated gold standard was constructed.

**No comparison against non-distillation baselines.** We do not compare against: prompt-engineered base MiniCPM5-1B with constrained decoding; a regex/rule-based extractor; commercial document-AI APIs; or other SLMs (Qwen2.5-1.5B, Phi-3-mini) under identical conditions. The claim "distillation is the right approach" is therefore unsupported relative to cheaper alternatives.

### 10.5 What This Work Does and Does Not Establish

**Supported by the evidence:**

- A 1.08B student can produce valid, schema-conformant JSON on representative enterprise documents at a small fraction of a 31B model's memory and latency cost.
- Prompt-template alignment between training and inference is decisive for small students β€” a 70-point effect, replicated under two independent perturbations.
- The three compatibility patches in [Β§5](#5-engineering-running-legacy-minicpm-under-transformers-5x) are required to run `openbmb/MiniCPM-1B-sft-bf16` under `transformers` 5.x, and are *not* required by the released `MiniCPM5-1B` checkpoint.
- The efficiency measurements (2.4 GB, 12.40 vs. 61.91 tok/s) are direct hardware measurements and are the most trustworthy numbers in the paper.

**Not supported by the evidence:**

- That SchemaForge-1B achieves 100 % accuracy on enterprise JSON extraction in general.
- That the KL distillation term is responsible for the gains, versus template-aligned SFT.
- That teacher scale is irrelevant in general (we tested one task at one difficulty).
- That 70.0 % is a precise estimate of out-of-domain performance.

### 10.6 Limitation Register

For traceability, we assign each limitation an identifier. The planned v2 run ([Β§11.3](#113-planned-v2-retraining-and-evaluation-run)) is organized around closing these in priority order.

| ID | Limitation | Section | Severity | Closes with |
|---|---|---|---|---|
| **L1** | In-domain eval is $n = 5$; public-benchmark subset size unrecorded | [Β§10.1](#101-evaluation-scale) | **Critical** | Scaled held-out eval, $n \geq 500$/domain, recorded |
| **L2** | No SFT control β€” distillation vs. format-conditioning confound | [Β§10.2](#102-training-scale-and-the-distillation-vs-formatting-confound) | **Critical** | $\alpha = 1.0$ ablation, all else fixed |
| **L3** | Two measurement harnesses; incomparable throughput and loss figures | [Β§10.3](#103-measurement-inconsistencies) | **High** | One unified harness; per-token-normalized loss |
| **L4** | Single seed; no variance estimates or error bars | [Β§10.4](#104-methodological-threats) | **High** | β‰₯3 seeds per configuration, mean Β± std |
| **L5** | Vocabulary projection is index-aligned, not semantics-aligned | [Β§4.4](#44-dual-tokenizer-cross-encoding-and-vocabulary-projection), [Β§10.4](#104-methodological-threats) | Medium | Ablation vs. OT / edit-distance alignment |
| **L6** | $\tau$ never ablated; $\tau$–$\alpha$ interaction unexamined | [Β§6.3](#63-loss-balance-alpha-and-temperature-tau) | Medium | 2-D sweep |
| **L7** | Synthetic, stylistically uniform in-domain documents | [Β§10.4](#104-methodological-threats) | **High** | Real documents with OCR noise, multi-column, non-English |
| **L8** | Teacher outputs used as ground truth; no human-annotated gold | [Β§10.4](#104-methodological-threats) | **High** | Human-labeled gold subset |
| **L9** | No non-distillation or competitive baselines | [Β§10.4](#104-methodological-threats) | **High** | Qwen2.5-1.5B, Phi-3-mini, base+Outlines, rule-based |

---

## 11. Conclusion and Future Work

### 11.1 Conclusion

SchemaForge demonstrates that sequence-level knowledge distillation from Gemma-4 teachers into a 1.08B MiniCPM5 student produces a structured-extraction model that matches its 31B teacher on JSON syntax validity and field-level F1 across a five-domain in-domain suite ($n = 5$ documents), while occupying **β‰ˆ2.4 GB of VRAM (16.0Γ— reduction)** and decoding at **61.91–76.27 tok/s (5.0–6.2Γ— faster)**. At 36 concurrent workers per 96 GB card β€” against 2 for the teacher β€” aggregate system throughput improves approximately **110Γ—**, changing the unit economics of document extraction by two orders of magnitude.

The methodological contributions β€” the balanced multi-task objective at $\alpha = 0.5, \tau = 2.0$, the numerically-stable log-space KL, the dual-tokenizer projection, and four non-obvious runtime patches including a depth-scaling fix that otherwise produces silent `NaN` β€” form a reproducible recipe for the sub-2B structured-output regime.

The finding we consider most transferable, however, is the negative one. **Three iterations differing principally in prompt header produced 0.0 %, 70.0 %, and 0.0 % zero-shot validity.** Neither more training data nor greater schema diversity rescued the failures; both failing iterations underperformed the *untrained* base model. Small distilled students learn format-conditioned policies bound to the literal training prefix, and they fail non-gracefully β€” to zero, not to a degraded-but-usable level β€” when that prefix changes. Anyone deploying a sub-2B structured extractor should treat the prompt template as a versioned API contract, assert conformance at the serving boundary, and monitor validity rather than latency, because latency will look perfectly healthy while the model returns nothing usable.

Finally, we would rather this paper be useful than impressive. The evaluation sets here are small ($n = 5$ in-domain, $n = 5$ training samples for the winning run), the runs are single-seed, and the SFT control that would isolate distillation's contribution from format conditioning was not run. The efficiency results are solid hardware measurements; the accuracy results are directionally strong but statistically imprecise. [Β§10](#10-limitations-and-threats-to-validity) states this in full, and the follow-up experiments listed below are the ones we would run before defending any stronger claim.

### 11.2 Future Work

**Priority 1 β€” Validate the current claims.**

1. **SFT ablation** ($\alpha = 1.0$, no teacher logits, everything else fixed). This is the decisive experiment: it determines whether the KL term contributes anything on this task.
2. **Scaled evaluation.** $n \geq 500$ held-out documents per domain with reported confidence intervals; the full 2,000-record `suneeldk/text-json` set with a recorded subset size.
3. **Multi-seed variance.** Minimum 3 seeds per configuration; report mean Β± std on every metric.
4. **Competitive baselines.** Qwen2.5-1.5B, Phi-3-mini, and prompt-engineered base MiniCPM5-1B + Outlines, under identical harnesses.

**Priority 2 β€” Strengthen the method.**

5. **Semantics-aware vocabulary alignment.** Replace index truncation with optimal-transport or minimum-edit-distance token alignment; ablate against the current projection.
6. **On-policy distillation.** Adopt GKD-style [8] student-sampled trajectories to eliminate the exposure-bias mismatch between teacher-trajectory training and student-trajectory inference.
7. **Template-robustness training.** Deliberately randomize prompt headers during distillation and measure the trade-off between in-template peak accuracy and out-of-template robustness. This directly tests the Template Binding Hypothesis ([Β§7.6](#76-analysis-the-traininference-template-gap)) as a causal claim rather than an observational one.
8. **$\tau$–$\alpha$ interaction grid.** A proper 2-D sweep.

**Priority 3 β€” Extend the scope.**

9. **Realistic document conditions.** OCR noise, multi-column layouts, non-English fields, scanned artifacts.
10. **Deeper schemas.** Recursive nesting beyond the single-level arrays in BMK-03/04; optional and union-typed fields.
11. **Quantization stacking.** INT8/INT4 on top of distillation β€” how far below 2.4 GB can the footprint go before validity degrades?
12. **Continual schema adaptation.** Adding a new target schema without full redistillation.

### 11.3 Planned v2 Retraining and Evaluation Run

The results in this paper are a **v1 release**. A second training and evaluation campaign is planned and is scoped directly against the limitation register in [Β§10.6](#106-limitation-register). Its purpose is to convert the directionally-strong-but-imprecise accuracy claims here into defensible estimates, and to determine whether the distillation objective is doing the work we attribute to it.

The v2 run will add:

- **A real-world evaluation corpus** replacing the synthetic five-document suite (L1, L7) β€” held-out documents at $n \geq 500$ per domain, including OCR-noisy scans, multi-column layouts, and non-English fields, with a human-annotated gold subset (L8) so that accuracy is no longer measured against teacher output.
- **The SFT control** ($\alpha = 1.0$, no teacher logits, everything else held fixed) (L2), which is the experiment that determines whether the KL term contributes anything on this task.
- **Competitive baselines** (L9): Qwen2.5-1.5B, Phi-3-mini, prompt-engineered base MiniCPM5-1B with FSM-constrained decoding, and a rule-based extractor β€” all under one harness.
- **A unified measurement harness** (L3) so that student and teacher throughput are like-for-like, and per-token-normalized loss so cross-run curves are comparable.
- **Multi-seed runs with reported variance** (L4) and confidence intervals on every metric.
- **An expanded metric set** beyond validity/F1/throughput/VRAM: per-field accuracy, schema-conformance rate, hallucinated-key rate, time-to-first-token, p50/p95 latency under concurrency, and cost per thousand documents.
- **Ablations** on temperature and vocabulary alignment (L5, L6).

We state this here rather than in a footnote because several headline numbers in this paper β€” the 1.000 F1 in particular β€” should be read as provisional pending that run. Results will be published as a v2 revision of this whitepaper with the v1 numbers retained for comparison rather than replaced.

---

## References

[1] G. Hinton, O. Vinyals, and J. Dean. "Distilling the Knowledge in a Neural Network." *NIPS 2014 Deep Learning Workshop*, 2015. arXiv:1503.02531.

[2] C. Buciluă, R. Caruana, and A. Niculescu-Mizil. "Model Compression." *Proceedings of KDD*, 2006.

[3] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio. "FitNets: Hints for Thin Deep Nets." *ICLR*, 2015. arXiv:1412.6550.

[4] Y. Kim and A. M. Rush. "Sequence-Level Knowledge Distillation." *EMNLP*, 2016. arXiv:1606.07947.

[5] V. Sanh, L. Debut, J. Chaumond, and T. Wolf. "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter." *NeurIPS EMCΒ² Workshop*, 2019. arXiv:1910.01108.

[6] X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu. "TinyBERT: Distilling BERT for Natural Language Understanding." *Findings of EMNLP*, 2020. arXiv:1909.10351.

[7] Y. Gu, L. Dong, F. Wei, and M. Huang. "MiniLLM: Knowledge Distillation of Large Language Models." *ICLR*, 2024. arXiv:2306.08543.

[8] R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. Ramos, M. Geist, and O. Bachem. "On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD)." *ICLR*, 2024. arXiv:2306.13649.

[9] B. T. Willard and R. Louf. "Efficient Guided Generation for Large Language Models." 2023. arXiv:2307.09702. (Outlines)

[10] S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, et al. "MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies." 2024. arXiv:2404.06395.

[11] M. Abdin et al. "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone." 2024. arXiv:2404.14219.

[12] Gemma Team, Google DeepMind. "Gemma: Open Models Based on Gemini Research and Technology." 2024. arXiv:2403.08295.

[13] M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr. "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design, or: How I learned to start worrying about prompt formatting." *ICLR*, 2024. arXiv:2310.11324.

[14] Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp. "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity." *ACL*, 2022. arXiv:2104.08786.

[15] M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, and G. Stanovsky. "State of What Art? A Call for Multi-Prompt LLM Evaluation." *TACL*, 2024. arXiv:2401.00595.

[16] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. "Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM)." *SOSP*, 2023. arXiv:2309.06180.

[17] T. Dao. "FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning." *ICLR*, 2024. arXiv:2307.08691.

[18] S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. "ZeRO: Memory Optimizations Toward Training Trillion Parameter Models." *SC*, 2020. arXiv:1910.02054.

[19] I. Loshchilov and F. Hutter. "Decoupled Weight Decay Regularization (AdamW)." *ICLR*, 2019. arXiv:1711.05101.

[20] T. Bray (Ed.). "The JavaScript Object Notation (JSON) Data Interchange Format." RFC 8259, IETF, 2017.

---

## Appendix A: Complete Hyperparameter Specification

```yaml
# SchemaForge-1B β€” schemaforge-1b-iter2 (production checkpoint)

model:
  student:            openbmb/MiniCPM5-1B
  student_arch:       LlamaForCausalLM   # no trust_remote_code needed
  student_params:     1080632832         # 679552512 non-embedding
  student_layers:     24
  student_hidden:     1536
  student_heads:      16 query / 2 kv (GQA)
  student_vocab:      130560
  teacher_primary:    google/gemma-4-31B
  teacher_secondary:  google/gemma-4-E4B-it
  teacher_vocab:      256000
  teacher_frozen:     true

distillation:
  alpha:              0.5          # CE weight
  tau:                2.0          # temperature
  kl_reduction:       batchmean
  kl_log_target:      true
  vocab_projection:   truncate_to_student_width
  label_mask_value:   -100
  mask_prompt_tokens: true
  mask_padding:       true

optimization:
  optimizer:          AdamW
  learning_rate:      2.0e-5
  betas:              [0.9, 0.999]
  eps:                1.0e-8
  weight_decay:       0.01
  lr_scheduler:       cosine
  warmup_ratio:       0.05
  max_grad_norm:      1.0
  epochs:             3
  early_stopping:     val_loss
  gradient_accumulation_steps: 4

runtime:
  precision:          bfloat16
  attention:          eager           # flash-attn install failed on py3.12
  sharding:           none            # single GPU
  max_seq_length:     1024        # as executed; 2048 was specified
  gradient_checkpointing: true

hardware:
  gpu:                NVIDIA RTX PRO 6000 Blackwell Edition
  vram:               96GB
  provider:           Nebius AI Cloud
  count:              1

software:
  python:             "3.12"
  torch:              "2.5"
  transformers:       "5.x"

inference:
  temperature:        0.0
  do_sample:          false
  max_new_tokens:     256
  trust_remote_code:  false        # stock Llama architecture
```

---

## Appendix B: Compatibility Patch Source

```python
"""
schemaforge/compat.py
transformers 5.x shims for LEGACY trust_remote_code MiniCPM checkpoints
(openbmb/MiniCPM-1B-sft-bf16). Python 3.12 / PyTorch 2.5.

NOT required for openbmb/MiniCPM5-1B, which is a stock LlamaForCausalLM.
Call apply_compat_shims() BEFORE from_pretrained();
call retie_lm_head(model) AFTER instantiation.
"""


# ---------------------------------------------------------------- Patch 1 ----
def _patch_import_utils():
    """Re-inject is_torch_fx_available removed in transformers 5.x."""
    import transformers.utils.import_utils as import_utils
    # Symbolic tracing is unused in the training path, so False is safe.
    import_utils.is_torch_fx_available = lambda: False


# ---------------------------------------------------------------- Patch 2 ----
def _patch_tied_weights(source: str = "model.embed_tokens.weight"):
    """
    MiniCPM-1B-sft-bf16 omits lm_head.weight from disk (tied to embeddings).
    transformers 5.x reports the absence as a CORRUPTED CHECKPOINT, and also
    rejects the legacy list-form _tied_weights_keys. Patch the expansion
    helper so both are handled.
    """
    import transformers.modeling_utils as mu
    _orig = mu.PreTrainedModel.get_expanded_tied_weights_keys

    def _patched(self, *args, **kwargs):
        keys = getattr(self, "_tied_weights_keys", None)
        if isinstance(keys, (list, tuple)):
            self._tied_weights_keys = {k: source for k in keys}
        return _orig(self, *args, **kwargs)

    mu.PreTrainedModel.get_expanded_tied_weights_keys = _patched


def retie_lm_head(model):
    """Re-establish the output-projection tie after instantiation."""
    model.lm_head.weight = model.model.embed_tokens.weight
    return model


# ---------------------------------------------------------------- Patch 3 ----
def _patch_dynamic_cache():
    """Restore legacy-cache interop methods on DynamicCache."""
    from transformers.cache_utils import DynamicCache

    if hasattr(DynamicCache, "from_legacy_cache"):
        return

    @classmethod
    def from_legacy_cache(cls, past_key_values=None):
        cache = cls()
        if past_key_values is not None:
            for layer_idx, (k, v) in enumerate(past_key_values):
                cache.update(k, v, layer_idx)
        return cache

    def to_legacy_cache(self):
        return tuple((layer.keys, layer.values) for layer in self.layers)

    def get_usable_length(self, new_seq_length: int, layer_idx: int = 0) -> int:
        seen = self.get_seq_length(layer_idx)
        return seen if seen is not None else 0

    DynamicCache.from_legacy_cache = from_legacy_cache
    DynamicCache.to_legacy_cache   = to_legacy_cache
    DynamicCache.get_usable_length = get_usable_length


# ------------------------------------------------------------------ Driver ---
def apply_compat_shims():
    """Call BEFORE from_pretrained(). Then call retie_lm_head(model)."""
    _patch_import_utils()   # 1. before dynamic module load
    _patch_tied_weights()   # 2. before class instantiation
    _patch_dynamic_cache()  # 3. before first generate()


# ----------------------------------------------------------- NaN diagnostic --
def install_nan_probe(model):
    """Localize which layer first produces NaN. Architecture-agnostic;
    worth leaving enabled in any long training run."""
    import torch

    def hook(idx):
        def fn(_module, _inp, out):
            h = out[0] if isinstance(out, tuple) else out
            if torch.isnan(h).any():
                raise RuntimeError(f"NaN first observed at layer {idx}")
        return fn

    for i, layer in enumerate(model.model.layers):
        layer.register_forward_hook(hook(i))
    return model
```

---

## Appendix C: Verbatim Prompt Templates

### C.1 Canonical Template (Iteration 2 β€” Production)

```
Extract structured JSON from the text:
{document_text}
JSON Output:
```

Python literal (whitespace is significant):

```python
CANONICAL_TEMPLATE = "Extract structured JSON from the text:\n{doc}\nJSON Output:"
```

### C.2 Iteration 1 Template (Failed β€” 0.0 %)

```
<start_of_turn>system
You are a JSON extraction assistant.<end_of_turn>
<start_of_turn>user
{document_text}<end_of_turn>
<start_of_turn>model
```

### C.3 Iteration 3 Template (Failed β€” 0.0 %)

```
You are an expert enterprise document parser specializing in structured
data extraction across multiple business domains.

Extract structured JSON from the text:
{document_text}
JSON Output:
```

### C.4 Schema-Explicit Variant (Model-Card Quickstart)

```python
prompt = (
    "Extract entity details into JSON with keys "
    "'invoice_number', 'vendor', 'date', 'amount':\n"
    "Invoice INV-99281, Acme Enterprise Solutions Inc, "
    "Date 2026-08-03, Total Amount $12,450.00.\n"
    "JSON Output:"
)
```

> ⚠️ **Deployment warning.** C.2 and C.3 differ from C.1 only in the instruction header. Both produced **0.0 %** validity β€” worse than the untrained base model's 34.2 %. Do not modify C.1.

---

## Appendix D: Benchmark Schemas and Source Documents

### D.1 BMK-01 β€” Enterprise Financial Invoices

```json
{
  "invoice_number": "str",
  "vendor_name":    "str",
  "invoice_date":   "YYYY-MM-DD",
  "subtotal":       "float",
  "tax":            "float",
  "grand_total":    "float"
}
```

**Source:** `INVOICE #INV-1001. Vendor: Acme Supply Co. Date: 2026-04-10. Item: Office Chairs x 4 @ $120.00 = $480.00. Subtotal: $480.00. Tax (8%): $38.40. Total: $518.40.`

**Expected:**

```json
{
  "invoice_number": "INV-1001",
  "vendor_name": "Acme Supply Co",
  "invoice_date": "2026-04-10",
  "subtotal": 480.00,
  "tax": 38.40,
  "grand_total": 518.40
}
```

### D.2 BMK-02 β€” Supply Chain Logistics and Freight

```json
{
  "bill_of_lading":  "str",
  "carrier_name":    "str",
  "ship_date":       "YYYY-MM-DD",
  "container_count": "int",
  "freight_cost":    "float",
  "total_amount":    "float"
}
```

**Source:** `TAX INVOICE 9942. Issued by: Quantum Logistics. Date: 2026-05-01. Shipping Container x 1 at $2500.00 ($2500.00). Insurance Fee x 1 at $150.00 ($150.00). Subtotal $2650.00. Tax $265.00. Total $2915.00.`

### D.3 BMK-03 β€” Commercial Hardware Bills of Sale

```json
{
  "receipt_id":       "str",
  "vendor":           "str",
  "transaction_date": "YYYY-MM-DD",
  "line_items":       "array<{description, quantity, unit_price, line_total}>",
  "tax":              "float",
  "total":            "float"
}
```

**Source:** `BILL OF SALE #771. Vendor: Tech Hardware LLC. Date: 2026-05-15. Server Rack x 2 @ $800.00 = $1600.00. Cable Pack x 5 @ $20.00 = $100.00. Subtotal: $1700.00. Tax: $136.00. Grand Total: $1836.00.`

*The only nested-array schema in the suite.*

### D.4 BMK-04 β€” Biomedical Supply Receipts

```json
{
  "order_id":    "str",
  "supplier":    "str",
  "order_date":  "YYYY-MM-DD",
  "items":       "array<{name, quantity, unit_price, subtotal}>",
  "total_price": "float"
}
```

**Source:** `COMMERCIAL INVOICE #INV-5502. Vendor: BioMed Supplies. Date: 2026-06-20. Centrifuge Tube x 10 @ $15.00 = $150.00. Pipette Set x 2 @ $75.00 = $150.00. Subtotal: $300.00. Tax: $24.00. Total: $324.00.`

### D.5 BMK-05 β€” Cloud Infrastructure Receipts

```json
{
  "receipt_id":          "str",
  "provider":            "str",
  "billing_date":        "YYYY-MM-DD",
  "service_description": "str",
  "total_charge":        "float"
}
```

**Source:** `PURCHASE RECEIPT #PR-8819. Vendor: Cloud Servers Inc. Date: 2026-07-04. Virtual Machine Host x 1 @ $1200.00 = $1200.00. Subtotal: $1200.00. Tax: $96.00. Total: $1296.00.`

---

## Appendix E: Reproducibility Checklist

| Item | Status | Note |
|---|---|---|
| Model weights released | βœ… | HuggingFace Hub, safetensors |
| Training hyperparameters | βœ… | [Appendix A](#appendix-a-complete-hyperparameter-specification), complete |
| Loss implementation | βœ… | [Β§4.3](#43-numerically-stable-log-space-formulation), full source |
| Compatibility patches | βœ… | [Appendix B](#appendix-b-compatibility-patch-source), full source |
| Prompt templates | βœ… | [Appendix C](#appendix-c-verbatim-prompt-templates), verbatim, all three |
| Evaluation schemas + documents | βœ… | [Appendix D](#appendix-d-benchmark-schemas-and-source-documents) |
| Hardware specification | βœ… | RTX PRO 6000 Blackwell Edition 96 GB, Nebius |
| Software versions | βœ… | Python 3.12, PyTorch 2.5, transformers 5.x |
| Training dataset released | ⚠️ | $n = 5$ teacher-generated samples β€” should be released |
| Random seeds | ❌ | **Not recorded.** Single-seed results |
| Evaluation subset size (`suneeldk/text-json`) | ❌ | **Not recorded.** See [§10.1](#101-evaluation-scale) |
| Confidence intervals | ❌ | **Not computed at experiment time**; retrospective Wilson intervals in [§8.1](#81-in-domain-benchmark-suite)/[§8.4](#84-out-of-domain-public-benchmark-suneeldktext-json) |
| SFT control ($\alpha=1.0$) | ❌ | **Not run.** Highest-priority gap β€” [Β§10.2](#102-training-scale-and-the-distillation-vs-formatting-confound) |
| Multi-seed variance | ❌ | Not run β€” [Β§11.2](#112-future-work) Priority 1 |
| Competitive baselines | ❌ | Not run β€” [Β§11.2](#112-future-work) Priority 1 |

---

## Citation

```bibtex
@techreport{ty2026schemaforge,
  title  = {SchemaForge: Distilling Ultra-Large Foundation Models into Edge SLMs
            for Real-Time Enterprise JSON Extraction --
            A Comparative Study of Gemma-4 Teachers and MiniCPM5-1B},
  author = {Ty, Arjhine A.},
  year   = {2026},
  note   = {v1.0. Model: SchemaForge-1B (schemaforge-1b-iter2)}
}
```

---

*SchemaForge v1.0 β€” August 2026. Author: Arjhine A. Ty. Production checkpoint: `schemaforge-1b-iter2`. Distilled from `google/gemma-4-31B` and `google/gemma-4-E4B-it` into `openbmb/MiniCPM5-1B` on NVIDIA RTX PRO 6000 Blackwell Edition / Nebius AI Cloud.*