Matthew Ford
chore: v11.6 + v11.5 + candidate-20260716a artifact reports and manifests
371c1ae
|
Raw
History Blame Contribute Delete
869 Bytes

Frozen genre-representation evaluation

The preregistered decision is negative. All three splits passed molecule and formula leakage audits, but the learned representation underperformed the composition baseline on every split.

Split Test n Composition Learned Learned - composition (95% CI)
seed 1 327 0.7890 0.6972 -0.0917 [-0.1376, -0.0459]
seed 2 309 0.8414 0.6893 -0.1521 [-0.2039, -0.1036]
seed 3 340 0.8088 0.7059 -0.1029 [-0.1500, -0.0559]

Under the frozen rule, support required the lower 95% confidence bound to be above zero on every valid split. The evidence therefore does not establish that the learned representation captures deeper formula semantics than elementary composition statistics. The benchmark's genre structure may be explained by those elementary statistics.