Matthew Ford
chore: v11.6 + v11.5 + candidate-20260716a artifact reports and manifests
371c1ae
|
Raw
History Blame Contribute Delete
869 Bytes
# Frozen genre-representation evaluation
The preregistered decision is **negative**. All three splits passed molecule and
formula leakage audits, but the learned representation underperformed the
composition baseline on every split.
| Split | Test n | Composition | Learned | Learned - composition (95% CI) |
|---|---:|---:|---:|---:|
| seed 1 | 327 | 0.7890 | 0.6972 | -0.0917 [-0.1376, -0.0459] |
| seed 2 | 309 | 0.8414 | 0.6893 | -0.1521 [-0.2039, -0.1036] |
| seed 3 | 340 | 0.8088 | 0.7059 | -0.1029 [-0.1500, -0.0559] |
Under the frozen rule, support required the lower 95% confidence bound to be
above zero on every valid split. The evidence therefore does not establish that
the learned representation captures deeper formula semantics than elementary
composition statistics. The benchmark's genre structure may be explained by
those elementary statistics.