Frozen genre-representation evaluation
The preregistered decision is negative. All three splits passed molecule and formula leakage audits, but the learned representation underperformed the composition baseline on every split.
| Split | Test n | Composition | Learned | Learned - composition (95% CI) |
|---|---|---|---|---|
| seed 1 | 327 | 0.7890 | 0.6972 | -0.0917 [-0.1376, -0.0459] |
| seed 2 | 309 | 0.8414 | 0.6893 | -0.1521 [-0.2039, -0.1036] |
| seed 3 | 340 | 0.8088 | 0.7059 | -0.1029 [-0.1500, -0.0559] |
Under the frozen rule, support required the lower 95% confidence bound to be above zero on every valid split. The evidence therefore does not establish that the learned representation captures deeper formula semantics than elementary composition statistics. The benchmark's genre structure may be explained by those elementary statistics.