| # Frozen genre-representation evaluation | |
| The preregistered decision is **negative**. All three splits passed molecule and | |
| formula leakage audits, but the learned representation underperformed the | |
| composition baseline on every split. | |
| | Split | Test n | Composition | Learned | Learned - composition (95% CI) | | |
| |---|---:|---:|---:|---:| | |
| | seed 1 | 327 | 0.7890 | 0.6972 | -0.0917 [-0.1376, -0.0459] | | |
| | seed 2 | 309 | 0.8414 | 0.6893 | -0.1521 [-0.2039, -0.1036] | | |
| | seed 3 | 340 | 0.8088 | 0.7059 | -0.1029 [-0.1500, -0.0559] | | |
| Under the frozen rule, support required the lower 95% confidence bound to be | |
| above zero on every valid split. The evidence therefore does not establish that | |
| the learned representation captures deeper formula semantics than elementary | |
| composition statistics. The benchmark's genre structure may be explained by | |
| those elementary statistics. | |