File size: 4,829 Bytes
edf3761 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 | ---
library_name: pytorch
license: apache-2.0
tags:
- mathematics
- modular-arithmetic
- learned-algorithms
- recurrent-neural-network
- cuda
- sair
---
# SAIR Modular Arithmetic Challenge — Learned Horner Weight Soup
This repository is a submission artifact for the
[SAIR Modular Arithmetic Challenge](https://github.com/SAIRcompetition/modular-arithmetic-challenge).
It implements the official `ModularMultiplicationModel` interface and emits
base-2 digits for `(a × b) mod p`.
> Evaluation status: the results below were produced independently with the
> published official evaluator. They are not an organizer-certified private-set
> leaderboard result.
## Result summary
The exact artifact in this repository was evaluated on an NVIDIA L40S on
2026-08-11.
| Evaluation set | H90 | Overall accuracy | Scored cases | Inference time |
|---|---:|---:|---:|---:|
| Published benchmark | 10 | 1.0000 | 1000/1000 | 248.4 s |
| Independent generator seed 1 | 10 | 1.0000 | 1000/1000 | 249.9 s |
| Independent generator seed 2 | 10 | 1.0000 | 1000/1000 | 248.9 s |
| Independent generator seed 3 | 10 | 1.0000 | 1000/1000 | 249.2 s |
| Independent generator seed 4 | 10 | 1.0000 | 1000/1000 | 246.0 s |
| Independent generator seed 5 | 10 | 1.0000 | 1000/1000 | 247.2 s |
All six runs passed the official static analysis, manifest validation,
preprocessing-isolation check, model loading, and determinism check. The five
additional sets use the published `generate_private_test_set` implementation
with independent seeds, but they are not the organizers' secret evaluation
set.
Additional diagnostics:
- algebraic metamorphic tests: 140/140 across scored Tiers 1–10;
- all 17 learned tensors randomized: 0/30 non-zero probes remained correct;
- 2048-bit modulus boundary: passed;
- 2049-bit modulus boundary: deliberately rejected with output zero.
The diagnostic Tier 0 is unscored. This submission declines primes wider than
2048 bits so that the diagnostic does not exhaust the shared 300-second budget;
the scored Tiers 1–10 are fully covered.
Raw result files, seed fingerprints, evaluator hashes, and the scope boundary
are recorded in [`evaluation_2026-08-11/`](evaluation_2026-08-11/).
## Architecture
The model is a width-generic, modulus-conditioned recurrent Horner cell with
91,840 learned parameters. Per-bit local features feed a shared bidirectional
associative scan: one direction propagates carry information and the other
propagates the learned modular-reduction decision. The same learned transition
is reused across positions, scan levels, recurrent steps, and register widths.
Inference performs two shared-weight passes over raw operand digits. On CUDA,
the recurrent state and model use FP16 and one complete three-round learned
transition is captured in a CUDA graph and replayed. The outer schedule does
not compute, correct, or look up the modular product; emitted answer digits are
produced by the trained parameters.
## Weight-soup provenance
`weights.pt` is an elementwise FP32 interpolation of two checkpoints from the
same learned-cell lineage:
| Parent | Git revision | Mixture weight |
|---|---|---:|
| 814,335-step harvest checkpoint | `c00027c6db90076e58bac25ce6c4c23a46ccfd40` | 0.75 |
| r15 champion (`17b8eb341153`) | `4a6cbbead597cdbafd618b28666a730313857dd1` | 0.25 |
Certified artifact fingerprints:
```text
weights.pt SHA-256:
2a245597f4499f83d0097d87801bd6dce79f014e5d1e3ed9cfa5f7a5e5c2363c
sorted tensor-content SHA-256:
5eb2e582891f59690cf719d8c44e040b6cb33e21356d3b62ff40c26c2ef79961
```
See [`provenance.json`](provenance.json) for parent file hashes and the full
machine-readable record.
## Submission layout
```text
manifest.json official entry point and model/training description
model.py ModularMultiplicationModel implementation
arch.py learned recurrent cell architecture
weights.pt certified weight-soup state dict
provenance.json parent and output fingerprints
evaluation_2026-08-11/
evaluation evidence and dataset fingerprints
```
This is an evaluator-specific PyTorch artifact, not a Transformers model and
not a Hugging Face hosted-inference endpoint.
## Reproduce the official interface check
Install the official challenge package, then run:
```bash
modchallenge check .
modchallenge evaluate .
```
To evaluate the immutable Hugging Face revision:
```bash
modchallenge evaluate-hf \
Dario9709/SAIR-Modular-Arithmetic-Challenge \
<40-character-commit-sha>
```
CUDA is strongly recommended. CPU execution is substantially slower and can
time out before the highest tiers even when predictions are otherwise correct.
## License
Apache-2.0. Competition acceptance and ranking remain subject to the
organizers' rules, secret-set evaluation, and manual compliance review.
|