Papers
arxiv:2609.24555

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

Published on Sep 28
ยท Submitted by
Muhan Zhang
on Oct 1
Authors:

Abstract

We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.

Community

Paper author Paper submitter

The Endless Exam: automatically verified mathematical constructions, with uncapped scores.

Models construct sets, graphs, codes and other mathematical objects. Valid answers earn continuous scores based on quality, so improvements still count after a reference is surpassed. Larger dimensions, graph diameters and code lengths create further challenges under the same mathematical rules.

The current release covers 69 instances across 14 families, with 17 evaluated configurations, including code and web access. The leaderboard combines 30 published-frontier references and 39 reproducible construction baselines. 100 marks average reference parity, and scores can exceed it. None of the evaluated models surpasses a published frontier.

Code, verifiers, reference constructions and model responses are available in the linked repository. New model evaluations and independently verified constructions are welcome.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.24555
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.24555 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.24555 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.24555 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.