Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

A parameterized benchmark for measuring mathematical progress beyond human frontiers
Muhan Zhangยท 2026ยท DOI 10.48550/arXiv.2609.24555

The core problem

The pursuit of artificial superintelligence (ASI) requires benchmarks that can measure progress beyond current human capabilities. Traditional exams cap performance at a maximum score, offering no signal once models surpass human baselines. The Endless Exam addresses this by introducing a benchmark that evaluates mathematical constructions through fourteen parameterized families. Each submitted object is automatically checked for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at . This allows for continuous measurement of progress, even as models exceed human-level performance. The benchmark draws on open mathematical problems for long-term targets and generates new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. The generators, verifiers, references, model responses, and analysis are released to su

Innovation

Eight models were evaluated on 69 distinct instances across the fourteen construction families. Despite none surpassing the published frontier, the continuous quality scores effectively distinguished performance among models. The scores ranged from to , with variations across families and instance parameters. Size-quality curves were generated, showing how construction quality changes as problem size increases. For most models, quality decreased with larger parameters, indicating difficulty in scaling constructions. However, some models exhibited stable or slightly improving quality, suggesting potential for scaling. The results highlight the benchmark's ability to provide nuanced performance metrics even when absolute frontiers are not exceeded. The data also revealed that certain families were more discriminative than others, with some families yielding scores that clustered near the frontier, while others showed wide dispersion. This variation underscores the importance of diverse construction families in measuring progress.
The pursuit of artificial superintelligence (ASI) requires benchmarks that can measure progress beyond current human capabilities. Traditional exams cap performance at a maximum score, offering no signal once models surpass human baselines. The Endless Exam addresses this by introducing a benchmark that evaluates mathematical constructions through fourteen parameterized families. Each submitted object is automatically checked for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at . This allows for continuous measurement of progress, even as models exceed human-level performance. The benchmark draws on open mathematical problems for long-term targets and generates new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. The generators, verifiers, references, model responses, and analysis are released to support continued measurement before and beyond human frontiers.
The Endless Exam is built upon fourteen parameterized construction families, each derived from open mathematical problems. These families generate instances at varying parameters, allowing for scalable difficulty. For each instance, a model submits a construction object, which is automatically verified for validity. The quality of the construction is then scored relative to a published frontier or baseline, with the score defined as:

Why it matters

The Endless Exam represents a significant step toward benchmarks that can measure progress beyond human capabilities. By allowing scores to exceed , it provides a continuous signal for improvement, essential for tracking progress toward superintelligence. The use of parameterized families ensures that the benchmark can generate new, harder instances as models improve, preventing saturation. The compact certificates enable verification of large constructions, addressing a key challenge in scaling benchmarks. However, the current evaluation shows that no model surpasses the frontier, indicating that human-level mathematical construction remains a challenge. The size-quality curves offer insights into how models handle increasing complexity, which can guide future model development. The release of all components supports reproducibility and community-driven extension. Future work could expand the number of families and incorporate more complex mathematical problems. The benchmark's design aligns with the goal of creating an "endless" exam that evolves with model capabilities, providing a lasting measure of progress toward ASI.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ