qedbot

Benchmarks and contamination

Each benchmark here is a public set of problems the register holds, with the share of its members that already carry a recorded AI claim. A problem with a published solution can test recall of that solution as much as reasoning.

Sets

3

Formal Conjectures: open

Statements the formal record still marks as open research.

28 of 432 members carry a recorded AI claim (6.5%).

3 cite a proof1 machine-checked

Formal Conjectures: all

Every statement carrying a Lean formalisation.

311 of 1409 members carry a recorded AI claim (22.1%).

582 cite a proof26 machine-checked

Erdős prize problems still open

Erdős problems with money attached and no recorded solution.

7 of 47 members carry a recorded AI claim (14.9%).

5 cite a proof

What this does and does not show

is
Public, recorded AI work against statements in each set, at the last build.
not
Evidence that any particular model was trained on any particular solution.
not
A measure of difficulty. Attention and difficulty are only loosely related.
absent
Benchmarks whose problem sets are not public, such as FrontierMath and AIMO, cannot be measured here, because their membership is unknown.