OpenAI’s Ten Math Advances Need the Slow Work of Verification
If the arguments survive expert scrutiny, several would reshape their fields. The important artifact is not the headline—it is a proof that specialists can challenge line by line.

Sources: OpenAI manuscript, Ten Advances in Mathematics and Theoretical Computer Science, OpenAI’s earlier First Proof submissions and review update, First Proof research-level mathematics challenge.
OpenAI released a 249-page manuscript presenting ten results that it says were obtained by an internal model. The collection ranges from high-dimensional sphere packing and coding theory to non-sofic groups, quantum parallel repetition, lattice hardness, Ramsey numbers, and extremal graph theory. These are not short benchmark answers. Each claim depends on a long argument in a specialized field.
The scale of the claims is extraordinary. The manuscript says its sphere-packing result improves the general high-dimensional exponent for the first time since 1978. It also presents an explicit non-sofic group and says other chapters settle, disprove, or sharply improve longstanding conjectures. Those statements are the authors’ claims until independent specialists have checked the definitions, reductions, hidden assumptions, and cited lemmas.
A manuscript is evidence, not certification
Publishing full proofs is the right starting point because it creates something falsifiable. Experts can identify a broken inequality, a missing case, a theorem applied outside its conditions, or an argument that quietly assumes the conclusion. A polished abstract cannot substitute for that process, and neither can the reputation of the lab that produced it.
OpenAI’s own First Proof experience shows why review matters. In February, the company published attempts on ten research-level problems and later updated its assessment after expert feedback, acknowledging that one proof it initially considered likely correct was wrong. That correction is not a failure of openness; it is the mechanism mathematics uses to separate plausible reasoning from durable knowledge.
Ten claims create ten different verification jobs
There is no single “AI solved math” check. A coding theorist, group theorist, operator algebraist, complexity theorist, and discrete geometer bring different background knowledge and will inspect different failure modes. Some chapters may be correct while another needs repair. Treating the document as one all-or-nothing demonstration would hide that essential granularity.
Verification should also distinguish novelty from correctness. A valid proof may reproduce an unpublished argument, depend on an overlooked result, or package known techniques in a genuinely new combination. Literature search is especially difficult when terminology differs between subfields. Establishing priority and contribution takes time after the logical core has been checked.
What responsible AI-assisted research looks like
Labs should preserve prompts, intermediate attempts, retrieval sources, human edits, model versions, and failed approaches. That provenance helps reviewers understand whether a result came from independent reasoning, hidden access to prior work, or a human-machine collaboration more substantial than the headline suggests. It also makes later corrections traceable.
The most credible outcome is not a model receiving ten instant trophies. It is a reviewable pipeline in which AI proposes arguments, humans interrogate them, formal tools verify suitable components, and the public record changes when an error is found. If these proofs hold, that process will make the achievement stronger. If some fail, it will still teach researchers where confident mathematical language outran the reasoning beneath it.
Quick questions
Did OpenAI prove ten major open problems?
OpenAI published a manuscript claiming ten advances produced by an internal model. The document provides detailed arguments, but each result still needs independent specialist review before it should be treated as settled.
Why is a 249-page manuscript not enough?
Length and detail make claims inspectable, not automatically correct. Experts must check every dependency, boundary case, reduction, citation, and novelty claim in the relevant field.
Can formal proof software verify all ten results?
Not automatically. Proof assistants can provide strong assurance after arguments are formalized, but translating advanced mathematics into a formal system is substantial work and does not by itself establish novelty.