AI Solved Complex Mathematics, Then OpenAI Uncovered a Striking Problem

News
OpenAI brengt interactieve ChatGPT-antwoorden nu ook naar gratis gebruikers
OpenAI has withdrawn three mathematical research papers produced by its AI model. A sign error invalidated one proof and undermined two other papers that built on it. The correction came just one day after the company made hundreds of AI-generated mathematical results public.
With the publication of its mathematical research results, OpenAI is giving scientists access to manuscripts and computer-checkable proofs. The company says it wants to make new knowledge available. The first corrections also highlight how much depends on rigorous review: researchers must assess not only individual proofs, but also which results rely on one another.

How can one mistake undermine three research papers?

OpenAI explains the issue in its record of corrections and retractions. The manuscript Algebraicity of Weil classes on split abelian eightfolds contained a sign error that invalidated a crucial step in its proof. Two other papers used the same mathematical construction. The mistake therefore affected not one isolated result, but an entire chain of arguments.
The company also revised 14 other manuscripts. OpenAI repaired proof steps, corrected theorems and clarified conditions, among other changes. The withdrawn papers include explanations of the missing justification. Earlier versions remain available, allowing researchers to track what the company changed.
An invalid proof does not automatically make the corresponding theorem false. It means the reasoning provided does not establish the conclusion. Another argument could still prove that conclusion later. Conversely, a convincingly written research paper does not make a mathematical claim correct.

What has OpenAI’s mathematical AI actually investigated?

OpenAI’s current catalog contains 719 manuscripts organized into 372 families. Each family groups related papers, such as a main result, supporting arguments and alternative proofs. The collection therefore does not represent 719 independently solved problems.
OpenAI gave its internal model roughly 4,000 problems. The company then grouped related outputs and selected results based on their scientific significance. Some findings also build on earlier work produced by the models.
That process makes a simple success rate misleading. The number of problems submitted and the number of published manuscripts measure different things. Without additional information, they do not show what share of the original questions the model solved correctly.
OpenAI compares the average computing effort per result with roughly three hours of reasoning in ChatGPT Pro. That comparison describes the amount of computing power used. It does not mean that a public ChatGPT subscription can solve the same problems within three hours. OpenAI is also publishing 10 summaries of the model’s reasoning to provide more insight into its research approach.
AI Wereld previously reported on a separate mathematical experiment with GPT-5 Pro in convex optimization. The new collection significantly expands the material available for study—from a single proof-of-concept to hundreds of interconnected manuscripts.

How does Lean verify mathematical proofs?

OpenAI says roughly 42 percent of the main results have now been formalized. This involves encoding theorems and proof steps in a precise computer language. The Lean proof assistant can then check whether a conclusion follows logically from the stated definitions and assumptions.
That is fundamentally different from asking an AI model whether an answer is correct. Lean does not assess persuasive wording; it checks formal proof steps. The software therefore offers a technical way to test mathematical arguments.
Researchers still need to establish exactly what such a check covers. OpenAI notes in its proof-verification instructions that some checks verify only supporting results. They do not verify the main theorem of the corresponding paper. A valid proof of one building block is therefore not the same as a complete proof of the construction that depends on it.
AI Wereld previously described this combination of solution-finding and proof formalization in its coverage of Harmonic’s Aristotle Agent. Such systems focus not only on producing answers, but also on making their reasoning verifiable.
The absence of a formalization does not, by itself, prove that a result is wrong. It simply means that this specific form of software verification is not yet available for that result. OpenAI warns that non-formalized results may contain errors and says it will publish additional formalizations as they become available.

Why do mathematicians also want access to the model?

The independent Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, advised OpenAI on the publication. In its October 6 statement, the group drew a clear line: its involvement does not give the results a scientific seal of approval. The mathematical community must judge their significance for itself.
The advisory group also wants to prevent scientists from simply chasing AI companies’ priorities. Mathematicians must be free to ask their own questions, develop their own methods and pursue research beyond the topics a company selects to showcase a model. That requires access to powerful systems and sufficient computing resources.
In its recommendations for responsible publication, AGMAI even urges AI labs to stop testing advanced mathematical problems on models that the broader scientific community cannot access. The group warns that this imbalance could create a divide between AI companies and other research institutions.
AGMAI is also calling for readable proofs, careful literature references and independent research archives. The group based its recommendations in part on more than 600 responses from the mathematical community. Scientific openness, it argues, is not just about making files available—it is also about enabling others to understand the work and build on it.
OpenAI says it will fund workshops, conferences and other research programs aimed at improving that understanding. The company is also working on a responsible release of the model used. It has not announced a date. For now, researchers can study the published results, but they cannot yet use the same model to explore their own research questions.
loading

Loading