For the past few years, we were told that AI can calculate, predict, imitate, and search. But it cannot really discover. That last part was supposed to be different. A calculator can give you the answer. A search engine can find the answer. An AI can even explain the answer.
But discovering something that nobody has discovered before? That was supposed to require a human mind. That argument is becoming harder to defend.
OpenAI recently announced that its internal Astra model produced results on ten long-standing problems in mathematics and theoretical computer science. These were not simple textbook exercises. Some of these problems had remained open for decades. The work covered areas including geometry, coding theory, group theory, quantum complexity, cryptography, and combinatorics. And this time there is something important behind the announcement.
The proofs were formalized in Lean and released as machine-checkable certificates. In other words, we don’t have to simply believe what OpenAI says. The formal proofs can be checked independently. That changes the argument – that blew my mind!
This is not the same old AI story
We have already seen claims that AI “solved” difficult mathematical problems that turned out to be much less impressive after someone looked closely. Finding an existing answer in a huge collection of mathematical literature is not the same thing as discovering a new answer. It is the difference between finding a book on a shelf and writing a new book. The Astra results appear to be different.
There is still plenty to question. The model is not publicly available. The work has not gone through the complete process of peer review. And there are legitimate questions about how much previous human work contributed to some of the results. Those questions matter but they do not erase the central fact that we now have mathematical results that can be formally checked.
The interesting part is not that AI can do math
I think we are looking at the wrong question. The question is not: Can AI do mathematics?
Clearly, it can do a lot of it. The more important question is: Can AI discover something that was not already known?
That is a much higher bar and if the answer is increasingly becoming ‘Yes’, then we have crossed an important line. A calculator is useful because it follows instructions. A discovery machine is different. It does not simply follow the road. It finds another road!
But don’t call this the singularity
There is a temptation to look at this and say: “This is it. The machines are becoming smarter than us.”
I don’t think that conclusion follows. Mathematics has something that most of the real world does not have: a perfect referee.
A mathematical proof can be checked, either it works or it doesn’t. That gives an AI system an unusually clean environment in which to search, try things, fail, try again, and receive an exact answer about whether it succeeded.
The real world is not like that. There is no Lean compiler for management. There is no perfect verifier for foreign policy. There is no machine that can tell you whether a business decision will still look intelligent three years from now.
Mathematics is therefore one of the easiest places for AI to demonstrate extremely high capability. That does not make the achievement less impressive. It just means we should not confuse one kind of intelligence with all intelligence.
The uncomfortable part
There is another issue here that I find more interesting. If a machine can discover a mathematical result that takes a human mathematician years to discover, what exactly is the human’s role?
We have traditionally treated intelligence as the ability to produce the answer. That definition may no longer work. A human may define the problem. The machine may discover the solution. Another human may verify what the machine did…
Who is the mathematician?
Maybe the answer is no longer as simple as we thought. It is similar to giving a mechanic a machine that can diagnose and repair an engine faster than the mechanic can. The mechanic doesn’t become useless but the definition of being a mechanic change.
The real revolution may be elsewhere
I don’t think the biggest consequence of AI mathematics is that mathematicians will disappear. The bigger consequence may be that the amount of mathematics humans can explore will increase dramatically. A human can spend years on one difficult problem. A machine can work on thousands of possible approaches. Most will fail. It doesn’t matter. If the machine can find one path that works, the failed attempts are almost irrelevant. That is where the economics of discovery start to change and this is why the Astra results deserve attention. Not because AI has suddenly become a human mathematician – it hasn’t. But because we are seeing evidence that a machine can sometimes move from knowing what humans know to finding something humans did not know.
That is a much bigger change.
We should be careful with the word “discovery”
There is still an important human responsibility here. A formal proof can tell us that the logic is correct for the statement that was formalized. It cannot automatically tell us that the machine formalized exactly the problem the mathematicians thought they were solving. That distinction is easy to miss. The machine can prove the wrong thing perfectly. So there are really two questions:
Did the proof work?
And did we prove the right problem?
The first can be checked by a machine. The second still needs humans. That is not a weakness of AI. It is simply a reminder that verification and understanding are not the same thing.
So what changed?
Something did. Not the singularity. Not artificial general intelligence. Not the end of mathematicians…something more specific and, in my opinion, more important.
AI is beginning to show signs of being a discovery system, not merely an information system. That distinction matters.
For decades, computers helped humans calculate faster. Then they helped humans search faster. Then they helped humans write faster. Now they are beginning to help discover things that humans did not know and once machines become good at discovery, the question is no longer whether humans can compete with machines at solving individual problems.
The question becomes: What happens when humans and machines stop solving problems separately?
That is where this gets really interesting 🙂