Shanmu Jin was not trying to make mathematical history. He is a neurosurgery resident at Peking Union Medical College Hospital in Beijing, he studied geology as an undergraduate before medicine, and he had been teaching himself matrix analysis because he needed it for research on transcranial ultrasound. Along the way he ran into a problem with an unusually simple statement, got interested, and pointed a language model at it.
He locked the model off the web, started an autonomous run, and left. Sixteen hours later it had produced a proof of Crouzeix’s conjecture, which had been open since 2004 and had survived twenty two years of attention from people who do this for a living.
The short version
- Who: Shanmu Jin, a resident physician in neurosurgery, self-taught in the relevant mathematics
- What: a proof of Crouzeix’s conjecture, proposed by Michel Crouzeix in 2004
- How: a single sixteen hour autonomous run of GPT-5.6-Sol on the ChatGPT Work platform, with web access disabled
- Who checked it: Anne Greenbaum at the University of Washington and Alex Townsend at Cornell, plus Michel Crouzeix himself
- Status: the reviewers say the argument holds. Formal peer review has not finished
What the conjecture actually says
You do not need graduate mathematics to follow the shape of it, which is part of why it attracted an outsider in the first place.
Take a square grid of numbers, a matrix. Every matrix has an associated region in the complex plane called its numerical range, which you can think of loosely as the territory the matrix occupies. Now apply some function to that matrix. Crouzeix’s conjecture says the result never grows more than twice as large as the biggest value that function reaches anywhere on that territory.
Twice. That is the whole claim. A clean factor of two, with no dependence on how big the matrix is, which is exactly the property that makes it useful and exactly the property that made it stubborn.
The number kept shrinking, and then it stopped
The reason mathematicians cared is that a bound of this kind was already known to exist. The argument was only ever about how small the constant could go, and the history of that argument reads like a stalled negotiation.
Crouzeix proved a bound of 11.08 in 2007. That was already a significant result, because the constant did not depend on the size of the matrix. Crouzeix and Palencia later brought it down to one plus the square root of two, roughly 2.414. Numerically, everyone could see the true answer looked like 2. Nobody could show it.
And the remaining gap was not the kind you close by grinding harder. Researchers working on it had established that the existing proof technique could not be pushed below 2.414 by tightening a lemma. Getting to 2 required a different idea, and for years nobody produced one.
The part that should have made everyone suspicious
Alex Townsend at Cornell had been trying to crack this himself, using various GPT models, for more than a year. Nothing worked. Then he asked a model about the problem and it pointed him at a preprint by someone he had never heard of, claiming the whole thing was done.
He and Anne Greenbaum at the University of Washington read it the way any working mathematician reads a claimed proof of a famous open problem, which is with the assumption that it is wrong. Townsend has been candid about this: the conjecture had been open for more than two decades, so they started from skepticism. A few hours in, they realized the argument was the real deal.
That sentence is the most important one in this story. Not the sixteen hours. Not the model version. The fact that two people who had every professional incentive to find a hole spent hours looking for one and came away convinced, and then Michel Crouzeix, the man whose name is on the problem, reviewed it and agreed.
What this is, and what it is not
The temptation is to file this under “AI solves famous math problem” and move on. That framing loses most of what is interesting.
| Settled | Not settled |
|---|---|
| Experts who wanted it to be wrong have read it and think it is right | Formal peer review, which is where subtle problems usually surface |
| The model produced the key argument during an unattended run | How much of the framing, setup and prompting was Jin’s contribution |
| Jin had no formal training in this area of mathematics | Whether this generalizes or was a favorable match of problem to method |
| The run had no web access, so it was not retrieving an existing answer | Whether anyone can reproduce the result with a comparable run |
Jin himself has been notably modest about it, saying there was certainly an element of luck involved in finding the key idea. Coming from the person whose name is on the paper, that is worth taking seriously rather than treating as false humility. Sixteen hours of autonomous work is an enormous number of attempted paths, and the run that finds the right one may look nothing like the ninety nine that do not.
Why an outsider, and why this problem
There is a pattern here that keeps showing up and rarely gets named. Crouzeix’s conjecture has an unusually simple statement and a very visual geometry, which is exactly what drew Jin to it while he was teaching himself the subject. He had no sense of which approaches the field had already exhausted, because he had not read twenty years of failed attempts.
That is normally a disadvantage. Paired with a system that can explore an enormous number of routes without getting bored or discouraged, it becomes something closer to an advantage. An expert knows which doors are locked. Someone who does not know can afford to have a machine try all of them.
It is not the first time this year that a language model has been credited with real progress on old problems. We covered it when a model cleared ten decades-old math problems for about $2,000 in compute, and then got flagged internally as a cyber risk for the same capability. What makes this one different is the author. That was a lab demonstrating its own system on problems it chose. This was a hospital resident with a geology degree, working alone on something he stumbled into.
The uncomfortable question underneath
Mathematics has spent centuries building a verification system that assumes a human wrote the argument and can explain why every step is there. That assumption is what makes peer review work. A referee can ask an author what they were thinking at line forty, and get an answer.
Nobody can ask a sixteen hour autonomous run what it was thinking. The proof either stands on its own or it does not, and checking it is now a matter of reading it line by line without any recourse to the reasoning that produced it. In this case that worked, because the proof turned out to be short enough and clear enough for three experts to verify by hand. That will not always be true, and the field has no settled answer for what happens when it is not.
The credit question is going to get uncomfortable too. Jin ran the model, framed the problem, recognized what he had, and wrote it up. The model found the argument. Neither did it alone, and there is no established convention for describing that split, which is a version of the same accounting problem that has followed automated work around for twenty years. Amazon named its human labor marketplace after a chess playing hoax and called it artificial artificial intelligence, which was funny right up until the distinction started to matter.
For now the honest summary is small and strange. A twenty two year old problem in matrix analysis appears to be closed. The person who closed it operates on brains for a living, the tool that did the heavy lifting ran unattended overnight, and the mathematicians who spent hours trying to break it could not.

