Correct Up to a Constant
![]() |
| Copyright: Sanjay Basu |
On Terence Tao, six ideas that have been quietly running the world for four thousand years, and what happens when the machines start doing the easy parts
I have had a long and complicated relationship with mathematics, and I want to be precise about which parts were love and which parts were hate.
The love was the concepts. I could see them. Somebody explained a derivative to me once as a rate of change and I did not need three worked examples afterward. I felt the thing tilt. Years later when I understood that a matrix is a machine that stretches and rotates space, and that eigenvectors are just the directions the machine cannot be bothered to turn, I sat with that for about a week and was insufferable at dinner. Probability made sense to me the first time somebody framed it as bookkeeping for ignorance. I have never had trouble with the ideas.
The hate was the exam hall.
My marks were a scandal. Not a catastrophe, which would at least have been dramatic. A scandal, which is worse, because it invites explanation. I would understand the entire architecture of a problem, set it up correctly, choose a sensible method, and then somewhere around line four I would drop a minus sign into the void and never see it again. My friends in physics used to console me by saying that a factor of two is a rounding error in astrophysics. In a three hour paper it is a zero.
I once lost seven marks for a missing dx. Seven. For a differential the size of a mosquito.
There is an old joke about a mathematician, a physicist, and an engineer being asked to prove that all odd numbers greater than two are prime. The physicist says three is prime, five is prime, seven is prime, nine is experimental error. I was the student who wrote nine is experimental error on an actual answer sheet and genuinely believed the examiner would recognize the intellectual honesty involved. He did not. He circled it in red with a small vertical stroke that I have thought about for thirty years.
What I would tell my younger self, if the transmission were possible and the boy were willing to listen, is that I was performing well. I had excellent asymptotic behavior. I converged reliably on the correct answer, just not inside the time limit. In the language of the discipline I loved and could not satisfy, I was correct up to a constant.
I bring all this up because I have just spent a weekend listening to a talk Terence Tao gave about his forthcoming book, Six Math Essentials, and it did something to me that mathematics has not done in a long while. It made me feel like the intuition was the point after all.
A word about Tao before we go further
I want to say something about the man and then get out of the way.
Tao got a gold medal at the International Mathematical Olympiad when he was thirteen years old. Thirteen. He had a PhD from Princeton at twenty one and a Fields Medal at thirty one. On any reasonable accounting of raw ability he is one of the two or three most gifted mathematicians alive, and probably the most versatile, in the sense that he can be dropped into a field he has never worked in and be useful in it within a month.
None of that is why I admire him. There are prodigies who curdle. The ones who become gatekeepers, who use the gap between what they see and what you see as a form of social leverage. Tao is the opposite of that. He has spent twenty years writing blog posts that explain things to people who are not him, running collaborative open problem projects where anyone can contribute, and, in the talk I heard this weekend, sitting in front of a camera and explaining what a negative number is. Without condescension. With actual pleasure.
The generosity is the thing. Genius is a lottery outcome. Generosity is a decision.
The six things
The book is built on six ideas. Numbers, algebra, geometry, probability, analysis, and dynamics. Every one of them is something a ten year old is exposed to. Every one of them has been developed into something that would be unrecognizable to that ten year old and that still, if you strip the notation away, is the same intuitive idea it started as.
That framing alone reorganized something in my head. Because what Tao is arguing, and I think he is right, is that the technical apparatus of mathematics is not the content. It is the packaging that makes the content transportable. The mathematics is a language for being careful about things you already half understand.
Which means my problem was never that I did not understand mathematics. My problem was that I was fluent in the intuition and illiterate in the transport layer. Anyone who has ever built a system with a beautiful architecture and a broken serialization format will know exactly the feeling.
Numbers, and the strange loyalty of a pattern
Start with counting sheep. You add sheep, you subtract sheep. Very quickly you notice you can always add and you cannot always subtract, because four sheep do not come out of a herd of three. And then something remarkable happens, which Tao describes almost casually and which I think is one of the deepest things in the whole talk. The arithmetic patterns are so regular that they insist there ought to be an answer. Not the world. The pattern. The pattern demands a number that the world does not supply.
So we invented negative numbers. And zero, which took an embarrassingly long time and which my own civilization (associated by ethnicity) has a certain proprietary feeling about. And fractions. And then irrationals, which arrived as a genuine scandal, and whose name is Latin for insane rather than for unreasonable, a detail I did not know and which improved my mood considerably. And then complex numbers, invented purely so that we could keep solving equations that had no business having solutions.
And then quantum mechanics turned up three hundred years later and could not be written down without them.
This is the pattern that repeats through the entire history of the subject. Somebody follows the internal logic of a formal system past the point where it corresponds to anything. They do it because the symmetry is pretty. And decades or centuries later the universe turns out to have been using that system all along.
I find this genuinely unsettling and I do not think anyone has explained it.
Algebra, and the wine merchants who had already solved calculus
The best story in the talk is Kepler at the wine market.
Wine was sold by the barrel, and barrels vary. Tall and thin, short and wide. So how do you price one. The market official had a stick with markings on it. He would push the stick through the bunghole in the side of the barrel down to the far bottom corner, read where the wine line hit the stick, and announce a volume. One diagonal measurement. Done.
Kepler could not let it go. He went home and did the algebra, radius r, height h, one measured diagonal. And he found that the diagonal alone does not determine the volume. It should not have worked. Then he added one more assumption, which is that the merchants were not idiots and would have converged on the barrel shape that holds the most wine for a given stick reading. He did some early and rough version of what Newton and Leibniz would later formalize as calculus. And the shape that fell out of the optimization matched the barrels actually being sold, to high precision.
Sit with that for a second. An illiterate guild of coopers, over a few centuries of trial and error and competitive pressure, had found the maximum of a function. They did not know it was a function. They knew which barrels sold.
I think about this constantly in my day job. I spend my working life around infrastructure, and the number of times I have found that some operations team has empirically converged on a configuration that turns out to be provably near optimal, for reasons nobody on that team could articulate, is not small. The market runs gradient descent whether or not anybody writes the gradient down. The mathematics does not create the answer. It explains the answer and, more importantly, tells you when the answer will stop working.
Geometry, and the axiom nobody liked
Euclid laid down five axioms. Four of them are the kind of thing you nod at. Given two points there is a line through them. Fine. The fifth one, the parallel postulate, was long and ugly and everyone hated it, including, it seems, Euclid. For two thousand years mathematicians tried to derive it from the other four so they could stop looking at it.
They failed, because it is not derivable. It is a choice. Change it and you get spherical geometry, where there are no parallel lines at all, because every pair of great circles meets. Change it differently and you get hyperbolic geometry, where lines that start parallel drift apart forever.
Once the floodgate opened, people built geometries where you can travel a loop and come back left handed. Riemann assembled the whole zoo into one language, for no practical reason whatsoever, because it was interesting.
And then Einstein needed to say that mass bends spacetime and had no vocabulary for it, and asked a friend, and the friend said there is this fellow Riemann. The equations of general relativity are almost trivial to state in that language. Curvature is proportional to mass and energy. That is the sentence. Solving it is a nightmare and we can barely simulate two colliding black holes on a modern supercomputer, but stating it takes one line, and the line only exists because somebody in the 1850s was annoyed by an inelegant axiom.
The lesson I take from this is not about geometry. It is that the thing in your system that everybody works around, the ugly component that nobody wants to touch, is load bearing. It is not ugly by accident. It is ugly because it is doing something the rest of the design cannot do, and the day you finally understand why is the day the whole architecture opens up.
Probability and analysis, or why the doubling strategy always sounds so good
Probability began, as Tao points out, with gamblers writing letters to their mathematician friends asking for help cheating more effectively. This is my favorite origin story of any field. Number theory started with somebody trying to win at dice.
Analysis is where I want to spend a moment, because analysis is the mathematics of error bars and infinities, and it contains the most useful practical warning I know.
Consider the doubling strategy at roulette. Bet a dollar on red. If you lose, bet two. Lose again, bet four. The first time you win you recover everything and profit a dollar. It cannot fail. It is arithmetically airtight and it has ruined an enormous number of people.
The flaw is not in the arithmetic. The flaw is that the strategy requires an infinite bankroll. What it actually does is take the risk of losing and compress it into a tiny, very unlikely, absolutely catastrophic event. You win a dollar a thousand times and then, on some Tuesday, you are betting your house.
I have watched this exact structure in system design. Retry policies that assume unbounded budget. Caching layers that work perfectly until the one cold start that takes down the region. Financial models that are correct on ninety nine percent of days. Analysis is the discipline that tells you where the tail lives, and the tail is always where the money is.
The other thing Tao does with infinity is the monkey. Give a monkey a typewriter and infinite time and it will eventually produce Hamlet. This is true and everyone knows it. What almost nobody has internalized is the timescale. A four letter word might take an hour. A seven letter word might take years. A single page of Hamlet takes more than the age of the universe. Infinity is not a big number. It is a placeholder for a number larger than any number you can name, and the difference between the two matters enormously when you have a budget.
And then there is the line in the talk that I have not stopped thinking about. Tao says that as a kid he played computer games, and that sometimes the right move is to play the game first with the cheat codes on, infinite ammunition, infinite health, just to see the shape of the solution. Then you play it properly and figure out how to do it with real resources.
He says this is how a lot of mathematics gets done. Assume no friction. Assume infinite energy. Solve the idealized problem. Then work your way back down to the finite world and see which parts survive.
And he says something else that I wish somebody had told me at nineteen. Mathematics has the freedom to fail, because failure is cheap. A surgeon who cuts wrong has committed a tragedy. A businessman who bets wrong has lost a company. A mathematician who assumes wrong has lost an afternoon.
Dynamics, and the traffic jam that is not there
Dynamics is the study of what simple rules do when you run them for a long time. Tao lives in Los Angeles and uses traffic, which is correct. Every driver is running a trivial rule. Speed up if there is room, slow down if there is not. Nobody is doing anything complicated. And out of that you get compression waves that travel backward through the flow, so that you can sit in a dead stop on the freeway with no accident, no obstruction, no cause visible anywhere, because something happened two hours ago and the wave has not finished dissipating.
Newton solved the two body problem exactly and it was one of the great triumphs of the human mind. He then tried the three body problem and reportedly said it was the only problem that ever gave him a headache. There is no closed form solution. There probably cannot be one. Run the numbers and the orbits look periodic for a long stretch and then drift, and then drift again.
We now think our own solar system used to have more planets than it does. Small gravitational nudges accumulating over millions of years, until something got flung out or two things met. The asteroid belt may be the debris. The most stable and predictable system any of us can point to is, on a long enough timeline, chaotic.
And the punchline is this. Even when you start with a system that is completely deterministic, with no randomness anywhere in it, the best available model for its long term behavior is probabilistic. Determinism does not buy you predictability. That is a philosophical result of the first order and it arrived as an accident of trying to compute where Jupiter would be.
Curiosity, and my great grand uncle
There is a thread running through all of this that I want to name, because it is personal.
Every single application in the list arrived decades or centuries after the useless research that made it possible. Complex numbers before quantum mechanics. Riemannian geometry before relativity. Sphere packing, which started with a British sailor wondering how to stack cannonballs efficiently in a hold and which Kepler answered with the arrangement you see in every supermarket orange display, packing about seventy four percent of available space. Kepler could not prove it was optimal. Nobody could, for nearly four hundred years, until a computer assisted proof in 1998 that the referees said they could not fully check, and then a formal machine verification in 2014 that finally settled it.
And then it turned out that if you take that problem and move it into thousands of dimensions over strings of bits, it becomes the problem of keeping wireless signals from being mistaken for each other. The entire telecommunications industry, and the pricing of spectrum auctions in the billions, sits on top of a sailor’s question about cannonballs.
Tao’s own contribution to this pattern is compressed sensing, which he describes with a modesty I find slightly maddening. A statistician and an electrical engineer, Emmanuel Candès and Justin Romberg, were trying to shorten MRI scans, which then took three minutes of a patient lying still, long enough that small children had to be sedated. They tried a nonstandard reconstruction method on undersampled data and got back an almost perfect image. Tao’s first reaction was that they had made a mistake, and he went home to prove that the information was simply not there. Halfway through writing the proof it collapsed and showed him the opposite. The method worked, and there was a reason.
Seismologists had stumbled on something similar. So had astronomers. Each field thought it had a local trick. Once the mathematics was understood, it stopped being three tricks and became one theory, and it is now taught next to least squares.
I have a family connection to Jagadish Chandra Bose, and I grew up hearing about a man who built millimeter wave apparatus in Kolkata in the 1890s and refused to patent any of it, and who spent the second half of his life measuring the electrical responses of plants because he wanted to know, not because anyone had asked. Almost none of it was useful in his lifetime. A great deal of it is useful now.
So when people ask what basic research is for, I have a family answer and a mathematical one, and they agree.
And now the machines
The last third of the talk is about AI, and it is the most careful thing I have read on the subject in months, mostly because Tao is not selling anything.
His central observation is that the discourse is one dimensional. There are easy problems and hard problems, humans go up to here and machines go up to there, so who wins. He thinks this framing is wrong and that the actual relationship is orthogonal.
Humans do depth. A mathematician picks one or two problems a decade, hard enough to be interesting and not so hard as to be hopeless, and grinds on them for years, and the byproducts of the grinding are worth more than the answer. AI does breadth. Point it at a thousand problems and it will fail on the deep ones, because it is essentially guessing, but somewhere in that thousand there are two hundred where a known method applies and nobody has had the time to try. Some obscure 1970 paper has the key. No human expert has the patience to check every combination. The machine has nothing but patience.
Five percent of a thousand is fifty solved problems. They may not be the fifty you most wanted. It is still fifty.
Tao’s description of how these things actually work is the most honest I have seen from someone at his level. They are curve fitting on the next word, and the surprise is not that they can do it but that iterating the process stays coherent instead of degenerating into the gibberish your phone keyboard produces. He compares working with one to working with a collaborator who knows an enormous amount and is slightly drunk, throwing out ideas, some of them worthless, and who with enough supervision produces something real. I have not read a better description and I doubt I will.
I want to add my own data to this, because I have some.
We ran a harness at work for a while, two engineers and four coding agent instances. The breadth and depth split showed up almost immediately and almost exactly as Tao describes it. The agents were extraordinary at the class of problem where a solution exists somewhere and needs to be located and adapted. They were useless on the problems where the difficulty was conceptual, where you had to decide what the system should be rather than how to build the thing already decided. And the bottleneck moved. It moved off generation entirely and landed on review.
Tao has a name for this from the mathematical side and it is the best phrase in the whole talk. Proof indigestion.
Generating proofs used to be hard. Verifying them used to be hard. Both are getting automated. What has not been automated, and what he thinks may not be automatable in any useful sense, is the digestion. Somebody has to read the thing, decide it matters, work out why it matters, reorganize it into a form that can be taught, and put it in a textbook. AI generated proofs, he notes, spend enormous effort on trivial steps and skate over the interesting one, because to a brute force process everything costs the same. A human who suffered at the hard step writes differently about the hard step.
He has stopped trying to keep up with his own field. Terence Tao. Stopped trying to keep up.
Change the nouns and that is my last eighteen months. Pull requests generated faster than they can be understood. Documents that exist and are correct and that nobody has metabolized. We solved production and left consumption exactly where it was.
The Kepler problem, restated for our moment
Here is the part that should worry people more than it does.
Kepler’s first theory of the solar system was that the six known planets sat on nested spheres with the five Platonic solids fitted between them. It is a gorgeous idea. It is completely wrong. He held it for years, and only gave it up because he got his hands on Tycho Brahe’s observational data, possibly by theft, and the numbers would not fit no matter how he pushed.
And before that, Copernicus. The heliocentric model, when first proposed, made worse predictions than the geocentric model it replaced. The Ptolemaic system had been tuned for over a millennium by Greek, Arab, and Indian astronomers and it was very good at its job. Copernicus was right and lost on the benchmark.
Tao’s point lands hard. If Kepler and Copernicus had been running model selection, the correct theory would have been pruned in the first round for underperforming.
That is what overfitting is. Not a technical footnote. The structural risk of any system that optimizes for fit against currently available data, which is precisely what we have built and precisely what we are now aiming at science. The right answer often arrives ugly and losing.
And then the seed corn question. The problems we hand to first year graduate students exist to make graduate students, not to make papers. If the machine writes those papers, we get the papers and no mathematicians. Nobody has proposed a serious answer to this and I do not have one either.
What I am not sure about
I do not know whether proof indigestion is a transitional problem or a permanent one. It is possible that in four years we have decent automated curation and this whole complaint reads like somebody in 1996 worrying that the web has too many pages.
I do not know how replicable the recent results are. As Tao says, the companies do not publish what they spent. When a model solves a genuinely open problem, we do not know if that was the only problem it was pointed at or the one success in three hundred attempts, or whether it cost a thousand dollars or a million. The FrontierMath results he mentions, where the best models handle five or six out of ten research level problems, are the closest thing we have to an honest measurement, and ten problems is ten problems.
And I do not know whether my discomfort with the helicopter is a real epistemic objection or nostalgia wearing a lab coat. Tao’s image is that you set out to hike to a waterfall, and you get lost, and on the way you find something else worth noting, and you see a second waterfall in the distance that you cannot reach yet but somebody will. The AI is a helicopter that takes you to the waterfall and back. You saw it. You learned nothing about how to get there.
I believe that. I also notice that I never complain about the helicopter when I am tired.
Back to the exam hall
I have been carrying around the idea that I was bad at mathematics for about thirty five years, and the talk did not exactly dissolve it, but it moved it.
Because what Tao says about how he works is that there is no thunderbolt. He tries something, it fails. He tries something else, it half works and then jams, and now he knows where one obstacle is. He finds a smaller problem with the same obstacle. He goes back and forth for months mapping the negative space, all the approaches that cannot work, until the path is clear by elimination. And when the answer finally comes, it does not feel like triumph. It feels like embarrassment. How did I miss this. I was so stupid.
He also points out the contradiction sitting at the center of mathematical education, which I have been personally acquainted with since I was fourteen. We grade outcomes with absolute severity. One sign error, zero marks. And we produce students who are terrified of being wrong. But the process of doing actual mathematics is the exact opposite. It is being wrong repeatedly and on purpose, because you cannot know why the clever thing works until you have watched the stupid thing fail. Niels Bohr said an expert is a person who has made every possible mistake in a very narrow field.
I made a lot of mistakes in a very narrow field. Nobody told me that was the job.
So here is where I have landed, on a Sunday evening, with an essay full of annotations and a slightly reorganized childhood. Kepler spent years defending nested Platonic solids before the data broke him and he found the ellipse. He was not being stupid. He was being wrong slowly, which is the only way anyone has ever been right.
The machines are very fast now. They are going to be faster. What they do not yet have, and what I am not sure how to give them, is the capacity to be wrong slowly. To hold a beautiful bad theory long enough to learn what specifically is bad about it. To lose seven marks for a missing dx and think about it for thirty years.
I do not know if that capacity is precious or just expensive.
Ask me again when the constant is resolved.

Comments
Post a Comment