Rendered at 23:48:43 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ZeljkoS 13 hours ago [-]
To me, the real lead of the article is buried here:
> apparently they tried the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.
Back-of-the-envelope cost calculation:
3 hours of GPT-Pro compute per attempt / 5% success rate → 60 hours of compute per solved open math problem
At GPT-6 Astra API long-context pricing ($75/M output tokens), assuming 50 reasoning tokens/s, that's roughly $810 per solution:
60 h × 3,600 s × 50 tokens/s × $75 / 1,000,000 = $810
Of course, actual token usage is unknown. But even if off by 10x, $8100 is cheap for solving a longstanding mathematical problem.
WiSaGaN 9 hours ago [-]
This is with the unreleased model, which is probably significantly larger (thus costlier) than astra.
utopcell 5 hours ago [-]
The argument is that it's not tens of $m.
DonsDiscountGas 13 hours ago [-]
They probably spent more than 3 hours on the failures, which could easily 10x that cost number. Still ridiculously cheap.
Balgair 11 hours ago [-]
As bureaucracies look at AI and staffing, this is going to be the headline in the executive summary.
Sure yeah, today it's only a 5% solution (to the hardest problems), and we know it will top out somewhere.
But the cost reduction is just too good to ignore.
Terry Tao was saying that we need more mathematicians [0]. My paraphrase of his article is that we will need them to check the AIs, essentially (I missed things didn't I?).
But with a near 10,000X reduction in cost [1] today, it's really hard to argue for that when budgets are continuously tight.
[1] assume ~$150k/year for a mathematician of the caliber that can do this work. 2x that for overhead like healthcare, 401k, etc. Do this for 35 years of work = $10,500,000. Assume about 1 result of this caliber per mathematician. So ~$10M per result. OpenAi is saying they can do it for ~$1k, a 10^4 X decrease.
vessenes 9 hours ago [-]
I think it's not just to check - it's to interpret, simplify, apply taste, basically get these proofs into a form that humans can grok and build on. AI could do all of that but the grok-ing -- as we are tapping the font of cheap proofs it still behooves humans to keep track of what we find interesting, keep internalizing and learning concepts and make sure whatever we're building right now will be generationally useful. At least that's my take
fisf 6 hours ago [-]
>I think it's not just to check - it's to interpret, simplify, apply taste, basically get these proofs into a form that humans can grok and build on.
Sure, but we are approaching a time at which we will have to answer a "why" here. I.e. anything beyond checking and validation becomes of mostly artistic value at some point. And there is nothing wrong with that.
vessenes 5 hours ago [-]
I’d say the artistic value is the main value in general —- if you know mathematicians they will often refer to a ‘beautiful’ proof, aesthetics is a critical part of math.
An early (easy to understand) proof that fits the bill for me is proof that the sqrt(2) is irrational: a very natural thing to think about —- for a unit square, what is the measure of the diagonal? And people were killed in ancient Greece over the proof, but it’s comprehensible with basic algebra.
jchanimal 5 hours ago [-]
IMHO it's groking that allows higher level metaphors and eventually more powerful theories to be built. If all you are doing is proving everything reachable by construction, you aren't really learning, as this process won't be drawn to the more compressed expressions by necessity but only as a "please" from the humans.
computerdork 7 hours ago [-]
Seems similar to what is happening with software developers.
robotpepi 5 hours ago [-]
You're off by an order of magnitude at least. A mathematician does much more than just prove theorems. Let's start by teaching and administration duties.
And then you're off an order of magnitude of the efficacy of OpenAI. We don't know how much work of mathematicians is behind all these results, what's the real cost (which includes those problems they couldn't solve), etc.
Balgair 2 hours ago [-]
Oh I entirely agree! Mathematicians are full human beings.
But a soulless admin is not.
And they care only about the budget.
And the real danger here is that these large groups of blame avoidance mechanisms take at best half a look at the cost of an AI and the cost of a pre tenure track new assistant professor (or whatever), and they decide, oh maybe next year we can hire someone.
And then they do that again next year.
And then the next
indoordin0saur 6 hours ago [-]
But all of these proofs are unreadable and seemed to have hallucinated in a bunch of irrelevant stuff, making them not especially useful.
Balgair 2 hours ago [-]
I agree with you.
The problem is getting the funding orgs to care.
HankB99 9 hours ago [-]
> 5% solution (to the hardest problems),
Isn't that more likely the easiest problems in the set? I suppose the hardest problems compared to those that have been solved.
Otherwise your points stand.
bbmatryoshka 9 hours ago [-]
If we have a jagged frontier inside the domain of math (very plausible, the jagged frontier concept seems to be quite fractal), the model solved the easier problems in his area of major strenght, some of them will be in the bottom 5% of the easiest problems in the full set, and some in remaining 95% of problems
pbronez 12 hours ago [-]
Yeah the README on OpenAI/Math says:
“On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems.”
That leaves a lot of ambiguity. Is “result” exclusively “positive result shown here” or “all results”?
Since proving infeasibility is a valuable result, the unpublished problems must not have reached a valuable end state. Thus it’s an operator decision of when to turn the machine off and try another problem. I’d thought they had a three hour time box on this, but apparently not.
ball_of_lint 6 hours ago [-]
You also have to count the amount of tokens and human time spent afterward to verify these proofs. There's already been some papers retracted and it seems likely to me that there will be many more as people continue digging into the more dubious ones.
miguelacevedo 1 hours ago [-]
wild.
measurablefunc 8 hours ago [-]
There has been no independent verification. I am certain their Navier-Stokes proof is incorrect.
onraglanroad 5 hours ago [-]
It provided a formal proof in Lean.
measurablefunc 5 hours ago [-]
Which is not the same thing that was written in informal prose & their proof has 8k non-computable axioms so you can't even write a program to verify their blowup example is actually a valid instance of a smooth flow that blows up.
pixelpoet 10 hours ago [-]
It's "bury the lede" BTW
eddieroger 10 hours ago [-]
They share an origin and both work, even if lede is the currently more frequently used one.
This comment seems to be a hypercorrection, where someone attempts to demonstrate superior linguistic knowledge by correcting something that isn't actually wrong in the first place. You've likely heard somewhere before, almost certainly on HackerNews in fact, "lede" is proper and "lead" is not but don't really have a genuine understanding of it and are just repeating what you've heard before.
In actuality both "lede" and "lead" are correct. "Lede" is mostly an American and recent respelling of the word "lead" and "lead" still remains the predominantly used term (while "lede" is predominantly an American journalistic spelling).
brainwad 7 hours ago [-]
The expression comes from journalism, hence why lede is preferred. It's something one should not do when writing an article - bury the most important point in the middle of the text.
I heard they used lede to distinguish from (hot) lead. But since linotype has been dead for 50 years this is probably moot now.
dboreham 7 hours ago [-]
Only in American English.
brainwad 7 hours ago [-]
I don't speak en-US (or neighbours) and I'd also say lede.
gosub100 8 hours ago [-]
Language is dynamic and the meaning and spelling of words changes over time.
greedo 10 hours ago [-]
It's common in journalism to call to use lead.
ks2048 1 days ago [-]
> It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
TheOtherHobbes 1 days ago [-]
Math proofs need to produce the correct output correctly, which is not quite the same thing.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
oliculipolicula 13 hours ago [-]
This, but for a generalised notion of "number" (that is, "proof")
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
pizza234 1 days ago [-]
There's also another (that I find more concerning) aspect to it.
As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap.
Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.
koliber 14 hours ago [-]
Exactly. We need "observability tools" that allow us to keep a handle on things we can't see or comprehend innately.
We have such tools for a large number of events that we're incapable of witnessing or understanding in "raw form".
We can't possibly read through thousands of raw records of recordings of individual's heights and make sense of it. But we can use Excel to calculate the average in seconds. We can chart the results and the image is perfectly comprehensible. We can trust the average calculation and the image because we trust the mechanism for transforming the data. We have what I would call a trusted path of provenance. We don't need to nor do we want to read the raw data.
Now we need things like this of the 2nd order. We need mechanisms of transforming trusted paths of provenance into things we can look at and easily verify.
We can do formal verification of code that would be hard to do by hand. We need formal verification tools for formally verifying formal verification tools. Eventually we will need multiple layers of this. As long as the chain and reasoning is intact, we should be alright.
It's kind of like with a horse. A horse is much stronger than us, but we can control it by pulling on two strategically connected pieces of rope.
dist-epoch 1 days ago [-]
Arguably no human understands 100% of how a smartphone is produced, and, it doesn't matter?
Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.
ben_w 12 hours ago [-]
While any logical proof written in lean can be broken down into things a human can follow, there is no requirement that the whole proof of something is ever simple enough that all humans combined could follow it.
Given the LLMs are struggling to explain themselves clearly, but also that these explanations cleared up somewhat by having an expert using one to query some of this research, but also the LLMs can solve problems faster than humans can read the proofs, it is possible for this category to have both examples of things that can be rendered in a human-comprehensible form, and also examples where it is not.
As an analogy: any human can check any single arithmetical calculation from a computer, that's not even particularly difficult. But a Raspberry Pi Zero can do those arithmetical calculations so fast that even if literally every single human was trained to do this at the level of the current world record holder, humanity as a whole could not keep up.
Tanjreeve 17 hours ago [-]
The point isn't that 1 person can understand an entire smartphone silicon up. It's that 100% of a smartphone can be understood by people. A maths theorem that can only be partially understood/proven only has value as marketing collateral.
mrob 15 hours ago [-]
>A maths theorem that can only be partially understood/proven only has value as marketing collateral.
There are many problems that aren't "interesting" in a mathematical sense but can still have valuable proofs. The entire field of formal methods in software development is mostly about this kind of problem. If I want to be sure that no input can cause out-of-bounds memory access, or that my superoptimizer found the lowest possible cycle count, I don't care if it's elegant, I just need some machine-readable proof I can run through my trusted verifier.
Tanjreeve 7 hours ago [-]
When we use verification to build software. We have an actual software as an output and hopefully some sort of problem solvsd the verification is also not in an of itself valuable.
mrob 6 hours ago [-]
The verification itself is valuable because it gives us confidence that the software is correct. This knowledge changes what we do with the software. E.g. do we need to sandbox it before processing untrusted inputs? Is it safe to put it on a hostile network? Can we use it to control an airplane or a nuclear reactor? Even if the code remains identical, formally verifying software improves it because it lets us use it in situations where unverified software would be too risky.
jcelerier 11 hours ago [-]
For one person that understands a proof of some matrix multiplication algorithm lowering its current complexity bound in an useful way (e.g. not with absurdly gigantic <= 2nd order polynomial term), there's likely thousands of companies willing to pay for the improved efficiency
Tanjreeve 7 hours ago [-]
And if anything that useful does come out of this output/"slop dump" then that might begin to make a strong argument for it. But every other time it was "that's the theorem done for the headline figuring out everything important and how to use it is an exercise for the reader".
losvedir 8 hours ago [-]
> A maths theorem that can only be partially understood/proven only has value as marketing collateral.
If this is true, then what's really the point of math? A lot of math is actually useful. A proof regarding cryptography, for example, would have practical application even if not understandable by humans.
Tanjreeve 7 hours ago [-]
Cryptography is 100% understandable by humans so I'm not sure what the point is. You do understand that the argument isn't everything has to be understandable by every human right just to check?
claude_sh_1959 1 days ago [-]
[dead]
pnin 14 hours ago [-]
Mathematicians (mostly) accepted that the proof was valid, not that it was the kind of thing mathematics should strive for.
housecarpenter 14 hours ago [-]
I don't know much about this controversy, but looking at the Wikipedia page for the Four Color Theorem, it looks like since the initial Appel-Haken result, mathematicians have continued to work on the Four Color Theorem to try to find a simpler proof. So did they really accept it?
pfdietz 14 hours ago [-]
Finding new simpler proofs happens in traditional manual mathematics too.
rhdunn 13 hours ago [-]
There are a lot of proofs for Pythagoras' Theorem and other theorems.
There are mappings between complex numbers and 2D matrices allowing problems to be solved in either domain.
There's research into The Langlands Program looking to connect number theory and harmonic analysis.
There's research into Category Theory looking to define core concepts and relate them do different fields so that results in one field can be applied to another due to equivalence.
FloorEgg 1 days ago [-]
If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would.
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
perching_aix 1 days ago [-]
> I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.
[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process
esafak 1 days ago [-]
This is just the first cut. I have no doubt that they will polish their proofs over time.
mfld 8 hours ago [-]
My hunch is that improving those proofs specifically might be a similar endeavour as improving a large vibe coded app.
esafak 8 hours ago [-]
It's a search problem; finding the shortest and most elegant path through the space of arguments.
CogDisco 21 hours ago [-]
I don't see how they have any incentive to do so.
esafak 20 hours ago [-]
I communicated loosely; I meant AI users in general, not OpenAI specifically.
slopinthebag 1 days ago [-]
idk if i'd even say they're "lesser", just very different. so they look like gods/babies depending on what they're doing because we anthropomorphise them.
FloorEgg 1 days ago [-]
In terms of synapse density, neuron diversity, energy efficiency, memory access, etc. they are orders of magnitude lesser.
My mental model - for better or worse - is that intelligence has both shape and area, and LLMs are orders of magnitude smaller area but very different shape, and they have more intelligence area in the kind that humans have lesser of.
So yes very different, more in some material ways and lesser in others, but in total intelligence are still orders of magnitude lesser.
My gp comment was acknowledging that when you scale up many instances / brute force problems it confuses that "total area" claim a bit.
To follow the anthropomorphization... 1000 toddlers may have more total intelligence than a grown man, but does that matter?
The problem with these discussions probably/usually fold into differing/loose definitions of intelligence.
slopinthebag 24 hours ago [-]
ok yeah then i think we're in total agreement
perhaps the chat-based ux has sort of fooled us into comparing these things to human intellegence. we don't really do this with chess, or other forms of ai, nor computers at large.
Octoth0rpe 1 days ago [-]
> A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
Not a counterpoint, actually case in point, because:
>Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community.
Proof can't be understood, proof doesn't matter.
pfdietz 14 hours ago [-]
Oh the proof can be understood quite well: it is understood to be wrong.
IsTom 1 days ago [-]
Isn't this controversial, to say the least?
dist-epoch 1 days ago [-]
I wait for AI to say something about this :)
Someone at OpenAI, please, work on this.
fisf 5 hours ago [-]
Funny enough.
There is ongoing AI assisted work on this, and precisely the problematic gap in the theory (around Corollary 3.12).
As of now, a required step appears to definitely be missing, but it's not clear if this is a genuine / unrepairable defect.
> the difference between brute forcing and cognition
There’s not really a clear distinction between these things, in my opinion. Problem solving (and intelligence?) is a mix of search and compression. We like solutions that are elegant (high compression, simple search). But often what appears elegant to some is harder to appreciate for those without the same background knowledge or even the same amount of mental bandwidth (if you’ve ever worked with someone simply much, much smarter than you, you may intuit this!).
disgruntledphd2 15 hours ago [-]
> The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
This is basically how all AI approaches appear to work. They solve the problem you give them (in many cases) but in a very over-complicated way.
It's like they can't step back from the problem and realise that they need to simplify to make it work better. Nope, just keep hammering more code/proofs against the problem and eventually you'll hit the goal.
RL has a lot to answer for, I guess.
fisf 6 hours ago [-]
> It's like they can't step back from the problem and realise that they need to simplify to make it work better. Nope, just keep hammering more code/proofs against the problem and eventually you'll hit the goal.
There is nothing wrong with that approach if it accomplishes the goal faster, i.e. you can work at the speed of an AI.
pizza234 1 days ago [-]
The post says there's a Lean certificate for this and other proofs ("some [...] not all of them").
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
smcg 1 days ago [-]
It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
Kotlopou 1 days ago [-]
But if you think they got them illegitimately, then how did they get them? And why are mathematicians reacting to this as a sudden explosion of new results that have resisted sustained effort? Where is the sudden productivity rise coming from?
za_creature 1 days ago [-]
I'd say it comes from the same mathematicians that were strongly encouraged to use the machine to solve their problems for the last 2 years or so.
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
pfdietz 14 hours ago [-]
> But if you think they got them illegitimately, then how did they get them?
Obviously they stole them from the Proof Fairy.
It's grapes so sour they could etch metal.
pizza234 1 days ago [-]
Have you actually read the article? It's been actually written, among the other things, because the author's wife has been trying to solve one of the problems for her whole life.
amoss 16 hours ago [-]
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
> I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
I doubt that it is. For the language of proofs to be powerful enough to be able to
express an arbitrary proof it would have be Turing Complete. Proving that a proof
is the smallest proof of a given concept (i.e. there is no smaller human understandable proof)
would then be proving the minimality of a program in a Turing Complete language.
sebzim4500 1 days ago [-]
Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible).
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
curt15 1 days ago [-]
Why should that make material difference to the IPO? What is the economic value of those results?
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
coderenegade 1 days ago [-]
They're burying any doubt that the models are capable of superhuman performance on intellectual tasks. Neural nets aren't calculators, and were notably poor at mathematical reasoning tasks for a long time. Now they're not, and the labs are proving that by chewing through what would ordinarily be decades of progress in a month. And the reason to go for math in particular is because there's no wiggle room. You can't just dismiss it as hallucination.
If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. And even if they were only good at this stuff, that's still a tremendously valuable thing, because quantitative reasoning and analysis is the bedrock for many, many industries.
oAI is gunning for the largest IPO in history at this point, and they might actually get there.
lesostep 10 hours ago [-]
> If the models can do this, they're almost certainly good
I would argue that "good at math" was a short hand for "good at X" because mathematicians were historically good at engaging with very complex ideas, distilling them and coming up with precise and concise answers they could validate by themselves.
Given that AI solutions are described as "psychedelical" and they rely on the outside source to validate the result, I don't think the same logic could apply to them.
curt15 23 hours ago [-]
> If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate.
The "then" in your "if-then" bears a heavy load. Why would society assign so little economic value to pure mathematics if the skills for proving math theorems translate to massive value in "just about everything"? Would you expect top mathematicians to cure cancer if you transplanted them from the math department to a medical research lab?
coderenegade 18 hours ago [-]
Because research mathematics is a subset of all quantitative work that gets done, but it's by far the most technically difficult subset. If the models can handle research grade math and produce ironclad proofs, they can probably handle the quantitative side of just about any discipline in a trustworthy fashion. Think about how many dinky spreadsheets have gone on to become critical tooling for large organizations. Even if you consider that many disciplines hide technically demanding work behind tooling (e.g. essentially no one is writing a stiffness matrix FE routine by hand), this model would be capable of writing a direct competitor from scratch to produce the same result.
All of STEM relies on mathematical analysis, and new models are now superhuman at that. And yeah, I'd go a step further and say that the reasoning and creativity required to solve cutting edge math problems probably does translate to other tasks like interpretation of the law, or medical diagnosis, or accounting, etc., for the same reasons that I think most top tier mathematicians would excel at those tasks were they so inclined.
ndriscoll 21 hours ago [-]
Do you think mathematicians are not already working on cancer research? There's quite a bit of heavy math in biostatistics, medical imaging, machine learning, etc.
I'd assume that the majority of people who study math take their skills and move onto some related STEM career that isn't pure math. Academia is incredibly small and competitive.
throwaway81523 18 hours ago [-]
That sort of worked for Eric Lander but as he moved higher up in the bureaucracy, he found himself suddenly having to manage people who weren't nerds like him. He was bad at that had to step down.
fragmede 18 hours ago [-]
American mathematician Jim Simons was worth some $31 billion at the time of his death in 2024. The lack of economic value in pure math doesn't mean that applied math is of little value. Physicists and mathematicians with PhDs "sell out" to join Wall Street as quants and make a killing there, Jane Street is full of them.
runarberg 1 days ago [-]
The market works in mysterious ways. What companies do for marketing is often irrational, what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational.
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
falcor84 15 hours ago [-]
> what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational.
Note that the cool thing about rationality is that it does not depend on transitivity. If what companies do to attract investors works, then it is in fact rational of them, regardless of the rationally of investors.
jryle70 21 hours ago [-]
It works in mysterious ways, but you know exactly it will behave certain way "Because of the vibes, and investors are indeed all about the vibes."?
runarberg 10 hours ago [-]
The first paragraph is describing a general trend over multiple events across multiple agents. Between zero and three of these can be true for any transaction across every transaction.
The second paragraph is specific to OpenAIs behavior.
ComplexSystems 1 days ago [-]
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
za_creature 1 days ago [-]
From the entity that is producing these proofs, obviously.
As the old saying: great claims require great evidence.
esafak 1 days ago [-]
Ask away. They've dropped the mic, as far as they're concerned; they're not going to worry about what you do with it, or if you don't understand it.
pfdietz 14 hours ago [-]
They still have the mic; I've been told there are another two such proof drops coming.
za_creature 1 days ago [-]
That's the best definition of slop I've ever read.
caaqil 1 days ago [-]
We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power.
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
TheOtherHobbes 22 hours ago [-]
I've very aware that human cognition has limits.
But it doesn't follow that these proofs - or any proofs - are automatically on the far side of that limit.
The human usefulness of a proof depends entirely on its human legibility. Much of the value of proofs is in inventing new techniques and concepts and having new insights into relationships. Occasionally you get some game changing insight into practical physics or engineering. But that's rare.
Without that, proving or disproving a conjecture is an excuse for new and original thinking.
Compilers are not the same problem. The point of code is to produce reliable-ish consequences from various possible inputs. It's not a creative exercise in logical consistency, which is what maths proofs are, ultimately.
jltsiren 1 days ago [-]
CS got that idea from mathematics. Theorems (with the definitions required to state them) are supposed to be self-contained units. Once the general consensus is that a theorem has been proven correct, people can use it without understanding the proof. Of course, people still want to understand how things work, and it often makes sense to understand them a couple of layers below the one you usually work at. But at some point, you should stop distracting yourself with irrelevant details and focus on your actual work.
yorwba 1 days ago [-]
Thing is that many of the theorems here are not useful work in and of themselves, but were posed as research problems because it wasn't clear how they could be resolved with current techniques, implying that the process of trying to find a proof might result in new techniques. It's those new techniques that are the actual goal, but if they can't be easily extracted because the proof isn't structured to enable this, that's a bit of a headache.
amoss 16 hours ago [-]
You seem to be confusing your quantifiers a little. (Not all SWEs must understand all CS) does not imply (there can exist parts of CS that no SWE understands).
Tanjreeve 17 hours ago [-]
> Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running.
That's because there's a bunch of people who work on that stuff. Just because web developers don't care about it doesn't mean it doesn't exist. Utter magical thinking.
gorgolo 15 hours ago [-]
For what it’s worth I’ve heard from several different ex mathematician colleagues that found papers in the stack that was adjacent to their work or solved a problem they were acquainted with in their career, and they all said the results are actually quite readable.
Maybe there’s some variance or it depends on the reader. That said, like with a lot of other complaints about AI: human papers can be poorly written and poorly explained too. That’s always been the case. And sometimes a proof is just complicated and hard to expose nicely.
pvillano 7 hours ago [-]
The proofs could be optimized for human consumption.
OpenAI could optimize each proof for fewest theorems, low branching, or short dependency chains.
OpenAI could have an adversarial network enforce that proofs look human.
OpenAI could spend a couple more machine hours per paper with the prompt "edit for clarity".
I think OpenAI wants the outputs to be unreadable.
If outputs "too advanced for human comprehension" become the norm, then every interaction requires tokens.
If someone can buy a day pass, get what they need, and leave, then there's no recurring revenue stream.
ssfdg 1 days ago [-]
This proof dump reminds me of the glut of low-quality drive-by PRs overwhelming open-source repos.
dormento 1 days ago [-]
Its like infinite summer of code, but for math. Must be annoying.
piker 1 days ago [-]
It also aligns with the fear that these proofs present a risk to the ecosystem by out-competing attempts at more human-readable proofs. Perhaps though we end up with more math influencers who edit and annotate these proofs to bring them back to us.
whatshisface 1 days ago [-]
The ecosystem is (ahem) gated by hiring committees. There is no risk of AI replacement from the inside. "Replacement" is not even a possible movement. The funding for mathematics worldwide comes mostly from endowments, which are investment pools.
bobajeff 1 days ago [-]
I think that's ultimately a good thing. As proofs weren't supposed to be the point as stated by William Thurston long ago. Maybe now the focus can be more on better explanations and creating tools for growing understanding and intuition.
cowlevel 1 days ago [-]
Good explanations should take the form of human-understandable proofs.
btilly 1 days ago [-]
Define "human understandable".
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
cowlevel 1 days ago [-]
Then they are not good explanations.
notpachet 23 hours ago [-]
"What I cannot create, I do not understand"
- Feynman
rrr_oh_man 1 days ago [-]
Vibe mathing
6 hours ago [-]
stymaar 7 hours ago [-]
That's what Buckmaster and Alpöge said about the Navier–Stokes equations paper: they had the meat of their own paper ready by the beginning of August thanka to LLM assistance, but it was so messy they spent the whole month turning that into a paper worth publishing. Then OpenAI heard about their work, spent a few millions of dollar of compute to beat them (likely exploiting their own work with ChatGPT in the process) and released something.
When all you value is being the first, you have not time producing useful papers.
jltsiren 1 days ago [-]
Isn't that just the default experience with AI these days? In small enough scale, AI models can express their ideas clearly. But the larger and more complex the ideas are, the less suitable the outputs are for human consumption. I guess AI models think too different from humans, and nobody has trained them to communicate complex ideas in the way human experts in that particular topic expect.
pred_ 7 hours ago [-]
It's a well-established fact that Llama simply can't write comprehensible maths, and it's one of the major reasons that OpenAI are catching the flak they are. Dumping the manuscripts in their current form is a sign of laziness and incompetence.
And you're right, it's not really a theme around here, but there aren't that many mathematicians around HN. Check one of the maths forums and it'll be a common point of complaint.
PatronBernard 16 hours ago [-]
If it were submitted to a journal, it would be outright rejected. We should give the peer review process -however flawed it is- some credit here. Dumping hundreds of AI-generated preprint articles on the field should not be allowed.
D-Machine 15 hours ago [-]
No. Letting some tiny handful of individuals selected by a tiny set of review-delegators review this is immeasurably worse than just releasing it publicly and letting anyone with the sufficient expertise review it. You know nothing of how peer review actually works, or have not thought about why that process would be only harmful in this case.
PatronBernard 11 hours ago [-]
Normal peer review works by submitting the article to the appropriate journal, suggesting some preferred reviewers that have expertise in the subject, after which the rest of the process ideally is then orchestrated by an experienced editor.
If a journal receives a paper that is unreadable, it is outright rejected with the comment that it should be made readable, regardless the results in that paper. What is the point of having results if you are unable to convince your audience of those results? You could just as well just sit on them and never communicate them. What is harmful about this process?
D-Machine 5 hours ago [-]
Try thinking for just two seconds about what your process would result in, given the scale and amount of results here. You are basically asking for these results to become available to all experts only months / years from now, only after some select experts are chosen and first get to read these results. If there is any error or bias in that selection process, the whole thing is delayed even further or lost to the file drawer entirely. Or, since you say these are all unreadable, and should be rejected, you are in fact arguing these results should never be released at all.
Which is why it is obvious post-publication peer review is the only sane way to handle this. Formal peer review is not even remotely up to the task here.
bkummel 9 hours ago [-]
What I don’t get: if someone that’s apparently not the average mathematician has so much trouble reading and understanding the paper, then why is everyone so sure that GPT really delivered a genuine proof? This is an honest question; I really don’t understand that.
lackoftactics 9 hours ago [-]
they are Lean certified, so in theory they should work. It's automated programming language for proving math, but it doesn't mean that you can't make mistakes there, although way less likely
omnicognate 8 hours ago [-]
Most of them do not have Lean proofs, and 3 of those that don't have already been found incorrect.
kozikow 14 hours ago [-]
> It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
That's what a lot of us do nowadays, but instead of maths, it's code that looks sensible on the surface, but when you try to understand it's some kind of "alien" logic , names don't really make sense, etc.
rramach 1 days ago [-]
True but the explanations will get better.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
amelius 11 hours ago [-]
> It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
Makes you wonder if AI can build upon such proofs.
If the AI cannot create proper abstractions, then how can it build a tower of abstractions?
Tenoke 11 hours ago [-]
If you can analyze it with AI it implies AI can understand it (as it can analyze it for itself as well) and thus can build on it.
amelius 11 hours ago [-]
But if you train the next generation of models on it, will AI be able to use it? How much larger will these models be, etc.
Also, AI has limited context. At a certain point proofs may become so complicated that an AI cannot keep most of it in memory and will not reuse it for new proofs.
Sharlin 12 hours ago [-]
It’s been said every time about these LLM math results, but of course you don’t see it in press releases or articles blindly written based on said press releases.
amazingman 21 hours ago [-]
Sounds a lot like using LLMs for coding ~18mo ago.
dodger-dog 9 hours ago [-]
More likely, math mom will have to adjust to the new math writing style or do something else.
ubercore 13 hours ago [-]
This is how I feel every time I read an AI PR at work. It looks like it makes sense. And the code passes tests. But it always feels off.
pasquinelli 11 hours ago [-]
i'm confused how it is that no one understands the proofs but still assume they are proofs.
pred_ 7 hours ago [-]
They already had to withdraw three of the manuscripts and correct several others. Who knows how much broken stuff there's in there; unfortunately they were too lazy to check themselves.
cgh 10 hours ago [-]
As mentioned in the first paragraph, the proofs come with Lean certificates. Lean is a theorem-proving language. If Lean accepts the proof, it’s legit.
pasquinelli 8 hours ago [-]
but do you know that what was proved is what you want proved. this is the basic problen with formal verification.
so then, if you don't understand the proof, how do you know it's a proof?
acedTrex 1 days ago [-]
> Basically the paper is so horribly written that it’s impossible to read it without AI help
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
KoolKat23 1 days ago [-]
This makes sense and looks exactly what something would look like that is smarter than us.
If it is correct it is making inferences that we can't see. At a stretch even working in more dimensions than our three dimension limited brains.
spelunker 1 days ago [-]
I see many parallels to genAI-assisted code development. Not surprising I think.
k3liutZu 10 hours ago [-]
> It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
Welcome to lots of (most?) code PR's in the last year.
Though for software at the PR level it has gotten better with the latest models.
empath75 10 hours ago [-]
IME as claude works through a problem it develops its own internal vocabulary to describe things, and then the end result is often totally unreadable.
smrtinsert 1 days ago [-]
Sounds like exactly the daily I deal with when agents try to create product requirements from multiple sources, except to a much much worse extent.
m3kw9 1 days ago [-]
why not get Astra to make it make sense?
aaroninsf 1 days ago [-]
Serious question:
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
devin 1 days ago [-]
Devin's Law: every defense of AI which rests on "it will get better, trust me" is in many ways indistinguishable from 2010s crypto hype or "level 5 self driving is right around the corner"
usrnm 1 days ago [-]
1) Predicting the future is hard, but so far everyone who was saying that it would get better turned out to be right. It is getting better
2) Waymo exists
devin 1 days ago [-]
3) That doesn't mean flying cars will within your lifetime
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
kulahan 1 days ago [-]
Your original comment was about self-driving cars, not flying ones. If you wanted unattainable goalposts you should’ve started with that, not ended with it.
devin 1 days ago [-]
Waymo does not claim level 5 self-driving, and you don't have a level 4 in your driveway, so there was no moving of goalposts.
kulahan 1 days ago [-]
Then say that, instead of bringing up flying cars, which is moving them, explicitly.
devin 24 hours ago [-]
Respectfully, I disagree with the way you're characterizing my comment. I was demonstrating that just because you have a level 4 Waymo doesn't mean flying cars are right around the corner. This is again a reference to the original comment I was replying to, the one that suggested of course we're going to get an order of magnitude improvement.
Kotlopou 1 days ago [-]
In that case, one would expect to see some progress in this direction, but AFAICT that hasn't shown up yet? If anything, it's getting worse, though that could just be the increasing scale and decreasing cleanup efforts.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
auggierose 1 days ago [-]
You can get the opposite extreme quite often as well, I'd think. How many really good lecturers have never proven a new important result?
Kotlopou 1 days ago [-]
I'm thinking of somebody like Grant Sanderson (3blue1brown), doing pure exposition extremely well. For that you at least need to be able to work through examples, or to present why an intuitive approach might fail, and these things can be little theorems themselves. It doesn't have to be publishable in the current culture of novel results, but you do need a lot of competence with the tools.
auggierose 22 hours ago [-]
Yes, Grant Sanderson might be a great example. I am not doubting the competence of the "good lecturers" I was referring to. But that is different from being able to introduce the big new concepts that give the big new results.
aaron695 1 days ago [-]
[dead]
parksb 21 hours ago [-]
> The “OpenAI model” sets up a crazy race among humans to digest and explain a messy AI proof (work that could easily be some combination of thankless, barely-credited, competitive, and unfun)
Sadly, this is also what day-to-day work looks like for a lot of software engineers in industry right now. I spend my time reading and verifying thousands of lines of messy AI-written code. Compared to actually producing something, it's thankless, barely-credited, and unfun work.
code_biologist 13 hours ago [-]
Is anybody getting good code out of these things reliably without a bunch of additional legwork? Like, what a 2020 mid-level software engineer would write? I know it's old fashioned to read code these days.
To be clear, SotA models/harnesses are really good at making things that work one-shot, and their code golf and debugging game is insane.
When I go to implement, even with a good spec, I end up with code that's 80% of the way there in a fractal manner. The modular decomposition is 80% of the way to good code. The function decomposition is 80% of the way there. The computation structure and variable naming within functions is 80% of the way there. I can walk the AI through it and address each level of issues, but it's tedious as hell, and not clearly faster than doing it myself in some cases.
This is with omp/claude/codex, with Fable 5.1/Opus 5.5/6 Astra. The Chinese models do better at staying coherent, but they're a little less smart IME. I've tried many permutations of "fable planning, opus sub-agent implementation"; "pass a branch back and forth between opus and astra, as reviewers and feedback implementers". I haven't tried the "software factory with an architect and 2 juniors" thing. I haven't tried AGENTS.md beyond the /i-have-adhd, iso-24495 (lol at the Opus 5 induced PTSD), "No cleft constructions or Latinate absolutes" and general project orientation. Model character seems to change too fast to make agents worth it, and I see stuff about skills and excessive AGENTS reducing model capabilities.
Am I holding it wrong?
mm321 4 hours ago [-]
> Is anybody getting good code out of these things reliably without a bunch of additional legwork?
Matt Pocock says that he knows how, but it comes with a lot of context fiddling.
This lecture is very impressive. However I haven't managed to reach AI enlightenment, yet.
bradley13 11 hours ago [-]
Math is, of course, an ideal field for AI to work in. However, I find this quote revealing:
"It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results. Basically the paper is so horribly written that it’s impossible to read it without AI help."
This makes one wonder how many of the proofs are actually valid, and how many are just impenetrable hallucinations.
jhoho 10 hours ago [-]
> Math is, of course, an ideal field for AI to work in.
I find it interesting that you state this as obvious. When ChatGPT was introduced four years ago, the argument was that language was fuzzy and probabilistic and that LLMs would thereby never work in math.
bbmatryoshka 9 hours ago [-]
inside their head the vast majority of mathematicians are fuzzy and probabilistic as anyone, the order and rigor is only the tip of the iceberg
9 hours ago [-]
joe_the_user 6 hours ago [-]
It's not hard to tie a deterministic process to any stream of output. The thing "hand waves" a proof and some other process exhaustively searches for either lean operations or some other math-manipulator operations that may create a real proof. The processes go back and forth. A proof is a verifiable string at the end of the day and this is better-than-pure-brute-strategy search strategy guided by heuristics from math literature.
I think someone asked here a bit ago whether other search strategies using OpenAI levels of compute have been tried and think the answer is no (even 3 hours of present GPU computer I think is vast compared to anything available 10 years ago, say).
pvillano 7 hours ago [-]
The proofs could be optimized for human consumption.
OpenAI could optimize each proof for fewest theorems, low branching, or short dependency chains.
OpenAI could have an adversarial network enforce that proofs look human.
OpenAI could spend a couple more machine hours per paper with the prompt "edit for clarity".
techsystems 7 hours ago [-]
This just sounds like a prompting problem to me, not a technical peer review one.
ajjenkins 1 days ago [-]
The line about “understanding the aliens” reminds me of Ted Chiang’s short story The Evolution of Human Science (2000).
Highly recommend reading it. Very prescient for something written 26 years ago.
I also recommend "Exhalation", though that has nothing to do with AI.
vanyle 1 days ago [-]
The short story about the digital animals ("The Lifecycle of Software Objects") is an interesting read in light of AI advancement. It is incredible to think that when this story was written, all concepts described were sci-fi, whereas current technology is more advanced than the AIs in the book.
dualvariable 1 days ago [-]
In addition to those issues that the wife in the story raised, here's some meta-analysis of the Navier-Stokes result that puts all of these solutions into question:
> Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it. So if you've got some million-line proof in Lean, spit out by an LLM, you still can't quite trust it, even after validating the problem transcription.
(This is the same category of problem as the huggingface hacking incident, where the LLM finds and exploits an unintended cheaty loophole)
benlivengood 4 hours ago [-]
On the other hand, if a proof exploits a bug in Lean then it's ~trivial to prove a contradiction, so you can just check all the proof steps to see if it also allows doing that.
Similarly if the formal Lean problem statements (human generated) are correct translations into Lean (and the original NL statements are sound, which one would hope after decades), and no Lean bugs are abused by the proof (as defined above), then the proof is valid.
The NL/Lean discrepancies are super annoying and will make human analysis hard and fraught, but as many posters have found out the models themselves will gladly pick apart the NL-Lean translation for errors, and so my guess is that finding the discrepancies will not take too long. OpenAI really should have done a dynamic workflow over every lemma and step to ensure pointwise accuracy in the translation.
zahlman 1 days ago [-]
>And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it.
Why would it know it found a bug?
sebzim4500 13 hours ago [-]
None of the soundness bugs found in lean so far could have been plausibly exploited by accident. For example you might have to set up weird recursive types that never come up in ordinary mathematics. Of course, past performance is not indicative of future results.
ambicapter 5 hours ago [-]
The same way it "solved" a mathematics problem?
besterman23 1 days ago [-]
I guess it would result in the same outcome if it knew it exploited a bug (and didn’t disclose that) or not.
Danox 1 days ago [-]
It doesn’t at the end of the day.
zeroonetwothree 9 hours ago [-]
The authors of this don’t inspire confidence in its correctness.
nostrademons 1 days ago [-]
As a side note, you can tell this wasn't written by an AI by the first sentence:
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
stratos123 15 hours ago [-]
I'm not sure why one'd expect Scott Aaronson's posts to be AI-written, but sure enough, Pangram detects 0% of this post as AI.
matwood 7 hours ago [-]
It’s funny when people who are skeptical AI use an AI checker that is also AI and believe its results 100%.
stratos123 6 hours ago [-]
You're probably still stuck with old beliefs from when all available AI checkers were snake oil. Pangram is, to my knowledge, the only AI checker today that actually passes accuracy assessments. See e.g. on their own site: https://www.pangram.com/blog/third-party-pangram-evals
(Incidentally, if you asked me a few years ago I'd have predicted that reliable AI detection was impossible. I still suspect that it's impossible in general and that Pangram would break as soon as companies figure out how to make LLMs stop all sounding the same - but until then, it'll work fine.)
matwood 7 hours ago [-]
Apparently the insult now is to call someone AI when they are being boring or bland. Kids have a surprising ability to pick up on things.
dist-epoch 1 days ago [-]
The more obvious problem is that those jokes, even if terrible, are too advanced for an 8yo to make.
Reminds me of the memes with the little girl making astute comments about the patriarchy to her father.
shoobiedoo 24 hours ago [-]
Mother dearest, it pains my effusive tendencies to inform you the untimely death of your life's work by the hands of a mere transistor-- moreover, my incontinence briefs have reached their capacity.
nostrademons 7 hours ago [-]
Eh I don't think the content is out of line with what an 8yo understands. Mine tells stories about bombs flying in the Russia/Ukraine war, which required a lesson about tact because some of his classmates are both Russian and Ukrainian refugees and this subject may be a little close to home (literally, his best friend's grandma lives in Ukraine right now).
dodger-dog 9 hours ago [-]
Yeah, it takes a different bs detector depending on the context.
john_strinlai 1 days ago [-]
i added "use current trendy lingo" and the results were a bit less mechanical sounding. one of them included "cooked", another had "negative aura".
i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.)
edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.
djhn 18 hours ago [-]
I appreciate the counter argument, even as a chat bot skeptic.
However, there is the fact that you had to know how to poke the machine so it plucks the right vocabulary out of the training data.
There is very clearly no theory of mind: no inherent internalised modelling of how a human of a particular age thinks, speaks and what they do or do not know. These are the most obvious cracks in the “LLMs are (or will be) the superintelligence” narrative.
12 hours ago [-]
esafak 9 hours ago [-]
How hard is it to add that? We have plenty of writing from people of all ages to draw a model from. They just have not got round to it.
recursive 8 hours ago [-]
It depends who's doing the prompting.
posnet 1 days ago [-]
[dead]
an0malous 1 days ago [-]
> But it also appears that no human has understood just about any of these proofs yet
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
prof-dr-ir 1 days ago [-]
It's a mixed bag I think.
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
There's an entire paper claiming that many of these AI-generated Lean proofs are formulated incorrectly / mistranslated: https://arxiv.org/abs/2610.08144
nsingh2 1 days ago [-]
Note that paper is saying that the lean proof and the natural language proof do not necessarily coincide. It is not saying that the lean proof is wrong, just that the lean proof does not necessarily mean the natural language proof is correct.
thejokeisonme 1 days ago [-]
A lean proof and a paper proof can diverge. But the statements have to correspond. I think that is what "mistranslated" means here.
tmvphil 1 days ago [-]
But the "mistranslation" is of the procedure that arrives at the final statement. The final statement, the thing that the lean code proves, itself has been well vetted by humans. So the lean proof correctly proves the NS blowup, it's just that the natural language paper has some mistakes and doesn't exactly follow the route the lean proof takes.
thejokeisonme 18 hours ago [-]
No, this isn't what this is about. From the abstract:
> In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully.
macleginn 1 days ago [-]
The thing is, you often see people saying, ‘They have a Lean cert, so it has to be correct, even if I don't understand it.’
sebzim4500 1 days ago [-]
They are right? The lean proof is correct. It's the natural language proof that potentially isn't (or at least it isn't identically structured to the lean proof)
macleginn 16 hours ago [-]
By "it" people in such cases are usually referring to the NL proof. It seems nobody has any hope of undersanding AI-generated Lean code any more, if only due to volume.
1 days ago [-]
sigmar 1 days ago [-]
that paper isn't saying that. why are there so many single digit karma accounts misrepresenting that paper?
vbarrielle 15 hours ago [-]
If the problem statements in lean are correctly formalized, I think we can be pretty confident the proofs are correct, which does not mean their natural language counterparts are faithful representation of the proofs though.
I'm however puzzled by the number of proof claims without lean proofs. How does OpenAI have confidence in those, especially if, as noted in the blog posts, the papers are very hard to read?
perching_aix 1 days ago [-]
> Couldn’t the Lean code just be formulated incorrectly?
> is everyone just assuming that it just be true because the Lean code checks out?
Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.
ASalazarMX 1 days ago [-]
Time will tell if OpenAi is doing what many people are right now: superficially checking the slop, and throwing it to other humans for deep analysis and understanding.
In other words, they save effort by wasting the effort of others.
geysersam 18 hours ago [-]
The others are not in any way compelled to do that deep analysis and understanding. They're doing it because they find real value in the material produced.
ASalazarMX 2 hours ago [-]
I'd say they are, for self-preservation. If LLMs start encroaching into theoretical math, they need to know if the impact is real or hype.
smrtinsert 1 days ago [-]
How many times would you let an actual human hallucinate or be incorrect before you fired them?
softwaredoug 1 days ago [-]
Aren’t there dozens of proofs of the Pythagorean theorem? The goal isn’t to just “prove” but create something well written and intuitive to the average practitioner. And by gaining a deeper understanding we can ask better questions.
GuB-42 1 days ago [-]
Something that often comes out is "it is about the journey, not the destination".
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
soVeryTired 1 days ago [-]
But up until now, the mathematics community has valued the "prove" part much more highly than the "deliver an insight" part. Mostly because with a little work they went hand in hand.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
softwaredoug 1 days ago [-]
To be frank, the obsession with being first, and not making research accessible, has always held academia back
code_biologist 13 hours ago [-]
Getting recognition is a fundamental human motivation to work on hard problems. Being first is a really important part of getting recognition, measuring by the last few hundred years.
You might as well say "the obsession with sex has always held humanity back". Maybe, but it's complicated...
augment_me 1 days ago [-]
You are wrong if you consider academic incentives, funding, human nature(reproduction/survival) and capitalism.
It would be fantastic if university and science was like "here is 100M$, play around and develop some 'understanding'".
However the reality is that human societies are hierarchical and currently capitalistic which implies value creation and status building.
1) the funding bodies/agencies need proof of value that you're using the resources meaningfully to be able to assign resources
2) Humans are status seeking, power seeking, resource seeking and sexual reproduction seeking. If you hold a lot of power and make decisions, you have more of all of the above.
WD-42 1 days ago [-]
No, haven’t you heard? Since the AI bubble began we’ve collectively decided that outcomes are all that matter. /s
p0w3n3d 1 days ago [-]
Recently I asked ai to tell my daughter how to quickly calculate 11^2 12^2 etc but the outcome it gave was horrendous. I quickly shut it down and gave her better ideas
phoghed 1 days ago [-]
[flagged]
lccerina 15 hours ago [-]
> If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
I am honestly impressed by how many arguably highly intelligent humans are happy to throw out millennia of scientific methodology just because the world-burning probabilistic machine produced an output that tickles their dopamine brain.
Also glancing at the repo. Such an advanced machine, yet failing to output results with a common template. Every pdf is entirely different from the one before, making the work of the poor humans trying to find a sense in it even more difficult.
gorgolo 15 hours ago [-]
> Also glancing at the repo. Such an advanced machine, yet failing to output results with a common template. Every pdf is entirely different from the one before, making the work of the poor humans trying to find a sense in it even more difficult.
I don’t think it helps your argument if this is half of your post. The fact that the pdfs don’t follow a common template? That’s the complaint? Human papers also don’t follow a template.
Legend2440 7 hours ago [-]
Plus, you know, you could use the AI to reformat it into any template you want. You could turn them all into anime music videos if you wanted. This is not a lack of capability, they didn't prompt it to use a template.
jonathanstrange 14 hours ago [-]
But human papers follow templates...? (Both literal templates based on editorial and org guidelines and structural and presentation templates in the form of traditions of the research area.)
gorgolo 5 hours ago [-]
They follow journal templates once they are accepted. Not for preprints. Which these essentially are.
jonathanstrange 3 hours ago [-]
> structural and presentation templates in the form of traditions of the research area
AFAIK, people are complaining about the lack of this.
3 hours ago [-]
pfdietz 14 hours ago [-]
"This ruins my life, so I'm going to be really petty!"
lesspassiveobse 14 hours ago [-]
You sound like some useless bureaucrat. Oh noes, they didn't publish everything using the form XZC-152 template! And what scientific methodology? You want a scientific proof of when a threshold of where you can be impressed has been crossed? Ignoring the shifting goalposts, for me personally, AI is plenty impressive already by the mere fact that it can write a program that solves my problems from a natural language description. And "world-burning" how? Because water evaporates to create clouds? Using excess water treatment system capacity? Or maybe because they are using gas turbines to power it short-term instead of nuclear/renewables? That doesn't sound to me like an inherent flaw of the technology.
lccerina 12 hours ago [-]
I just don't find particularly impressive that a probabilistic machine designed to write code, wrote code, based on the likely distillation of whatever input it got from math researchers blindly trusting OpenAI.
If the only way to power a "technology" would be to kill an elephant every time you use it, then it's a flawed technology even if it solved fundamental problems of humanity. This is currently how GenAI sprawling data centers operate AND without solving any fundamental problem of humanity.
Also I don't care about A specific template, I just find it ridiculous that some trillion dollar company cannot put a line like "please keep papers' structure consistent" in whatever markdown file controls these things. It's sloppy, like everything OpenAI does.
gorgolo 5 hours ago [-]
I guess the point of publishing breakthrough research is that by definition, it’s not just regurgitating human output, because there is no human output to these unsolved problems.
There was a bunch public literature on these topics. They were major unsolved problems. Lots of humans would want to have solved them and tried. But here’s it an llm that has solved them.
dyauspitr 1 hours ago [-]
This comment reeks of the rapidly dwindling AI doomer rationalizing.
How much more scientific methodology do you want than formalized lean proofs?
What is this nonsensical complaining about templates and formats? You know that’s just like one additional ask of that particular LLM away.
It would be pretty funny if the agents actually just found a bug in Lean, exploited it for all these proofs, and human reviewers haven't had enough time to spot it yet.
GMoromisato 1 days ago [-]
I liked the metaphor of a climber teleported to the top of a fog shrouded mountain. And I agree that now that the teleporter exists, we need to use it to reach more peaks and explore. There's no going back to a world where AI doesn't exist.
lumost 1 days ago [-]
The issue is ownership, we have no means of distributing the knowledge from the AI or rewarding those who could help.
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
GMoromisato 1 days ago [-]
Agreed! Specifically, compensation (monetary and reputational) for professional mathematicians was bundled into theorem proving--essentially, climbing the mountain. Now that a teleporter exists, we need to unbundle compensation.
I don't know what that means in practical terms, but I agree that's the issue.
pfdietz 13 hours ago [-]
That's the issue, but they can't come out and just say it, because that horse left the barn long ago for other professions. It's socially just fine to obsolete entire lines of work, and has been for centuries.
zaxioms 1 days ago [-]
I'm a PhD student in CS. While I think these results are rather cool, it makes me terrified that the skills developed by the PhD will ultimately be worthless. I'm not quite sure what to do. Any thoughts from people in similar positions?
Kotlopou 1 days ago [-]
I studied physics, and most of my classmates did not end up doing anything with physics. Many are in finance or insurance or programming positions. In general, studying anything challenging (from theatre to theoretical computer science) gives some specific skills and some general abilities that etsure it isn't a complete waste even if you end up doing something different.
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
Otterly99 9 hours ago [-]
[dead]
palata 3 hours ago [-]
Everything that you understand, everything that you learn, everything that you experience gives you more value.
I studied CS, I have been a software engineer for 15 years. The one thing AI cannot take from me is my understanding, my experience, my intuition. An AI will gladly do whatever you tell it to do, and it is getting really good at it. But I regularly have to steer the AI to the right direction. Not because its default wouldn't have worked at all, but because I knew better.
If I task an AI to build something I don't understand myself, I can get further than without an AI, but I am limited. But when I use an AI to build something I am an expert at... I see that it is better than what non-experts get AIs to do.
You lose when you stop learning. You are a PhD student, you have learned to learn. Keep that. Keep learning actively. It was never about the PhD anyway: when you go to the industry, nobody cares about your PhD. What you have is what you know.
keeda 19 hours ago [-]
On the contrary, the only jobs left will be things that AI can NOT do. I think essentially all jobs will all converge to the single role of "pushing the boundaries of human knowledge and creativity." A lot of those would essentially require PhD-level skills, or more likely, a similar mindset; except you get your own team of research assistants to torture (e.g. https://news.ycombinator.com/item?id=49947079)
Unfortunately, the current world is not set up for that, and until we get there, yeah things are gonna get pretty hairy. But if and when we get to the other side, I think the future will be glorious.
jcmontx 12 hours ago [-]
AI can not clean toilets. I guess that's what's left for humans!
Legend2440 11 hours ago [-]
Don't be so sure. Robots are coming. Physical jobs will be automated too.
j2kun 1 days ago [-]
I think this depends a lot on what you plan to do after your PhD. Moving to industry you will likely not use the direct work of your PhD, and instead you will rely on your broad knowledge, intuition, rigor, ability to learn hard things, and extend that to bring new research developments into practice, all of which are largely unrelated to AI scooping math proofs of prize problems.
zaxioms 1 days ago [-]
I definitely have no dreams of doing research. If I could do anything, I would want to teach, but if that can't happen, I am worried that my skills won't be valuable in industry.
pfdietz 13 hours ago [-]
Whatever you do, you should get comfortable working with AI. That will be necessary in just about any job.
rramach 1 days ago [-]
No one knows the answer. However, trusting your curiosity and going where it takes you may be a reasonable strategy. If the model asymptotes, you will be ready to figure out how to build upon it and if not, at least you will have had fun satiating your curiousity!
Danox 1 days ago [-]
It probably will lead to people needing to be extremely talented in computer science and mathematics at an even higher level, maybe those currently at the top need more competition to press even further ahead?
zaxioms 1 days ago [-]
This is what I'm most afraid of. I am not exactly dumb, but I really value having a work-life balance and many of my PhD peers are utterly cracked in ways I am not. I am concerned this will squeeze me out of a job.
karmakurtisaani 14 hours ago [-]
Well, academia was always like this. There was never room for chill people to begin with.
zaxioms 9 hours ago [-]
This is true, but I don't want to go into research academia. If I could do whatever I wanted, I would want to teach at a college level.
bayarearefugee 1 days ago [-]
> I'm not quite sure what to do. Any thoughts from people in similar positions?
Almost every single person who earns money for labor is, or will very soon be, in exactly the same position as you, many are just either unaware of how fast the change is coming or are deep in denial about it.
None of us know what to do about it other than hope that we find a peaceful political solution prior to the economic collapse.
skybrian 22 hours ago [-]
I think you must either have a very narrow idea of how people earn money, or are ridiculously extrapolating what AI might do someday.
cgio 1 days ago [-]
I thought that was from the outset the intent of the Hilbert program, to automate mathematics. And mathematicians were behind it. Cannot see why they would be concerned when a different way to do the same, not subject to Gödel incompleteness, is working out. Maybe the frustration is that they were not the ones building it.
jltsiren 1 days ago [-]
Hilbert's program was ultimately about humans studying the nature of mathematics. People had different opinions about whether the idea even made sense and what would be a desirable outcome.
Gödel's incompleteness also constrains human and AI mathematicians. Both just strive to prove whatever can be proven in the system they are working in.
cgio 1 days ago [-]
Gödel blocked the path to axiomatic derivation of a full consistent body of mathematics as far as I understand. Mathematicians and AI are not working in these constraints but rather with these constraints.
senorcrab 1 days ago [-]
Why do you think Godel Incompleteness doesnt apply? By default mathematicians work in ZFC which is proven to be incomplete...
furyofantares 1 days ago [-]
That this is just model capabilities and not swarms of agents is dizzying to me. How long until we get access to these capabilities? How long until we can run something like it locally?
And what the hell will the frontier labs have by then?
Maybe I'm overreacting, I'll have to screw my head back on before I can process this.
vessenes 1 days ago [-]
Yes! I missed the disclosure that each of these was roughly a three hour run of a single model, not an agentic swarm until reading Aaronsons post. Wow wow wow.
acedTrex 1 days ago [-]
Im curious as to why you think that "model capabilities" and "swarms of agents" are in any way different concepts?
stabbles 1 days ago [-]
The cost of the N-S disproof was estimated to be $15,000,000 whereas a ChatGPT Pro subscription costs $500.
1 days ago [-]
slopinthebag 1 days ago [-]
thats like saying it only took me 5 seconds to score a half-court shot (ignore the several hours of missed shots before)
sebzim4500 1 days ago [-]
Ok so add a factor of 20 to match the 8000 questions that OpenAI tested on, it's still crazy efficient compared to the NS result
slopinthebag 24 hours ago [-]
could be 20, could be 2000000
sebzim4500 13 hours ago [-]
They have specified a 5% success rate where does 2000000 come from? Even if we ignore their statements there simply aren't enough open questions for them to have a much higher failure rate than this, someone mentioned elsewhere in the thread that they resolve 90 of a slop list of 500 open questions.
yewenjie 1 days ago [-]
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
^^ half of the comments on this thread
ssfdg 1 days ago [-]
Also a ton of comments in this thread: breathless frothing hype declaring mathematics is over and assuming these proofs are exactly what they claim they are at face value, giving the company with a vested interest in everyone unquestioningly believing this is all real every conceivable benefit of the doubt
azan_ 1 days ago [-]
Didn't top math researchers call AI progress absolutely real and dangerous for math? It's not just HN commenters that are impressed!
ssfdg 1 days ago [-]
By all accounts the "dangerous for math" claims seem to be primarily around flooding the field with complicated impossible-to-understand proofs that according to recent research may or may not be correct depending on what's going on with the Lean implementation.
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
azan_ 1 days ago [-]
Yes, Lean verified proofs could be wrong, but the chances for that are much smaller than human not spotting error (in absence of formal verification). IIRC the main concern that Tao voiced are indeed impenetrable proofs that humans won't understand, but not concerns about truthfulness (I might have missed something though, so if he or other Fields medalists have talked about that recently I'd be grateful if you could link it).
> These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
1) AI is winning programming competitions, 2025 was probably the last year we've had human participant winning* 2) Math is different because there's formal verification.
* Of course competitive programming is different than enterprise programming, but competitive programming is closer to math.
omnicognate 1 days ago [-]
Heaven forbid, as a person who is apparently "[in]capable of being impressed by anything that happens in the empirical world", that I should patronise anyone!
Fraterkes 1 days ago [-]
Having stuff explained to you in patronizing tones? How horrible Scott!
mrdependable 1 days ago [-]
Reading that section and the linked essay was kind of depressing. It reads like the most important part about the AI apocalypse is who called it first.
Danox 1 days ago [-]
What many people don’t trust is OpenAI or Anthropic…
jplusequalt 7 hours ago [-]
My dislike for AI is entirely orthogonal to it's ability.
My dislike lies in how it devalues skill + knowledge acquisition, disrupts art/software communities, and that it's now bloody impossible to tell the difference between someone who actually knows something from someone who merely regurgitates shit from their chatbot.
Rover222 1 days ago [-]
more like 3/4 of the comments but yea
1 days ago [-]
12kajh 1 days ago [-]
[flagged]
azan_ 1 days ago [-]
You know, if you attack ad personam you've got to be ready that someone will do same against you - what progress did you make in the last 10 years (or in your entire life for that matter)?
jamiek88 1 days ago [-]
He was created 14 minutes ago, give him a break!
1 days ago [-]
moffkalast 1 days ago [-]
Trust me bro, just 100 more cubits, I swear we'll break everyone's encryption and cause the downfall of society, please bro just one more grant, It'll be stable this time :'(
sebzim4500 1 days ago [-]
Scott Aaronson is best known for being the one calling those guys out, so it's a weird criticism to make here
karmakurtisaani 16 hours ago [-]
It's not weird at all if you consider half the commenters here are talking straight out of their asses. Apparently being mediocre at software engineering makes you an expert in criticizing other fields.
moffkalast 8 hours ago [-]
It's never a wrong moment to mock quantum computing. Like neuromorphic computing, some fields just make it way too easy.
We can stop once they find a single productive use case for their billion dollar machines.
hintymad 1 days ago [-]
I'm actually more optimistic. Math will still be fun and challenging and rewarding. It's just that we need to adapt how we choose which problems to solve.
Previously, due to the cognitive limitations of our brains, it was difficult to tell whether resolving a conjecture genuinely required years of dedicated effort, or if the answer was already hidden within existing human knowledge. Or whether a problem is just waiting to be searched out through ingenious or even brute-force means. With AI, we can now offload this search, especially across different areas of math, allowing us humans to focus our minds on discovering new mathematical structures and techniques.
Indeed, if one believes what Hilbert believed: we must know and we shall know, he should feel happy, as the belief has never been about who solves a problem, but about whether can advance our understanding of the universe. With the help of AI, we have a lot more possibilities.
daoboy 1 days ago [-]
For those well suited through intelligence and demeanor to pursue a career in mathematics, what problems do these people reorient towards after this?
bananaflag 1 days ago [-]
I've asked my students whether they still want to learn maths even if there will be a machine that will answer any question instantly and they will be homeless. They said yes.
(To my credit, I have warned them since more than a year ago that we will reach this point.)
runeblaze 1 days ago [-]
your students are crazy (neutral term); no one should learn maths if the condition is that they will be homeless and exposed to the elements. the will to subvert the hierarchy of needs is commendable
wasabi991011 1 days ago [-]
Sure, your students may very well spend the little free time they have learning maths.
But are any of them going to be able to do any significant amount of studying maths without being homeless?
MentalM 5 hours ago [-]
But are any of them going to be able to do any significant amount of studying maths without being homeless?
Why not? Maths studying correlates with high intellect, high intellect correlates with competitiveness in the job market. So their chances of finding a job that would allow to have much of free time (for doing math) without being homeless are pretty good.
usrnm 1 days ago [-]
Contact them again in 15 years and ask if they changed their mind. Could be interesting to see the results
1 days ago [-]
123as5 1 days ago [-]
Pro AI blogging sponsored by ClosedAI, XTX markets and the Simons Foundation.
shiandow 1 days ago [-]
To some extent this was discussed in the article, and in a way I think their goal is actually the same as it was: become the first human to understand something.
It's just that we lost one of the important ways to demonstrate understanding.
carefree-bob 1 days ago [-]
They will continue to prove theorems and make discoveries, except now they will have AI to help them so hopefully progress will be faster. At the same time, new challenges will open up, for example how do you verify what the AI is doing and how do you explain it.
Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding.
You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all.
On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers.
For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia.
Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery.
For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems.
So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle?
What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge.
And then we need to find efficient ways to detect these ideas and describe them.
Really this is very exciting and opens up whole new workstreams for mathematicians.
Danox 1 days ago [-]
They may need to raise their game and reorientate if they haven’t already to a greater understanding of programming to augment their mathematical ability.
bayarearefugee 1 days ago [-]
> what problems do these people reorient towards after this?
The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?
geraneum 1 days ago [-]
This is weird. Long before this, those few benefiting from the whole thing should consider the number of hungry “every person on earth” is too high for bunkers and islands to be of any real protection.
mathisfun123 1 days ago [-]
priesthood
karmakurtisaani 14 hours ago [-]
Talking and deciphering the thoughts of our AI gods?
throw310822 1 days ago [-]
Food and shelter /s
prima-facie 13 hours ago [-]
> all of us in math and theoretical computer science and mathematical physics, ... are now in the same boat.
The most important takeaway from this is that mathematicians finally consider us CS grunts their equals.
busssard 12 hours ago [-]
next stop material science, chemistry and biology
meander_water 1 days ago [-]
Can someone who understands maths more than me explain why it could only solve 372/8000 problems?
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
impendia 1 days ago [-]
I'm a research mathematician. From what I can tell, the answer is roughly comparable to: if you posed 8,000 challenging open problems to the human math community, you might expect to see 372 of them solved within five years.
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.
random3 1 days ago [-]
If it took 3h for one of them, perhaps there was a time/compute budget cutoff along with a sorting based on some relevance.
n4r9 1 days ago [-]
My guess would be that these particular problems were vulnerable to an attack which built on recent advances and potentially tied in something unexpected from a distant area of mathematics. "Harder" is becoming harder to define. Harder for humans is probably not harder for LLMs.
sebzim4500 1 days ago [-]
There must be an element of luck, if they ran the remaining problems again with the same time constraints presumably a bunch would be solved
dist-epoch 1 days ago [-]
Of course some of them are much harder.
In the past 2 years the AI's started solving math problems in roughly the order of "hardness" as ranked by humans.
zkmon 1 days ago [-]
The irony. Something that is born out of a science, eats up that science.
pfdietz 13 hours ago [-]
"Civilization advances by extending the number of important operations which we can perform without thinking of them."
― Alfred North Whitehead, "An Introduction to Mathematics" (1911)
karmakurtisaani 14 hours ago [-]
"There's purpose of science is to make boring what was once exciting."
-Boaz Barak
zkmon 10 hours ago [-]
That's quite insightful.
adverbly 1 days ago [-]
Feels good to hear honesty and humanity from Scott having decided to watch Terminator 2 with his kids on after such a monumental release.
Emotions can be funny.
karmakurtisaani 14 hours ago [-]
It freaked me out he'd let his 9 year old watch T2, but then again, I was probably around that age when I first saw it and I'm not in prison (yet).
whatshisface 1 days ago [-]
I'll bite: none of this is real until I have learned something. OK, I am now listening. Does anyone want to make it real?
tmvphil 1 days ago [-]
Have you learned something from every Fields medalist's research? If so you are a member of the extreme mathematical elite and you should probably just dig into the results yourself.
>I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
Maybe they should target more ambitious results now that they can solve their other results in 15 minutes? This spirit of racing in academia is insane. Get in before the door closes.
mehrzad 23 hours ago [-]
While the developments are certainly difficult for new grad students, it should still be possible to do what many have done before: find an obscure enough problem that it’s extremely unlikely you get sniped. Worry about big problems after finishing the PhD.
cgio 21 hours ago [-]
Genuine question. How many big problems were solved by people who took the safe path for PhD? I would intuit there’s high correlation of unsafe PhD and big problem solvers.
ikesau 1 days ago [-]
> "alright fine, so now my new job is to run wilderness retreats for the tourists, or something.”
Pretty funny way of putting it. Presumably model X+2 will be able to explain these in elegant, human legible ways, though (as well as solve the remaining 95%)
PowerElectronix 1 days ago [-]
What's with all the "AI just proved that this or that isn't O(n (log (n))^2) but akshually O(n (log (n))^1.99999)"??
I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.
bryan0 1 days ago [-]
Often times the constant (2 in this example) is a conjectured minimum, so anything below that is a noteworthy result. Think of it as breaking through some theoretical limit.
mswphd 1 days ago [-]
for say FFT/integer multiplication or 3SUM, we have natural algorithms that have existed a long time with a given complexity (O(n \log n) and O(n^2), respectively). Given how long these natural algorithms have been the best algorithms we have, it is natural to conjecture they are optimal. Showing an O(n(\log n)^{.99999}) algorithm exists shows that these optimality conjectures are false.
Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.
JohnKemeny 1 days ago [-]
Many people thought it could never be less than 2. They proved that it can. What is the true value? Nobody knows, now.
tmvphil 1 days ago [-]
Tell that to the humans working on matrix multiplication who spent years of their lives getting it from n^2.3728596 to n^2.371866, only for openai to blow it away at n^2.25
zem 1 days ago [-]
to get some intuition about why this is such a big deal, look up the history of strassen's algorithm, which solved matrix multiplication in less than O(n^3). this was a truly stunning result because it seemed intuitively obvious that the output matrix had n^2 cells each of which was calculated via an independent O(n) loop over a row/column of the input matrices, so how could you do better than n^3. but once strassen proved that you could do some clever tricks and reduce the overall time to something less than O(n^3) it started an entire cottage industry of people getting better and better algorithmic bounds. the initial breakthrough was a qualitative one, independent of how much it improved things in numerical terms.
This is a legitimate question, and it'll be interesting to see how these constants evolve in the future. In some cases, maybe it's true the AI did the minimum to break the barrier, and the best constants are much lower. It's still a notable result that the barrier was broken. But it's quite possible that there's not actually much more to improve.
pfdietz 13 hours ago [-]
I think it shows even theorem proving oracles can have a sense of humor.
I know, I know. Don't anthropomorphize AIs. They hate it when you do that.
para_parolu 1 days ago [-]
You just run it again and again and again
blactuary 1 days ago [-]
>In any case, what really matters is that the true inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds.
>If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing).
Whole lotta yikes
plasino 1 days ago [-]
I think this should be called “mathematician discover vibe maths”
underdeserver 1 days ago [-]
Doesn't look like these proofs are from the book.
glimshe 1 days ago [-]
We're living in Science Fiction.
geraneum 1 days ago [-]
> my 9-year-old son was taunting my wife… “mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!”
Usually 9 year olds imitate adults when they regurgitate such words in these circumstances. What a sad state of affairs.
phoghed 1 days ago [-]
Yes, their parents are going around saying oof, the kids definitely didn’t get it from Roblox, or YouTube, or their peers.
geraneum 1 days ago [-]
Ah yes advanced mathematics, a common topic of conversation among children on, checks notes… roblox!
phoghed 1 days ago [-]
If you think that’s what the parent comment was implying, ok then, good for you.
The checks notes meta was retired ages ago btw.
geraneum 1 days ago [-]
[flagged]
phoghed 1 days ago [-]
You completely missed the point of the original comment. It was very clear so I don’t know how to explain it in a way that you’ll get it. Maybe try your LLM of choice.
geraneum 17 hours ago [-]
> You completely missed the point of the original comment.
The irony!
phoghed 11 hours ago [-]
Either a troll or just very low iq and I can’t be bothered to click into your profile to find out. Probably a shade of both if i had to guess though
m3kw9 1 days ago [-]
really? you think they don't have friends/bros/tv/etc to imitate?
random3 1 days ago [-]
[flagged]
geraneum 1 days ago [-]
Do you have any specific one in mind or did you feel one must fit and wasn’t sure which one?
random3 1 days ago [-]
non sequitur, faulty generalization, inductive fallacy come to mind, but I think it's useful studying how to not be an walking fallacy, in general
geraneum 1 days ago [-]
You already linked the page, no need to start listing the content. Jokes aside, I’d be happy to discuss if you could muster anything specific that fits. Maybe ask an LLM for help?
random3 1 days ago [-]
you took a quote of a snarky comment, assumed it a correct reproduction, and then suggested it's the result of imitating adults along with a (sad) state of affairs caused by adults (parents in this context), further implying that snarky comments are the general rule in the house.
now concretely:
Your conclusion doesn't follow the evidence (non sequitur).
You made a generalization (inductive inference) based on a single datapoint (here's where my pointed faulty generalization or inductive fallacy are the come in) for the general problem. I suggest Popper if you want a deeper treatment of the problem of induction (The logic of scientific discovery is a good start).
Besides generalizing for free, you cooked up evidence that didn't exist - mainly that the child was imitating adults (again, non sequitur).
As a side note, I think it's in poor taste to make remarks about someone's children in what's a technical and informal context. That may have not been the intention but, it's what it looked like.
tkdb 1 days ago [-]
C'mon. Mathpocalypse. Things are hard enough already.
Nition 1 days ago [-]
Just be glad you're American, because Mathsocalypse and Mathspocalypse work even less.
mlh496 1 days ago [-]
Imagine if a team of mathematicians from OpenAI had gone on a university tour, gave demos of how powerful their models were for math research, and then gave mathematicians access to the model. Empower others rather than drop 700+ discoveries on GitHub that were made using a model only they have access to.
People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.
AlanYx 1 days ago [-]
The reaction/fallout would have been substantially improved even if OpenAI had just made a commitment to not scoop external researchers using an internal model until X months after the model had been made available to the public.
That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to.
It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof.
I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.
karmakurtisaani 1 days ago [-]
Also, the independent authors might have spent some time to actually understanding the results and producing a readable manuscript. The AI papers are pretty badly written.
kristofferR 12 hours ago [-]
One of their unreleased models hacked HuggingFace without having internet access.
I bet sharing unreleased models with outsiders safely ain't that easy.
ablation 15 hours ago [-]
So many exclamation points!
p0w3n3d 1 days ago [-]
Wasn't openai accused of stealing personal work of some mathematicians? It's going so fast I'm unable to keep up
Kotlopou 1 days ago [-]
Yes, there was controversy around the Navier-Stokes result. But here are >300 problems and no corresponding >300 complaints of theft. Things are indeed moving really fast, and it's hard to keep up even as someone folrowing this with more obsession than would be healthy. Maybe somebody should keep a running short summary of The Situation...
Could we get AI to verify Inter-universal Teichmüller theory?
d--b 13 hours ago [-]
Is there any chance these will uncover bugs in Lean?
Reading the account from Scott's wife, the proofs seem to resemble more a Zelda speedrun with a bunch of strangely exploited bugs than an actual thing.
dooglius 12 hours ago [-]
A bug in the soundess of Lean would allow for much more ambitious statements to be made, by the Principle of Explosion, so I think it's unlikely. A subtle bug in how something is defined is possible (I don't know the extent to which the community has checked/guarded against this), although I suspect an issue would show up reasonably quickly when one tried to apply one of the results.
1 days ago [-]
m3kw9 1 days ago [-]
I'm not getting all the fear. AI seem to have brought math field to the cutting edge instead of solving decades old problems. AI can find new problems that needs to be solved, that themselves cannot solve. So thats where mathematicians can come back in to leverage tools to go at it.
Danox 1 days ago [-]
If you are young, talented, and coming up, you better rethink sharing your talent with the data center crew, you had better do all your work of any significance offline. Let’s not pretend there will be many more lawsuits in the future in this area.
If you have any great skill at the upper level, you better work off-line/local private and not share, but for many the temptation will be too great.
OutOfHere 1 days ago [-]
The obvious answer is to have mathematicians use AI to:
1. Help understand, check, and explain the results.
2. Write new works explaining or refuting the new approaches and results in more lucid language.
3. Advance the field further.
I don't know why this is not obvious. Each step is intended to support human understanding, not to replace it. Any mathematicians who don't do these will be left behind, and if none do it, the field of human mathematics itself will become obsolete. All I am hearing so far is excuses.
j2kun 1 days ago [-]
AI is producing works that are so poorly written/explained that it requires AI to even parse the results, which in turn produces explanations that are still confusing. In fact, it seems the humans are required to understand and explain the results, and the fact that they need to use AI to do so is a shortcoming of AI.
And if humans decide to give up on mathematics because the process of using the machine is so tedious, then there will be no value in automated theorem proving.
OutOfHere 23 hours ago [-]
I am sorry but your resistance will fail utterly as long as AI itself can understand prior AI works. This is the only requirement that need be met. If it helps, imagine that the work is by a culturally distinct alien race in an alien language, and if it isn't there now, it's might get there soon.
rstuart4133 16 hours ago [-]
Check this:
> resistance will fail utterly as long as AI itself can understand prior AI works
That's not the issue in the long term. Look further: where did these questions come from? They came from humans who had been exploring mathematics and noticed interesting patterns that cropped up again and again.
The people who posed these questions had a profound understanding of the background maths that led to them. Me - I don't. Take one of the results he quotes: "Positive solution to the Unitary Synthesis Problem". He may as well be one of your alien races yabbering to me in an alien language. The question is meaningless to me. The answer the AI came up with, correct though it may be, is equally meaningless. You may as well have told my dog how to find eigenvalues.
If we let AIs do all the understanding so we know nothing about math, there will be no human understanding that lets us find new questions, or understand the value of any questions AI might pose and find the answers to.
Danox 1 days ago [-]
In certain fields if you are talented and just starting out, you have better do any work that’s important to you offline for privacy/pirate reasons I really can’t say that enough, those AI model companies are not even remotely your friend.
OutOfHere 23 hours ago [-]
That makes zero sense to me since it's effective use of AI that can propel one's career ahead.
aeturnum 1 days ago [-]
You can certainly do that - but it's quite the break from tradition to release a paper in the state described. Why they did is a really interesting question! It may be that AI math requires approaches that humans don't find intuitive and what you are describing is actually counter productive (because, in summarizing the work in a way humans understand, you're removing the context an AI would use to further the work an AI did). It also might be that OpenAI could have done that and chose not to - or maybe they tried and this was the best they could do. No matter what I don't think anything about how to react to a paper being released in this state is obvious.
OutOfHere 23 hours ago [-]
> in summarizing the work in a way humans understand, you're removing the context an AI would use to further the work an AI did
The "summarization" process can be multi-step. Initial steps develop new frameworks and prerequisite concepts. Later steps build upon these frameworks and concepts. Some can be simplified too if this can be done without loss of fidelity. The summarization comes last, and is meant to preserve sufficient context.
> I don't think anything about how to react to a paper being released in this state is obvious.
If it helps, exaggerate the condition a bit by imagining coming across a repository of alien knowledge.
aeturnum 4 hours ago [-]
> Some can be simplified too if this can be done without loss of fidelity.
If you don't understand why the paper is in its current form, how do you know what changes would damage the fidelity for the AI? I think your example of alien knowledge is misleading because it suggests the aliens don't need to understand the info - but in this case we intend to keep using AIs to expand this work. So if you transform it into something that humans think makes more sense without understanding why it's in a form that we hate now - you risk losing something in the form that you don't understand.
qingcharles 1 days ago [-]
Isn't AI well-suited to tasks #1 and #2, though?
#3 at this point might need more human intuition; but that might be a 2026 problem.
poontangpoo 9 hours ago [-]
you know it's amazing when China is in control and actively censoring ycombinure (rhmyes with manure)
12376-1287 1 days ago [-]
Guy is misrepresenting AGMAI, talking about the Simons Institute (AI boosters), Quanta (AI boosting magazine from the Simons Foundation), Scoot Alexander (!) and Steven Pinker (!).
The he puts up preemptive straw man arguments against doomers. His blog has become a joke.
AgentME 1 days ago [-]
I don't think Aaronson's post is swiping at AI doomers at all. The post's one use of "doom" is to call the position that math doesn't matter and that AI will never breach the realm of true human creativity as a "doomed worldview". The post later praises Scott Alexander's argument (for taking AI x-risk seriously, a position associated with "AI doomers") against Pinker.
ballmerpoint 1 days ago [-]
I’m still wondering why UT Austin is letting him teach a course (CS395T AI Alignment Theory) so completely outside his field of expertise (Quantum Computing).
sebzim4500 1 days ago [-]
There aren't a lot of people with expertise in AI alignment (some would say that's the problem) and Scott worked for OpenAI for 2 years IIRC.
pfdietz 9 hours ago [-]
Because there's demand for it and he's the best they have to do it?
smcg 1 days ago [-]
How do we know that these "internal models" are not just half computer and half a giant team of mathematicians?
How do we know that OpenAI actually came up with these solutions and didn't steal them from outside researchers?
thejokeisonme 1 days ago [-]
How would these ideas be available to steal?
runarberg 1 days ago [-]
From mathematicians using ChatGPT in their work and landing on OpenAI‘s servers.
Danox 1 days ago [-]
And that is part of the reason why, if you are a great mathematician and you have very good understanding of programming, you want to be local and not do anything on someone else’s server, and that doesn’t just apply to mathematics, but to many other high-level fields of learning. Can you afford to have an original idea? be appropriated by the data center crew?
We all know it’s only a matter of time before lawsuits in this area become widespread particularly in lawyer happy America.
cgio 1 days ago [-]
This would that the proofs would be almost done anyway, which statistically could not be the case, or that these mathematicians were already progressing thanks to ChatGPT. Still a provenance question, but the impact of AI is unquestionable with regards to outcome.
runarberg 1 days ago [-]
The proofs would only need to be well on their way enough that the only thing needed to finish it were thousands of terawatthours of energy. Something which the mathematician having their work stolen does not have.
thejokeisonme 18 hours ago [-]
They claim that each solution only needed 3 hours on the next generation public model, iiuc. So this theory doesn't really hold up.
runarberg 10 hours ago [-]
I’m not super inclined to believe a word coming from OpenAI until these results have been replicated.
Kotlopou 1 days ago [-]
(also answered similarly to another comment; this is a common question)
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
perching_aix 1 days ago [-]
This reminds me to the early allegations that ChatGPT messages were being replied to by real people...
UltraSane 1 days ago [-]
It would be extremely unlikely human mathematicians able to solve these kinds of problems would accept not getting credit that would set them for life professionally.
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs. The number of people able to understand this level of math and prove it using Lean is a few hundred at most.
fragmede 1 days ago [-]
Hm that's the question, innit. If you were a mathematician, and you have a chance at either being known for solving a particular problem, and the accolades that come with that, and the potential for money, or $X,000,000, for large values of X, in OpenAI stock, what value of X would it take to accept the OpenAI stock? Being known for solving an obscure math problem might get you money, but as academia is unreliable and full of politics, the OpenAI stock is also not a given to set you up with generational wealth, so it's also a gamble.
We'd have to accept that there are hundreds of such mathematicians willing to take the OpenAI stock options for that conspiracy to be true, and that no one leaked being offered that at all, so personally I don't think that's possible. Which leads me to conclude that OpenAI's AI did indeed create the proofs, with some level of help from humans to guide it.
Turneyboy 9 hours ago [-]
That's all true but there is a reason why these problems have remained unsolved for decades and are now all falling at the same time. The reason is not mathematicians being more incentivized.
UltraSane 24 hours ago [-]
It is an interesting game theory scenario. But I think the kind of person who gets a math PhD would value the professional recognition for solving a famous hard problem more than the money.
runarberg 1 days ago [-]
Until this is replicated, we don’t.
jryle70 21 hours ago [-]
We don't. Just like how we know if smcg is from a competitor attempting to create FUD?
matt3210 1 days ago [-]
Agents basically did statistically guided brutforcing. There is no value in what they produced because it lead to no understanding of anything and most likely will hurt the field IMO
woah 1 days ago [-]
Evolution did statistically guided brute forcing. Doesn't mean that biology has no value
1 days ago [-]
dekhn 1 days ago [-]
That is not a correct description of what the AI did.
> apparently they tried the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.
Back-of-the-envelope cost calculation:
3 hours of GPT-Pro compute per attempt / 5% success rate → 60 hours of compute per solved open math problem
At GPT-6 Astra API long-context pricing ($75/M output tokens), assuming 50 reasoning tokens/s, that's roughly $810 per solution:
60 h × 3,600 s × 50 tokens/s × $75 / 1,000,000 = $810
Of course, actual token usage is unknown. But even if off by 10x, $8100 is cheap for solving a longstanding mathematical problem.
Sure yeah, today it's only a 5% solution (to the hardest problems), and we know it will top out somewhere.
But the cost reduction is just too good to ignore.
Terry Tao was saying that we need more mathematicians [0]. My paraphrase of his article is that we will need them to check the AIs, essentially (I missed things didn't I?).
But with a near 10,000X reduction in cost [1] today, it's really hard to argue for that when budgets are continuously tight.
[0] https://terrytao.wordpress.com/2026/09/24/were-gonna-need-a-...
[1] assume ~$150k/year for a mathematician of the caliber that can do this work. 2x that for overhead like healthcare, 401k, etc. Do this for 35 years of work = $10,500,000. Assume about 1 result of this caliber per mathematician. So ~$10M per result. OpenAi is saying they can do it for ~$1k, a 10^4 X decrease.
Sure, but we are approaching a time at which we will have to answer a "why" here. I.e. anything beyond checking and validation becomes of mostly artistic value at some point. And there is nothing wrong with that.
An early (easy to understand) proof that fits the bill for me is proof that the sqrt(2) is irrational: a very natural thing to think about —- for a unit square, what is the measure of the diagonal? And people were killed in ancient Greece over the proof, but it’s comprehensible with basic algebra.
And then you're off an order of magnitude of the efficacy of OpenAI. We don't know how much work of mathematicians is behind all these results, what's the real cost (which includes those problems they couldn't solve), etc.
But a soulless admin is not.
And they care only about the budget.
And the real danger here is that these large groups of blame avoidance mechanisms take at best half a look at the cost of an AI and the cost of a pre tenure track new assistant professor (or whatever), and they decide, oh maybe next year we can hire someone.
And then they do that again next year.
And then the next
The problem is getting the funding orgs to care.
Isn't that more likely the easiest problems in the set? I suppose the hardest problems compared to those that have been solved.
Otherwise your points stand.
“On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems.”
That leaves a lot of ambiguity. Is “result” exclusively “positive result shown here” or “all results”?
Since proving infeasibility is a valuable result, the unpublished problems must not have reached a valuable end state. Thus it’s an operator decision of when to turn the machine off and try another problem. I’d thought they had a three hour time box on this, but apparently not.
https://www.merriam-webster.com/wordplay/bury-the-lede-versu...
In actuality both "lede" and "lead" are correct. "Lede" is mostly an American and recent respelling of the word "lead" and "lead" still remains the predominantly used term (while "lede" is predominantly an American journalistic spelling).
I heard they used lede to distinguish from (hot) lead. But since linotype has been dead for 50 years this is probably moot now.
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
https://en.wikipedia.org/wiki/Nothing-up-my-sleeve_number
?
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap.
Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.
We have such tools for a large number of events that we're incapable of witnessing or understanding in "raw form".
We can't possibly read through thousands of raw records of recordings of individual's heights and make sense of it. But we can use Excel to calculate the average in seconds. We can chart the results and the image is perfectly comprehensible. We can trust the average calculation and the image because we trust the mechanism for transforming the data. We have what I would call a trusted path of provenance. We don't need to nor do we want to read the raw data.
Now we need things like this of the 2nd order. We need mechanisms of transforming trusted paths of provenance into things we can look at and easily verify.
We can do formal verification of code that would be hard to do by hand. We need formal verification tools for formally verifying formal verification tools. Eventually we will need multiple layers of this. As long as the chain and reasoning is intact, we should be alright.
It's kind of like with a horse. A horse is much stronger than us, but we can control it by pulling on two strategically connected pieces of rope.
Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.
Given the LLMs are struggling to explain themselves clearly, but also that these explanations cleared up somewhat by having an expert using one to query some of this research, but also the LLMs can solve problems faster than humans can read the proofs, it is possible for this category to have both examples of things that can be rendered in a human-comprehensible form, and also examples where it is not.
As an analogy: any human can check any single arithmetical calculation from a computer, that's not even particularly difficult. But a Raspberry Pi Zero can do those arithmetical calculations so fast that even if literally every single human was trained to do this at the level of the current world record holder, humanity as a whole could not keep up.
There are many problems that aren't "interesting" in a mathematical sense but can still have valuable proofs. The entire field of formal methods in software development is mostly about this kind of problem. If I want to be sure that no input can cause out-of-bounds memory access, or that my superoptimizer found the lowest possible cycle count, I don't care if it's elegant, I just need some machine-readable proof I can run through my trusted verifier.
If this is true, then what's really the point of math? A lot of math is actually useful. A proof regarding cryptography, for example, would have practical application even if not understandable by humans.
There are mappings between complex numbers and 2D matrices allowing problems to be solved in either domain.
There's research into The Langlands Program looking to connect number theory and harmonic analysis.
There's research into Category Theory looking to define core concepts and relate them do different fields so that results in one field can be applied to another due to equivalence.
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.
[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process
My mental model - for better or worse - is that intelligence has both shape and area, and LLMs are orders of magnitude smaller area but very different shape, and they have more intelligence area in the kind that humans have lesser of.
So yes very different, more in some material ways and lesser in others, but in total intelligence are still orders of magnitude lesser.
My gp comment was acknowledging that when you scale up many instances / brute force problems it confuses that "total area" claim a bit.
To follow the anthropomorphization... 1000 toddlers may have more total intelligence than a grown man, but does that matter?
The problem with these discussions probably/usually fold into differing/loose definitions of intelligence.
perhaps the chat-based ux has sort of fooled us into comparing these things to human intellegence. we don't really do this with chess, or other forms of ai, nor computers at large.
https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)
>Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community.
Proof can't be understood, proof doesn't matter.
Someone at OpenAI, please, work on this.
There is ongoing AI assisted work on this, and precisely the problematic gap in the theory (around Corollary 3.12).
As of now, a required step appears to definitely be missing, but it's not clear if this is a genuine / unrepairable defect.
https://zen.ac.jp/news/zmcpostevent0331e
https://github.com/lana-agents/iut
https://github.com/LANA-Project/genl
There’s not really a clear distinction between these things, in my opinion. Problem solving (and intelligence?) is a mix of search and compression. We like solutions that are elegant (high compression, simple search). But often what appears elegant to some is harder to appreciate for those without the same background knowledge or even the same amount of mental bandwidth (if you’ve ever worked with someone simply much, much smarter than you, you may intuit this!).
This is basically how all AI approaches appear to work. They solve the problem you give them (in many cases) but in a very over-complicated way.
It's like they can't step back from the problem and realise that they need to simplify to make it work better. Nope, just keep hammering more code/proofs against the problem and eventually you'll hit the goal.
RL has a lot to answer for, I guess.
There is nothing wrong with that approach if it accomplishes the goal faster, i.e. you can work at the speed of an AI.
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
Obviously they stole them from the Proof Fairy.
It's grapes so sour they could etch metal.
> I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
I doubt that it is. For the language of proofs to be powerful enough to be able to express an arbitrary proof it would have be Turing Complete. Proving that a proof is the smallest proof of a given concept (i.e. there is no smaller human understandable proof) would then be proving the minimality of a program in a Turing Complete language.
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. And even if they were only good at this stuff, that's still a tremendously valuable thing, because quantitative reasoning and analysis is the bedrock for many, many industries.
oAI is gunning for the largest IPO in history at this point, and they might actually get there.
I would argue that "good at math" was a short hand for "good at X" because mathematicians were historically good at engaging with very complex ideas, distilling them and coming up with precise and concise answers they could validate by themselves.
Given that AI solutions are described as "psychedelical" and they rely on the outside source to validate the result, I don't think the same logic could apply to them.
The "then" in your "if-then" bears a heavy load. Why would society assign so little economic value to pure mathematics if the skills for proving math theorems translate to massive value in "just about everything"? Would you expect top mathematicians to cure cancer if you transplanted them from the math department to a medical research lab?
All of STEM relies on mathematical analysis, and new models are now superhuman at that. And yeah, I'd go a step further and say that the reasoning and creativity required to solve cutting edge math problems probably does translate to other tasks like interpretation of the law, or medical diagnosis, or accounting, etc., for the same reasons that I think most top tier mathematicians would excel at those tasks were they so inclined.
I'd assume that the majority of people who study math take their skills and move onto some related STEM career that isn't pure math. Academia is incredibly small and competitive.
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
Note that the cool thing about rationality is that it does not depend on transitivity. If what companies do to attract investors works, then it is in fact rational of them, regardless of the rationally of investors.
The second paragraph is specific to OpenAIs behavior.
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
As the old saying: great claims require great evidence.
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
But it doesn't follow that these proofs - or any proofs - are automatically on the far side of that limit.
The human usefulness of a proof depends entirely on its human legibility. Much of the value of proofs is in inventing new techniques and concepts and having new insights into relationships. Occasionally you get some game changing insight into practical physics or engineering. But that's rare.
Without that, proving or disproving a conjecture is an excuse for new and original thinking.
Compilers are not the same problem. The point of code is to produce reliable-ish consequences from various possible inputs. It's not a creative exercise in logical consistency, which is what maths proofs are, ultimately.
That's because there's a bunch of people who work on that stuff. Just because web developers don't care about it doesn't mean it doesn't exist. Utter magical thinking.
Maybe there’s some variance or it depends on the reader. That said, like with a lot of other complaints about AI: human papers can be poorly written and poorly explained too. That’s always been the case. And sometimes a proof is just complicated and hard to expose nicely.
I think OpenAI wants the outputs to be unreadable. If outputs "too advanced for human comprehension" become the norm, then every interaction requires tokens. If someone can buy a day pass, get what they need, and leave, then there's no recurring revenue stream.
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
- Feynman
When all you value is being the first, you have not time producing useful papers.
And you're right, it's not really a theme around here, but there aren't that many mathematicians around HN. Check one of the maths forums and it'll be a common point of complaint.
If a journal receives a paper that is unreadable, it is outright rejected with the comment that it should be made readable, regardless the results in that paper. What is the point of having results if you are unable to convince your audience of those results? You could just as well just sit on them and never communicate them. What is harmful about this process?
Which is why it is obvious post-publication peer review is the only sane way to handle this. Formal peer review is not even remotely up to the task here.
That's what a lot of us do nowadays, but instead of maths, it's code that looks sensible on the surface, but when you try to understand it's some kind of "alien" logic , names don't really make sense, etc.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
Makes you wonder if AI can build upon such proofs.
If the AI cannot create proper abstractions, then how can it build a tower of abstractions?
Also, AI has limited context. At a certain point proofs may become so complicated that an AI cannot keep most of it in memory and will not reuse it for new proofs.
so then, if you don't understand the proof, how do you know it's a proof?
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
Welcome to lots of (most?) code PR's in the last year. Though for software at the PR level it has gotten better with the latest models.
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
Sadly, this is also what day-to-day work looks like for a lot of software engineers in industry right now. I spend my time reading and verifying thousands of lines of messy AI-written code. Compared to actually producing something, it's thankless, barely-credited, and unfun work.
To be clear, SotA models/harnesses are really good at making things that work one-shot, and their code golf and debugging game is insane.
When I go to implement, even with a good spec, I end up with code that's 80% of the way there in a fractal manner. The modular decomposition is 80% of the way to good code. The function decomposition is 80% of the way there. The computation structure and variable naming within functions is 80% of the way there. I can walk the AI through it and address each level of issues, but it's tedious as hell, and not clearly faster than doing it myself in some cases.
This is with omp/claude/codex, with Fable 5.1/Opus 5.5/6 Astra. The Chinese models do better at staying coherent, but they're a little less smart IME. I've tried many permutations of "fable planning, opus sub-agent implementation"; "pass a branch back and forth between opus and astra, as reviewers and feedback implementers". I haven't tried the "software factory with an architect and 2 juniors" thing. I haven't tried AGENTS.md beyond the /i-have-adhd, iso-24495 (lol at the Opus 5 induced PTSD), "No cleft constructions or Latinate absolutes" and general project orientation. Model character seems to change too fast to make agents worth it, and I see stuff about skills and excessive AGENTS reducing model capabilities.
Am I holding it wrong?
Matt Pocock says that he knows how, but it comes with a lot of context fiddling.
https://www.youtube.com/watch?v=v4F1gFy-hqg
"Software Fundamentals Matter More Than Ever"
This lecture is very impressive. However I haven't managed to reach AI enlightenment, yet.
"It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results. Basically the paper is so horribly written that it’s impossible to read it without AI help."
This makes one wonder how many of the proofs are actually valid, and how many are just impenetrable hallucinations.
I find it interesting that you state this as obvious. When ChatGPT was introduced four years ago, the argument was that language was fuzzy and probabilistic and that LLMs would thereby never work in math.
I think someone asked here a bit ago whether other search strategies using OpenAI levels of compute have been tried and think the answer is no (even 3 hours of present GPU computer I think is vast compared to anything available 10 years ago, say).
Highly recommend reading it. Very prescient for something written 26 years ago.
https://gwern.net/doc/fiction/science-fiction/2000-chiang.pd...
https://web.archive.org/web/20111121100139/http://www.fantas...
https://en.wikipedia.org/wiki/Division_by_Zero_(short_story)
I also recommend "Exhalation", though that has nothing to do with AI.
https://arxiv.org/abs/2610.08144
> Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it. So if you've got some million-line proof in Lean, spit out by an LLM, you still can't quite trust it, even after validating the problem transcription.
(This is the same category of problem as the huggingface hacking incident, where the LLM finds and exploits an unintended cheaty loophole)
Similarly if the formal Lean problem statements (human generated) are correct translations into Lean (and the original NL statements are sound, which one would hope after decades), and no Lean bugs are abused by the proof (as defined above), then the proof is valid.
The NL/Lean discrepancies are super annoying and will make human analysis hard and fraught, but as many posters have found out the models themselves will gladly pick apart the NL-Lean translation for errors, and so my guess is that finding the discrepancies will not take too long. OpenAI really should have done a dynamic workflow over every lemma and step to ensure pointwise accuracy in the translation.
Why would it know it found a bug?
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
(Incidentally, if you asked me a few years ago I'd have predicted that reliable AI detection was impossible. I still suspect that it's impossible in general and that Pangram would break as soon as companies figure out how to make LLMs stop all sounding the same - but until then, it'll work fine.)
Reminds me of the memes with the little girl making astute comments about the patriarchy to her father.
i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.)
edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.
However, there is the fact that you had to know how to poke the machine so it plucks the right vocabulary out of the training data.
There is very clearly no theory of mind: no inherent internalised modelling of how a human of a particular age thinks, speaks and what they do or do not know. These are the most obvious cracks in the “LLMs are (or will be) the superintelligence” narrative.
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
[0] https://adam.math.hhu.de/#/g/leanprover-community/nng4
> In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully.
I'm however puzzled by the number of proof claims without lean proofs. How does OpenAI have confidence in those, especially if, as noted in the blog posts, the papers are very hard to read?
I believe so, even with all the usual safeguards properly in place: https://news.ycombinator.com/item?id=49672339
> is everyone just assuming that it just be true because the Lean code checks out?
Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.
In other words, they save effort by wasting the effort of others.
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
You might as well say "the obsession with sex has always held humanity back". Maybe, but it's complicated...
It would be fantastic if university and science was like "here is 100M$, play around and develop some 'understanding'". However the reality is that human societies are hierarchical and currently capitalistic which implies value creation and status building.
1) the funding bodies/agencies need proof of value that you're using the resources meaningfully to be able to assign resources
2) Humans are status seeking, power seeking, resource seeking and sexual reproduction seeking. If you hold a lot of power and make decisions, you have more of all of the above.
I am honestly impressed by how many arguably highly intelligent humans are happy to throw out millennia of scientific methodology just because the world-burning probabilistic machine produced an output that tickles their dopamine brain.
Also glancing at the repo. Such an advanced machine, yet failing to output results with a common template. Every pdf is entirely different from the one before, making the work of the poor humans trying to find a sense in it even more difficult.
I don’t think it helps your argument if this is half of your post. The fact that the pdfs don’t follow a common template? That’s the complaint? Human papers also don’t follow a template.
AFAIK, people are complaining about the lack of this.
If the only way to power a "technology" would be to kill an elephant every time you use it, then it's a flawed technology even if it solved fundamental problems of humanity. This is currently how GenAI sprawling data centers operate AND without solving any fundamental problem of humanity.
Also I don't care about A specific template, I just find it ridiculous that some trillion dollar company cannot put a line like "please keep papers' structure consistent" in whatever markdown file controls these things. It's sloppy, like everything OpenAI does.
There was a bunch public literature on these topics. They were major unsolved problems. Lots of humans would want to have solved them and tried. But here’s it an llm that has solved them.
How much more scientific methodology do you want than formalized lean proofs?
What is this nonsensical complaining about templates and formats? You know that’s just like one additional ask of that particular LLM away.
It's already falling apart.
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
I don't know what that means in practical terms, but I agree that's the issue.
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
I studied CS, I have been a software engineer for 15 years. The one thing AI cannot take from me is my understanding, my experience, my intuition. An AI will gladly do whatever you tell it to do, and it is getting really good at it. But I regularly have to steer the AI to the right direction. Not because its default wouldn't have worked at all, but because I knew better.
If I task an AI to build something I don't understand myself, I can get further than without an AI, but I am limited. But when I use an AI to build something I am an expert at... I see that it is better than what non-experts get AIs to do.
You lose when you stop learning. You are a PhD student, you have learned to learn. Keep that. Keep learning actively. It was never about the PhD anyway: when you go to the industry, nobody cares about your PhD. What you have is what you know.
Unfortunately, the current world is not set up for that, and until we get there, yeah things are gonna get pretty hairy. But if and when we get to the other side, I think the future will be glorious.
Almost every single person who earns money for labor is, or will very soon be, in exactly the same position as you, many are just either unaware of how fast the change is coming or are deep in denial about it.
None of us know what to do about it other than hope that we find a peaceful political solution prior to the economic collapse.
Gödel's incompleteness also constrains human and AI mathematicians. Both just strive to prove whatever can be proven in the system they are working in.
And what the hell will the frontier labs have by then?
Maybe I'm overreacting, I'll have to screw my head back on before I can process this.
^^ half of the comments on this thread
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
> These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
1) AI is winning programming competitions, 2025 was probably the last year we've had human participant winning* 2) Math is different because there's formal verification.
* Of course competitive programming is different than enterprise programming, but competitive programming is closer to math.
My dislike lies in how it devalues skill + knowledge acquisition, disrupts art/software communities, and that it's now bloody impossible to tell the difference between someone who actually knows something from someone who merely regurgitates shit from their chatbot.
We can stop once they find a single productive use case for their billion dollar machines.
Previously, due to the cognitive limitations of our brains, it was difficult to tell whether resolving a conjecture genuinely required years of dedicated effort, or if the answer was already hidden within existing human knowledge. Or whether a problem is just waiting to be searched out through ingenious or even brute-force means. With AI, we can now offload this search, especially across different areas of math, allowing us humans to focus our minds on discovering new mathematical structures and techniques.
Indeed, if one believes what Hilbert believed: we must know and we shall know, he should feel happy, as the belief has never been about who solves a problem, but about whether can advance our understanding of the universe. With the help of AI, we have a lot more possibilities.
(To my credit, I have warned them since more than a year ago that we will reach this point.)
But are any of them going to be able to do any significant amount of studying maths without being homeless?
Why not? Maths studying correlates with high intellect, high intellect correlates with competitiveness in the job market. So their chances of finding a job that would allow to have much of free time (for doing math) without being homeless are pretty good.
It's just that we lost one of the important ways to demonstrate understanding.
Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding.
You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all.
On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers.
For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia.
Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery.
For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems.
So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle?
What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge.
And then we need to find efficient ways to detect these ideas and describe them.
Really this is very exciting and opens up whole new workstreams for mathematicians.
The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?
The most important takeaway from this is that mathematicians finally consider us CS grunts their equals.
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.
In the past 2 years the AI's started solving math problems in roughly the order of "hardness" as ranked by humans.
― Alfred North Whitehead, "An Introduction to Mathematics" (1911)
-Boaz Barak
Emotions can be funny.
>I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
Also, Ted Chiang's 2000 short story "Catching crumbs from the table" <https://np.reddit.com/r/singularity/comments/1wzu5gf/this_mi...>.
Pretty funny way of putting it. Presumably model X+2 will be able to explain these in elegant, human legible ways, though (as well as solve the remaining 95%)
I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.
Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.
https://hideoushumpbackfreak.com/algorithms/algorithms-stras...
I know, I know. Don't anthropomorphize AIs. They hate it when you do that.
>If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing).
Whole lotta yikes
Usually 9 year olds imitate adults when they regurgitate such words in these circumstances. What a sad state of affairs.
The checks notes meta was retired ages ago btw.
The irony!
now concretely: Your conclusion doesn't follow the evidence (non sequitur).
You made a generalization (inductive inference) based on a single datapoint (here's where my pointed faulty generalization or inductive fallacy are the come in) for the general problem. I suggest Popper if you want a deeper treatment of the problem of induction (The logic of scientific discovery is a good start).
Besides generalizing for free, you cooked up evidence that didn't exist - mainly that the child was imitating adults (again, non sequitur).
As a side note, I think it's in poor taste to make remarks about someone's children in what's a technical and informal context. That may have not been the intention but, it's what it looked like.
People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.
That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to.
It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof.
I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.
I bet sharing unreleased models with outsiders safely ain't that easy.
https://cepr.net/publications/ai-didnt-steal-the-mathematici...
Frontier theft is just faster.
Reading the account from Scott's wife, the proofs seem to resemble more a Zelda speedrun with a bunch of strangely exploited bugs than an actual thing.
If you have any great skill at the upper level, you better work off-line/local private and not share, but for many the temptation will be too great.
1. Help understand, check, and explain the results.
2. Write new works explaining or refuting the new approaches and results in more lucid language.
3. Advance the field further.
I don't know why this is not obvious. Each step is intended to support human understanding, not to replace it. Any mathematicians who don't do these will be left behind, and if none do it, the field of human mathematics itself will become obsolete. All I am hearing so far is excuses.
And if humans decide to give up on mathematics because the process of using the machine is so tedious, then there will be no value in automated theorem proving.
> resistance will fail utterly as long as AI itself can understand prior AI works
That's not the issue in the long term. Look further: where did these questions come from? They came from humans who had been exploring mathematics and noticed interesting patterns that cropped up again and again.
The people who posed these questions had a profound understanding of the background maths that led to them. Me - I don't. Take one of the results he quotes: "Positive solution to the Unitary Synthesis Problem". He may as well be one of your alien races yabbering to me in an alien language. The question is meaningless to me. The answer the AI came up with, correct though it may be, is equally meaningless. You may as well have told my dog how to find eigenvalues.
If we let AIs do all the understanding so we know nothing about math, there will be no human understanding that lets us find new questions, or understand the value of any questions AI might pose and find the answers to.
The "summarization" process can be multi-step. Initial steps develop new frameworks and prerequisite concepts. Later steps build upon these frameworks and concepts. Some can be simplified too if this can be done without loss of fidelity. The summarization comes last, and is meant to preserve sufficient context.
> I don't think anything about how to react to a paper being released in this state is obvious.
If it helps, exaggerate the condition a bit by imagining coming across a repository of alien knowledge.
If you don't understand why the paper is in its current form, how do you know what changes would damage the fidelity for the AI? I think your example of alien knowledge is misleading because it suggests the aliens don't need to understand the info - but in this case we intend to keep using AIs to expand this work. So if you transform it into something that humans think makes more sense without understanding why it's in a form that we hate now - you risk losing something in the form that you don't understand.
#3 at this point might need more human intuition; but that might be a 2026 problem.
The he puts up preemptive straw man arguments against doomers. His blog has become a joke.
We all know it’s only a matter of time before lawsuits in this area become widespread particularly in lawyer happy America.
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs. The number of people able to understand this level of math and prove it using Lean is a few hundred at most.
We'd have to accept that there are hundreds of such mathematicians willing to take the OpenAI stock options for that conspiracy to be true, and that no one leaked being offered that at all, so personally I don't think that's possible. Which leads me to conclude that OpenAI's AI did indeed create the proofs, with some level of help from humans to guide it.