Hacker Newsnew | past | comments | ask | show | jobs | submit | nilkn's commentslogin

None of these incidents involve single instances of commercially or publicly available systems. They all involve large swarms of internal models. The stuff you're describing is not the research frontier. It's really not even close.

I think it's easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we're also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.


The work is either: (1) the act of concealment: generating an AI book that cannot be detected as AI generated; or (2) actually just writing the book yourself.

It's easy to come up with new open problems. It's hard to come up with new open problems that seem to teach us something fundamentally new about the world. Our current batch of problems went through a complex selection process over decades (or centuries) based not purely on difficulty but also on perceived insightfulness.

I studied math, but I am not a mathematician, so I think I have a slightly different perspective on this than Tao overall. This is certainly the definitive end of an era in mathematics, but I think he's wrong that insightful new open problems are truly non-renewable. They might be non-renewable by humans at the rate at which they are being closed, but I see no reason why AI systems could not also discover insightful new open problems. In fact, once we have Riemann-capable AI mathematicians, I'd personally love to see what the next Riemann hypothesis is, which even these AI systems cannot solve with any amount of available compute.

I think we're about to find that, on the spectrum of mathematical intelligence, the best human mathematicians were only a fraction of a percent forward from the very beginning, and there's a vast universe of mathematical depth that's beyond our ability to imagine or work on directly in any way. We're used to feeling like we're able to directly perceive the Platonic realm, but we're almost certainly going to discover that our own minds, even when joined together over centuries of deliberation, can only interact with a tiny little shadow within it.


I haven’t been following the AI proof stuff very closely, but the impression I got was that these models are producing massive Lean programs that prove the statement one way or another, but are quite difficult to fully understand.

Actually, I have to admit I don’t really know what math is. With physics we suspect there’s a universe, and when we study physics we’re improving our description of the behavior of that universe, right? The universe exists whether or not we know how it works.

Eventually, as you suggest, maybe we’ll hit math that won’t fit in anybody’s head at all. What is the nature of mathematics that doesn’t fit in any human’s head? Does it even exist in some sense?


I think math is compressible structure. That's why we care about something like the Riemann hypothesis but, to use Tao's example, we really couldn't care less about computing the 10^10^10th digit of pi. The first compresses a vast amount of information about the primes, while the second decompresses information that we've already compressed (a few lines of code can define every digit of pi).

Most patterns that exist are incompressible. Math is basically a search for those compressions that do exist. An example I personally really like is the amplituhedron: a geometric structure that humans have just barely been capable of recognizing compresses information about scattering amplitudes and Feynman diagrams. That one happens to be within our reach, but it's right at the edge, and we can only imagine what glorious, wondrous compressions exist in abundance beyond the edge. Math accessible only to superintelligence would exist entirely beyond that edge, compressing patterns whose existence we cannot even detect using objects and constructions that we cannot grasp.

As an aside, I also think this is why AI is quickly becoming superhuman at math: intelligence is essentially a form of pattern compression.


I think part of mathematics is taking things that don't fit in our head and giving them human abstractions so they can.

Take infinity. Infinity can't fit in your head, hell, it can't fit anywhere, but you can abstract away the endlessness and look at infinities of different sizes, et al.

Now, is there a single formula for something actually represented in this world that would take most of a humans life just to read it, no idea.


The models produce both Lean code for formal verification and a traditional-style narrative proof. Like the general long-form output of frontier models, the math papers produced appear to be generally correct technically, but written in an ungraceful and sometimes hard-to-follow style, so they are often polished by a human mathematician as of today.

What you are saying implies that by some technique that hasn't been discovered yet, we can make the models to have the capabilities of extrapolate the information they are trained on and also interpret that what they are extrapolating are Riemann-capable hypothesis. I do believe it will accelerate the discovery of that "vast universe of mathematical depth that's beyond our ability" but at the cost of removing the "fun part" of solving the problems. Not sure if the community is willing to do that.

What's just as interesting is this morning Axiom Math announced 212 and OpenAI then appears to have rushed out their 186 announcement just 1-2 hours later followed by Astra. Did they accelerate the release of Astra itself? Not necessarily, but it definitely looks like they ended up pushing much harder and faster than planned on their 186 result. X activity suggests Anthropic had a similar result as well but wasn't as fast as OpenAI in packaging it up and sharing it in response to Axiom, so they mostly just bolted onto OpenAI's messaging.

The reason I think this is interesting is that Axiom is a tiny lab in comparison that wouldn't have had access to Astra at all. I'd be curious to learn how Axiom is able to effectively compete at this frontier with vastly fewer resources.


My most controversial opinion by far in tech circles is that I still just use a standard Windows gaming PC as my home desktop. My current machine I just bought pre-built from Microcenter, complete with a 5090 and everything.

I can fire up a Linux terminal with WezTerm and WSL2 at any point. It's customized and beautiful and totally fine. I have Codex running in one right now. I can listen to Dolby Atmos music through Apple Music or fire up a game with zero compatibility issues and full RTX support. It's just versatile like nothing else. I pair it with a gigantic 48" LG OLED TV as my monitor.

The only thing that might tempt me away from this is a fully loaded Mac Studio with 512GB of unified memory. That would be a real capability gap from my current machine. But I've contemplated wiping Windows and installing Omarchy, and I just can't figure out really what I'd gain, but what I'd lose is quite clear.


Windows has also come a long way from a terminal perspective. Sure, the UI is a bit of a mess, but Powershell can do anything in the UI from the command line, and agents are very capable with PoSH. If you really care about ricing the UI, there's hundreds of utility apps to do almost anything you want.

Agents are also able to tweak and debug Windows errors, since the registry, group policy, event log, and other Windows internals have been largely unchanged for 25+ years and are well documented. All have old command line tools or modern Powershell to manage.


Has it restarted losing your session to install an "Intel Corporation - Extension - 22.1120.5.12" yet?


you have to understand that 90% of HN use macs and the only time they see windows is once every 5 years when a relative asks to setup a new or clean an infected one

The rank of a rational elliptic curve can be seen as a measure of how complex its arithmetic structure is. Roughly speaking, you can imagine that a curve of rank r has a substructure of dimension r. So a curve with a high rank is a pretty exotic object. This curve here has a 30-dimensional (or greater) lattice substructure. You can think of it as being possible to arrange the rational points on this curve into a predictable structure in a vector space of dimension at least 30. In that space, there would be at least thirty independent directions in space that could be combined together to produce distinct rational points on the curve.

To really quantify how exotic, it's conjectured that curves with rank 2 or greater have an asymptotic density of zero. That doesn't mean they don't or can't exist, but it does mean they become vanishingly rare, so finding even individual examples of high-rank curves has been absurdly hard.


I do laugh at mathematician's (assuming you are one) explanations to laypeople.


It's not really quite that simple. You wouldn't hire Jeff Dean to do bug fixes in your mobile app. I'm sure he's capable, but I honestly doubt he'd stay interested and focused on it enough to really do a good job. That doesn't mean he's not a superstar.

What it comes down to is that there are different types of high performance. Some people are good at just executing tasks given by their manager. Some people are good at being generative, thinking across boundaries, acting autonomously, creating value without direction, etc. A term like "superstar" will get disproportionately applied to someone really good at the latter and rarely someone really good at the former, because the potential impact of the former is typically strictly capped, while the latter is uncapped.


I think if we’re calling someone a “superstar” there can’t be a “but” but I also acknowledge this isn’t that important of a nitpick lol


Math that humans don't understand but nonetheless allows AI systems to develop breakthroughs in various fields of science, technology, physics, engineering, medicine, etc., would have great value to humanity even if it doesn't help humans understand abstract truth at all.

Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it, it initially seemed useless, then another AI system found a predictive model of electromagnetism using it.


But at that point you have full AGI and its not just today's models. Today's models still need humans to understand things since it builds upon human knowledge.

When you have full AGI of course you no longer need humans to understand math.

> Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it

Developing multivariable calculus requires much more than just solving problems though, it requires defining an entirely new system and space. That is not the situation mathematicians face today, modern AI cannot do that.

When talking about mathematicians and AI don't use fictive examples, we can look at what AI can do today and extrapolate that they can do more of that tomorrow, that is what we have to work with.

In the case you posit where AGI exists there is no reason to even discuss what is left for humans to do, since AGI is defined as when humans are no longer needed for anything, the AGI can do every bit of thinking humans can.


I don't find AGI to be a useful technical term, as nobody can agree on what it means. For instance, you used it at least five times here, but you never defined it, and I could point to intellectually credible people who would say we've already reached AGI.

Anyway, if we put the AGI framing aside, I think the main point you're making is that AI mathematics hasn't yet demonstrated the ability to theory-build in the way that the great human mathematicians have (Grothendieck, Scholze, etc.). And I'd agree with you on that. Where we disagree, I suppose, is I think that capability is coming -- I don't see anything that would prevent its development.


It's reasonable to say that current or near future AI can find novel mathematical results and develop applications from those.

The idea of AI stepping from a graph theory/combinatorics innovation to some new and useful algorithm isn't crazy.


Exactly. Soon, human brains are going to be obsolete.


So then it sounds like you agree that math has additional utility beyond just human comprehension.

If I understand you correctly, you're just qualifying that that will only be the case when AGI exists. To be clear, I actually disagree with you here because I think it's very plausible to find a use case for human-incomprehensible math proofs before AGI exists. I'm just saying it sounds like you're agreeing with the parent comment that math is not purely about human comprehension.


I believe this rule of thumb will come to fail. The combination of superhuman mathematical reasoning and synthesis in upcoming AI models plus the rapid build-out of scalable formal verification infrastructure means this exponential in math is going to take off quite explosively, and we've barely seen anything yet. Mathematics is going to decisively move beyond human ability fairly soon (within our lifetimes, if not much more abruptly). It seems abundantly clear to me that much of the work will only be immediately accessible to AI, and rather than trying to explain all of it back to humans we will rather focus on explaining the portions that humans would benefit disproportionately from understanding.


Maybe that will be true when it's math with practical applications, but most theoretical math isn't like that. If it's not practical and it's not for mathematians to understand, what good is it?


You could have one really hard to understand proof of a theorem and then a lot of interesting human-understandable stuff that relies on that theorem. We already have lots of proofs with oracles, where you can work out consequences of what kind of structures and solutions could exist if you had some magic thing to solve a hard part, so it just seems like a variation on that. Many people learn calculus or even the real numbers without understanding the complete formalization from set theory.


We have thousands of years of precedent that suggests that breakthroughs in mathematics tend to accumulate into broader technology breakthroughs in other domains.

Why does this tend to be the case, even when some of the smartest people in the world have historically predicted incorrectly that certain branches of math would forever be useless (e.g., number theory)? I can only offer my own theory on that, but my guess is that mathematics is simply a predictive framework based on pattern compression. A more powerful pattern compression framework accelerates every single field that relies on pattern recognition or prediction of the unknown based on patterns.


It sounds like the idea is to turn on a math generator and keep running it until it generates something interesting. And it might be fun to try it. But if it’s too much output to read and we don’t understand the output either, how does anyone recognize when it’s done something that’s practically interesting?

The output might make a cool screen saver as-is, but we probably need a way to evaluate it somehow.


This is relevant, but only after AI has solved all the open problems including Millenium problems. Until then, as AI keeps solving harder open problems, people will pay attention and be interested.


Yes, of course we'd need a way to evaluate it. I don't right now have a fully conceived answer to what that will look like. But I'm confident at least in saying we would not evaluate it, like Tao is suggesting, by only accepting something once a human can easily teach it unassisted to another human. That sets the bar dramatically too low and would quickly become an extraordinary impediment to progress. You'd have to think of yourself less like a researcher and more like the director of the world's largest research institute. It's highly unlikely you'll understand or even care about every single paper every one of your researchers is producing, but you'll care about the overall research direction and whether the intermediate results are accumulating into outcomes you consider meaningful. How to do this where the institute is based on superhuman AI mathematicians is an unsolved problem, but I see no reason to imagine it's unsolvable.

Let me make up an example of where I could imagine this going. Something we essentially cannot do right now is predict coarse-grained phenomena from systems that involve millions or trillions or more of interacting components. Over hundreds/thousands of years of experiment and theory we've derived laws that essentially do this in a few special cases, but we have no systematic theoretical way of doing it in general, and frankly I think it's beyond human ability. Whatever deep patterns or structures exist for doing this in a general way I think are simply out of reach for us.


We have thousands of years of precedent that suggests that breakthroughs in mathematics tend to accumulate into broader technology breakthroughs in other domains.

That's a misconception. Only a tiny percentage of mathematics has seen any applications whatsoever. There are vast libraries full of mathematics no one (in this discussion, anyway) has ever heard of that no one reads anymore and has never been applied to anything.

This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."


> There are vast libraries full of mathematics no one (in this discussion, anyway) has ever heard of that no one reads anymore and has never been applied to anything.

And that's an issue why? It would seem to me that producing that also produced the mathematics that revolutionized the world repeatedly for centuries. I would go further and claim that, if you want the mathematics that revolutionizes the world, there's no way to get it without advancing mathematics as a field broadly. Those are not two separate activities, and thinking that they are is indeed a misconception.

> This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."

You're right: "prove all the math" does not make sense on any level, and nobody serious would phrase any of this in that way. I certainly didn't.


And that's an issue why?

The issue is SNR: signal to noise ratio. Generating exponentially more mathematics, particularly if the process is indiscriminate or optimized for something other than usefulness or mathematical relevance (such as optimizing for machine-provability), does not imply that we get exponentially more applications. We may end up halting the progress of applications altogether as the entire capacity of the world's mathematical apparatus is consumed by the interpretation and investigation of machine-generated proofs.

You can already visit arXiv and find vast numbers of not-yet-published mathematical papers. Most should never be published. None of this junk is benefitting humanity in the slightest.


It could just as easily be the opposite: it could end up being far easier to reasonably direct and evaluate the research direction and output of AI systems than human mathematicians, who are forced to specialize over decades and essentially cannot pivot and often can't even meaningfully evaluate each other's work.

Moreover, the disdain you have for low-value output in mathematics is not unique to you. Talented mathematicians don't like it either. Your mistake is assuming that AI will cause math to be dominated by low-value outputs. In fact, the opposite is likely the case: the marginal value of proofs will fall so low that the bar for meaningful research will become dramatically higher, not lower. I expect the goals of research mathematics to become extremely ambitious relative to the past, organized around substantial and enormous goals, not mass-generated slop as you're imagining.

Of course, yes, there will still be lots of slop, just like GitHub is full of AI coding slop, LinkedIn is full of slop, etc. But that's a generalized issue of the AI era, not unique to math.


Moreover, the disdain you have for low-value output in mathematics is not unique to you

I didn't say anything about low-value output. No one actually knows the value of any particular piece of mathematics within that deluge. Mathematicians don't have a magical ability to differentiate high-value mathematics from low-value merely by reading paper titles and abstracts.

The dirty secret in the mathematical world -- that has been going on for a long time already -- is that papers get attention based on the reputation of the authors, not on the rigour or validity of the proof. The big headline-grabbing papers are getting read by mathematicians because AI researchers have leveraged media exposure to bypass the reputation network, but media exposure doesn't scale.

When everyone is using LLMs to generate proofs, only reputable mathematicians will be able to get their work read. And herein lies the crux of the problem: an exponential takeoff in the volume of output from respected mathematicians will leave a critical shortage of readers.

it could end up being far easier to reasonably direct and evaluate the research direction and output of AI systems than human mathematicians

That's baseless speculation. All indications so far are that LLMs produce proofs far longer and far more complicated than humans are capable of, such that only machines can check the proofs for validity. Digesting them into a human-readable interpretation of the results is an open problem.


> I didn't say anything about low-value output.

False. You very plainly did. You simply used the term “junk” instead.

> That's baseless speculation.

It might be speculation (as is much of what you’re writing), but it’s not baseless. Obviously, it’s quite easy to direct AI agents, a single one of which can pivot across all of mathematics, unlike all human mathematicians.

> All indications so far are that LLMs produce proofs far longer and far more complicated than humans are capable of, such that only machines can check the proofs for validity.

I’m unaware of any clear evidence of this. Hence, it appears to be baseless speculation.

> Digesting them into a human-readable interpretation of the results is an open problem.

I’m unaware of any clear evidence of this. Hence, it appears to be baseless speculation. Moreover, and more importantly, to my knowledge there hasn’t been any meaningful result in AI mathematics so far that has posed any kind of blocking issue on understanding it yet.


One day it might be for the AI's pleasure, the same way it has heretofore been for ours. Or if you prefer, as a byproduct of its programming to acquire knowledge.


People say similar things about automation of software engineering. Different, but similar.

I'm deeply suspicious. I do not yet have a concise statement for why, but a lot of literature on the sociology of knowledge work sort of points at my thoughts.

Section 5 of the Thurston article cited by Tao touches the elephant. Raduchel's article on the economics of software [2] also touches it.

I've tried to put words to this for a few years. I think I'm just going to start writing versions of it as see if that helps me shape the thought into something more concise.

So, in the spirit of this article's style, here are some postulates:

1. There is a sociological process happening in the production function during knowledge work.

2. That production function and the associated sociological process spans years or even decades, and must outlast many of the artifacts that are produced during the early years of the function.

3. You cannot get the right lines of code or the right theorems proved without running that sociological process alongside the artifact production process.

4. It is impossible to completely separate the sociological process from the artifact construction process. If you just iterate on artifacts then too much of the required hidden state is lost to make progress in the right direction. This is true even if you include distilled artifacts capturing pieces of the sociological process (eg meeting notes, documentation, commit logs, prompts).

5. So you need that sociological process, or something like it, to still happen.

6. For a lot of knowledge work that process plays out in extremely high-fidelity social interactions [3] that we have not yet captured in the datasets that would be required to reproduce those dynamics.

7. And even if we do collect that data, our current architectures and training algorithms and hardware would be useless given the size of the datasets.

So: the technology today gives us the ability to iterate on the production of artifacts. But it does not sufficiently simulate the social process which gives rise to the Right artifacts.

This isn't exactly what I actually think, but it's a version of the thing that I intuit when I watch heavy use of AI in both software projects and formalization projects. And simulating that process feels way harder than people are currently assuming.

[1] https://arxiv.org/pdf/math/9404236 Section 5.

[2] https://www.nationalacademies.org/read/11587/chapter/11 pp 166-168.

[3] there is a reason we still gather in-person around white boards, and why doing so is more crucial for some types of work than others.


Many eminent mathematicians did their best work while not talking about it with anyone, sometimes in isolation. Newton's calculus, Perelman's poincare, much of Grothendiek's work, Wiles's fermat, Ramanujan's earlier days. That's not all top mathematicians as you can look at Von Neumann as a sociable counter-example. But it shows that discussion of your current ideas is not a requirement. Grothendiek goes so far as to say it is a net negative for mathematical creativity because it is difficult to resist thinking like the herd without some level of seclusion.


> eminent

I would use a different adjective: exceptional.

There are not not times where lone geniuses produce amazing output. But most of humanity's progress over the last few thousand years (or at least certainly the last few hundred) resulting from a different type of work.


That hypothesis requires substantiation. Just because 10 people in a room came up with a breakthrough doesn't mean a sociological process contributed positively to that breakthrough. People generally like socializing, so the base rate is going to be high by default. The breakthroughs might have been more numerous if they had each followed Grothendieck's advice (who is a contender for the world's most talented theory builder) and intentionally decorrelated from one another to free themselves from convention.

Referring to my list of examples as exceptional in the sense of being rare is rather unfair, given how numerous these examples are relative to the body of mathematical work we would consider incredible and especially relative to the desire of the average human to socialize. The fact that such a large % of that body of work occurred while the individual was in relative isolation is something we should pay attention to.


> People generally like socializing

If people don't like socializing then your argument not only falls apart, but actually leads us to something like the opposite of the conclusion you found.

And people definitely don't like socializing! Because we're not talking about "socializing" in the sense of parties or even water cooler conversations.

We're talking about socializing in the context of a group of professionals working together on under-specified problems/solutions.

The English word for that type of socializing is:

Meetings.

Saying that "people like socializing" in the context of professional knowledge work is like saying "people like woodworking" in the context of stick framing cookie cutter developments in Georgia in the summer.

I don't think I've ever worked with a productive scientist or mathematician who didn't loathe meetings. In fact, I think I can probably count on one hand the number of people I've met who seem to not loathe meetings.

Also: there are pretty powerful social/professional forces -- in Mathematics and related disciplines in particular -- which incentivize people to play into the perception that they are a "lone genius" or at least the "main character".

So the group of people we're talking about both hate meetings and have a strong incentive to be perceived as an independent genius/main contributor.

And yet the vast majority of good science and mathematics involves large groups of people (who mostly complain about the resulting necessity of meetings, lol).

I'll also submit that most of the "lone genius" instances are confounded by the strong social incentive to be perceived as either a lone genius or at least the main contributor. There are at least a few prominent historical examples of math/math-adjacent giants who "worked alone" but were in fact in deep conversation with both the literature and their peers.

But, again, there's a lot of gray area, and that's what I meant in my original post when I said "This isn't exactly what I actually think". Two examples of gray area:

1. I'll certainly concede that there are some types of work that do not require that sort of social process to happen. That's the sort of work you seem to be pointing toward here. But opportunity to do that type of work is rare in general, and those opportunities are far more scarce today than it was in 1690 or even 1990.

2. There are also some types of work where the corresponding sociological process is pretty simple and is probably sufficiently captured.

So I'm not bearish about what's possible, at least per se. But technology executives and analysts in particular have some really fundamental misunderstandings about how this type of labor works in most cases, and I see some of those same misunderstandings among lots of professional mathematicians.


I'm just barely old enough to have experienced a couple waves of technology being introduced and creating these sorts of life optimizations. It almost never actually improves quality of life. It just makes life more optimized, and since everyone else is doing the exact same optimization as you, that optimization goes from novel and fun to table stakes to a basic necessity to compete and function in society professionally, at which point it's really nothing but pure efficiency, not genuine improvement in relative happiness.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: