
Does OpenAI’s Astra Mean AGI Has Arrived? | AI Reality Check
August 6, 202629 min · 5,405 words
Show notes
Cal Newport takes a critical look at recent AI News. Video from today’s episode: youtube.com/calnewportmedia (0:00) Does Open AI’s Astra mean AGI has arrived? (3:07) What actually happened? (11:35) What does this mean for mathematics? (21:35) What does this mean for OpenAI? Links: Buy Cal’s latest book, “Slow Productivity” at Sponsor: Thanks to Jesse Miller for production and mastering and Nate Mechler for research and newsletter.
Highlighted moments
Astra is a combination of some sort of underlying LLM and a very complicated, what they would call an orchestration program. I would call a harness or an orchestration harness, but a very complicated program, not machine learned, written by people, that allows, makes many calls to the LLM.
“They all three share the following properties. They're construction-based. They're either construction-based existent proofs or counterexamples, quantitative bound improvements, or generalizations or extensions of a known result.”
“It's not doing new math, it's using our math, and mathematicians are running it and having to try to understand what it's doing.”
Transcript
Introduction to Astra
0:00Over the weekend, OpenAI announced that their new pre-release AI system, Astra, had produced 10 math results that, and I'm quoting here, resolve or make substantial progress on long-standing open problems. Now, Noam Brown, who's actually one of my favorite AI researchers, tweeted out a list of what these 10 results were, and they included things like a better bound for high-dimensional sphere packing, a new lower bound example for arithmetic circuits, and two counter examples
0:30from extremal graph theory. Now, I was personally excited by these results because many of them touch on areas of applied mathematics that I actually use in my research as a theoretical computer scientist. And it was sort of neat, right, to see that you have this giant AI company happening to turn its attention like a sort of GPU-powered eye of Soarin' on this little narrow field of discrete mathematics where small communities like the ones I'm in actually work
1:03on. We're like, whoa, we're in the spotlight now. So that was exciting. But then, predictably, that sort of online AI is a force that gives us meaning crowd couldn't just let us nerds be excited by a new tool. They had to try to connect it like they do with every AI announcement to some sort of demented eschatology where this time, for sure, we've just launched ourselves into an imminent new world of massive disruption. Gary Marcus actually did a good job of rounding up some of these reactions
1:39into a newsletter he published. So I'm going to read a few quotes that he highlighted. Dean Ball said in the aftermath of the Astro announcement, quote, everybody in the world will soon be able to use the model that made these breakthroughs for every problem they face in life, no matter how mundane. Kevin Roos said, almost nobody is pricing in the possibility that the models just keep plowing through every discipline the way they're plowing through math. Matt Schumer said, looks like the GPT next is going to make Fable look like a toy and usher in a golden age of science. All right,
2:14so what's really going on here? What is Astra? Does it represent a major leap over other existing AI systems? Do these new math results it produced represent humanity crossing an event horizon towards inevitable artificial superintelligence, or are they merely evidence of a more narrow evolution of math tools? Well, it's Thursday, so it's time for an AI reality check episode of this podcast, which is the perfect opportunity to go searching for some measured answers, which is exactly what we're going to do. As always, I'm Cal Newport, and this is Deep Questions, the show for people
2:49seeking depth in a distracted world. All right, so I want to break up my discussion of what's going on
What Happened with Astra
3:01with Astra here into three big questions. Question number one, what actually happened? All right, so let's get into the basics here. Perhaps the most basic question of all is, what type of system is Astra? Isn't that obvious? If you look at the OpenAI announcement, they describe it as, and I'm quoting here, our next major model. Now, this gives the impression that it's essentially an LLM, maybe an LLM that has a sort of minimal chat harness on top of it so that you can have interactions with it. That's
3:34what I think about when I think about model. But there's leaked details from a recent Washington, D.C. briefing where Sam Altman came to brief the U.S. government on Astra, and some leaked details from that briefing makes it clear that actually Astra is a combination of some sort of underlying LLM and a very complicated, what they would call an orchestration program. I would call a harness or an orchestration harness, but a very complicated program, not machine learned, written by people,
4:04that allows, makes many calls to the LLM. It can spawn also multiple agents, each making their own calls to an LLM. So a lot of logic, a lot of structure. So what we're really looking at here is not a chatbot type system, but instead something much more like the alpha proof or the, what's it called? Alpha zero? Alpha proof? You know, I forgot what it's called. Alpha proof, I believe. So there's a deep mind system called alpha proof that's also incredibly structured and orchestrated where
4:39you spawn different agents to try out different mathematical techniques and you come back and you evaluate. Actually, in alpha proof, you write your results into a formal proof language called lean, which can then be automatically verified. So it's a much more complicated sort of math control program that's using some sort of underlying LLM. So that's our best understanding of what is happening here. This is different, for example, than what OpenAI used to recently disprove the unit distance conjecture. We did an episode about that a couple months ago. There, they made pains to indicate that this was actually just a general reasoning model that they prompted like a chatbot
5:13and then poured through its response to find a proof. Now we have a much more structured math solving program that's calling an LLM. All right. So that's just the best we know about what type of system we're actually running here. Okay. Another basic question about what just happened. Why did
Why Focus on These Problems
5:28OpenAI focus on these particular problems of all the problems you might focus on? These are relatively important problems in their narrow subfields, some more than others, but it seems at first glance to be a pretty random collection of areas that these 10 problems are drawn from. Well, there are some properties they all share. I'm actually going to read these from a tweet that Gary Marcus retweeted. They all three share the following properties. They're construction-based. They're either construction-based existent proofs or counterexamples, quantitative bound improvements,
6:00or generalizations or extensions of a known result. I would also add, just putting on my own mathematics hat that they seem to be largely in a discrete mathematics space with graph theory and combinatorics being very well represented. This would contrast with like the more continuous space in which scientific math from physics to biology to engineering occurs, right? So we're in this sort of discrete math, especially with graphs and other combinatorial structures. This makes sense because these sort of discrete structures and the logic that surrounds their descriptions and properties are,
6:31I think they're well suited for LLMs, for discrete token-based representations. It's probably the, these are like the right areas in math to go looking for problems you can solve with LLMs as opposed to like going into physics and trying to simplify massive equations or find new ways into approximating differential equations. Another key answer to why they looked at these particular problems is that these were the problems that they got solutions for, all right? See, I think there's a sense when you read the online reaction to these results from the AI is a force that gives me meaning crowd that open AI
7:04just had this beautiful new model and they started throwing problems at it, solved every problem. And after they got through 10, they were like, well, we got to stop and tell the world about this before returning to our efforts of solving all of science. In reality, however, that's not what happened. Noam Brown, who worked on this project, actually tweeted the following, And yes, we did try other major problems without success. Sadly, no Millennium Prize problems yet. Millennium Prize problems, by the way, is a list of major open problems, each of which carries with it a million-dollar bounty to incentivize people to try to solve them. All right, so, you know, there's this general
7:37space where these type of systems work well within the broader mathematics, and then they sort of are searching within that space, trying different problems to see which ones they're actually able to solve. And that's why it looks like a sort of random collection of areas, as opposed, for example, to saying, hey, let's just turn to a chapter of the handbook of combinatorics and just solve all the open problems. And now that problem, that question has been resolved. It's more scattershot than that. Another basic question
Astra's Capability
8:03about what actually happened. Does Astra represent a major leap in AI capability? So, sort of more generally, there is, again, this sense from the AI as a force that gives us meaning crowd that Astra is this new leap forward, maybe even a crossing of an event horizon, right? So, something changed that now allows us to say something like artificial superintelligence is imminent. Here, I think there's bad news for open AI. Their PR department's very good at these announcements, but it's really unclear to the rest of us who kind of know this world of
8:38mathematics. What is this Astra system doing that, you know, this announcement they put out on Saturday, what in there is new that we weren't able to do on Friday? This is an important question. Okay, so within 24 hours, this was bad news for open AI. Within 24 hours, a mathematician who works for Anthropic announced that he had already gotten Fable, which is the sort of the cutting edge LLM that Anthropic currently has released. So, it's out there on the market. He had used Fable and he'd
9:10already reproduced half of those 10 problems, presumably without this sort of fancy math solving harness that open AI was using. More generally, as I mentioned, DeepMind has a system called Alpha Proof that has a very smart math solving harness on top of it that spawns agents that call LLMs to more systematically search through the space of possible solution paths. And it's a very smart way to leverage what it is that LLMs do best in math. They wrote a paper in May, they published it,
9:40that announced that Alpha Proof, which is, I don't know what LLM it's running on, but certainly one that is smaller and dumber than whatever pre-release LLM that Astra is using. They announced that it solved nine of the 353 open air-douche problems and 44 of 492 OEIS conjectures. So, 50 open problems. That was back in May. So, it's not as if we're now able to solve math problems of a type we couldn't before.
10:10It's just like, hey, they're saying we are continuing with math harnesses able to solve these type of problems like other 2026, circa 2026 models are able to do as well. Now, please don't jump on me as saying that I'm downplaying the importance or the difficulty of solving these problems. I'm just saying we don't have evidence that Astra has a new capability in this way, some sort of fundamental new capability that didn't exist last week. Hey, I need to take a real quick break here to tell you
10:41about the presenting sponsor that made this AI reality check episode possible. They're called Done Daily. They're an online service that connects you with a real coach that helps you build a custom productivity system designed to fit your life. The coach will help you actually get important stuff done. Look, this is not some AI agent or over-featured productivity tool. It's a real person working with you to cut through distractions, face your productivity dragons, and lock in habits
11:12that actually get results. So if you want to find depth in our increasingly distracted world, you need to check this service out. You can find out more at done daily.com. That's done D-A-I-L-Y.com. All right, let's get back to our episode. All right, so let's move on to our second
Significance for Mathematics
11:32big question. So now that we kind of know what happened, the second big question is one of significance. What does this mean for mathematics and mathematicians? I'm going to start on sort of with a bit of a negative approach, and then I'll move on to the more positive approach, okay? So if we ask, has AI solved mathematics now, like whether with Astra or if we want to combine Astra with Alpha Proof and Fable, like in general, are we basically at a place, as was implied in those tweets I read at the beginning of this episode, that like we're now going to just, we're able to
12:06basically like plow through all the math, AI, and then all the science, or whatever the big claims will be. No, we have not. Noah Brown, again, let's return to Noah Brown, who again worked on this project. He tweeted the following, we still haven't solved math. Astra isn't building new branches of mathematics or posing interesting new conjectures. We could also look to the mathematician Thomas Bloom, who was involved in verifying Opens.ai's effort to disprove the unit distance conjecture, and he said, hey, these are, yeah, these are good results. These are big news that they solved these
12:37results, but he still described those results as being based on constructions, which is what I was trying to imply before, is that, again, there's a certain type of math that these models are good at, in particular conjectures where you construct a counterexample or construct a positive example, or apply or generalize an existing mathematic result to get out of it a new bound, right, this sort of construction-based approach. He's emphasizing these are constructions, and none of these, again, are as big of a news as if, for example, we had actually proved instead
13:07of disproven the unit distance conjecture, which would have required a whole new argument and not just a construction or application of a tool. Thomas Bloom went on to heavily push back on the idea that systems like Astra would be replacing mathematicians. Here's what he said, it's not right to call proving one conjecture made by a mathematician using theory developed by over a century of work by mathematicians with an AI built by mathematicians and trained by reading everything ever written by all mathematicians as, quote, replacing mathematicians, end quote. A little bit of pro-human chauvinism there that I'm on board for. It's not doing new
13:42math, it's using our math, and mathematicians are running it and having to try to understand what it's doing. All right, so as noted previously, the right way to think about it is these LLM-based new math tools work on certain types of problems some of the time. Problems that have a certain character, they're usually discrete and have proof based on constructions or applications of existing objects. You can throw a lot of results at these systems, and you kind of look for the ones that it's actually able to make progress on. So this is not, as many in the AI is a force that gives us
14:14meaning crowd implies a sign that AI can basically do all math now and will soon also devour other fields of science. So certainly, I think those tweets from the intro are just dead wrong. This idea that we're in a golden age of science because there are some discrete math construction proofs we can do and others we can't, I think is just its sci-fi futurist, completely head in the clouds type of over extrapolation and not really that useful. The reality, of course, is, is, you know, AI progress
14:46is jagged. There are certain jags you can go really far on and other ones you don't make much progress at all. The right analogy to use here, I keep saying, is tributaries on a river. So you have these tributaries feeding into a river, each of them representing a different capability of potential capability of AI. Now to find out or exploit that capability, you have to mount an expedition down that tributary, which requires a lot of resources, time and attention and expertise. And sometimes you
15:17find the tributary, if you put enough resources at it, it's very navigable, which would correspond to like, oh, we get a big jag there. We're able to like build pretty impressive capabilities here with AI systems. A lot of other tributaries turn out to be not very navigable at all, either because we don't have the time, attention, or resources to explore them, or we do, and we don't get very far. That is the current state of AI. So we found computer coding as a tributary with a huge amount of resources and expertise at this, and we were able to make a lot of progress on that tributary.
15:48It's like the Hudson off of the bay. Like it goes really far, and that's a place where AI is doing well. We're having a lot of effort now put on, again, these sort of discrete math construction-based conjecture-proving or disproving, and we're finding, hey, we can navigate this. Maybe it's not as deep water, if we're going to continue this metaphor, as computer coding. It's not so like universal, like every mathematician now will only be using these tools, but it's like, I would say, a pretty well-navigable river. But the thing about all of this, here's the key thing,
16:19exploring one tributary doesn't necessarily help you explore others. You still have to go tributary by tributary, mount an expedition, and see where you get. Sometimes you can use the tools or discoveries of another exploration to kind of help get one going. Sometimes you have to invent these tools entirely from scratch. What they're using to do this mathematics is very different than what they're using, for example, for computer coding harnesses. It's just different types of training and systems. All right? So again, it's really the wrong way to think about this, of every time we make progress in one of these tributaries to say, we now are going to make progress automatically in
16:51all tributaries, AGI is coming. That's just not the way this works. That's more of a a Max Tegmark style model of AI capabilities as being measurable by some sort of like intelligence number that rises like a water level. And that you have these different mountain peaks where the height of the mountain represents the complexity of the task that humans currently do. And as the water rises to a certain level, it's covered all mountain peaks that require that much intelligence. And in that mindset, when you get over what seems like a high mountain peak, like working on discrete math conjectures, like, wow, that means AI can do everything that's that hard.
17:27And therefore, as this water level rises, soon there'll be no mountain peaks left and we can do everything. Again, that's the wrong model. It's tributaries on a river. So another way to look at this is we've been exploring this river really with hundreds of billions of dollars worth of resources and untold hundreds of thousands of hours of human effort. And we're still relatively limited in what tributaries. We finally, after multiple years, made the coding tributary work. We're getting some navigability on discrete math conjecture, proving, disproving. Certainly, there's some
17:59tributaries involving the production and processing of text that have been proven navigable. And that's kind of it. Like, it's proving really hard to find the new tributary. And we're not seeing the negative results of how many of these that we have tried to explore. So I think that's the right way to think about it. So if we want to step back, let me put on the positive hat now. This is, for people in those fields of mathematics, I think going to be cool or at least interesting. I think the field
18:29could use a shaking up. In fact, let me make a series of predictions and put on my positive hat now. So please don't tell me that I'm trying to downplay all AI stuff. I'm just trying to be realistic. Let me put on my positive hat here. Here's my predictions. I think in the future, as these tools get more usable and cost-effective, they will significantly improve the quality of research in certain mathematical fields in the sense that the depth and quality results per paper will go up because some of these, like, disproving or applications of some existing results that might be hard or be tripping people up can be really helped by these tools, which allows us to keep
19:02going forward. Case in point, within a week of the unit distance conjecture being disproved from that OpenAI chat transcript, a human mathematician sort of took the general structure of that proof and was like, oh, I can now, I can make this into a much better paper. I can extend this. I can improve this. I can make it more elegant and I can find some other applications, right? So it's this, like, symbiosis of you're able to get unlocked by an LLM-based math tool and then that
19:33allows you to apply your reasoning to extend that further. So I think we're going to get better quality papers in the areas where these tools apply. I think these tools will be integrated pedagogically in the graduate level. You're not going to learn how to use these tools ever at the undergraduate level because you can't use them unless you've mastered all the mathematics. So undergraduate and intro graduate education will still be on understanding the math so that you can then use these tools, right? It's not like vibe coding a application, you know, a JavaScript game where, like, the game works. I don't care how the code works. None of these tools are usable if you don't
20:04understand the underlying mathematics. So I think they will be taught. It'll probably be at the graduate level. I think we'll quickly exhaust in the various fields where these tools exist. We'll quickly exhaust the obvious one-shot results, you know, where, hey, we just kind of fed at this a few times, a few different ways, and got an answer we can publish. But we will get good at learning where they're useful and where they're not, where we're going to waste our time, and where they might actually help us make progress. I think more subfields of math will come in play. In particular, as these harnesses get better, the orchestration layers, right now we're kind of
20:35in combinatorics graph theory. There'll be other areas where we're going to make progress. I think that's going to be exciting. I also think we're going to see a shift to the focus more on the harness, and we're going to get away with lower cost LLMs, maybe even like open weight LLMs that are small or research LLMs that have been tuned on specific areas of mathematics combined with a really smart harness. That's probably going to be the combination, not that we're going to be using some trillion parameter Fable-style model to do our math, because too much of that training and size of Fable is dedicated to things that do not help us solve math. We need these things to be
21:10very cost-effective for them to have traction in the world of mathematics, because mathematicians have no money. Any grant dollars we get go towards paying for our graduate students. We do not have big lab or equipment budget, so we need these things to be something we can run in the server closet in our department, but I do think that's imminently possible here. All right, so I think there's a very useful mathematical story here, but that broader story that math has been solved or science has been solved, I think is just ridiculous. All right, question number three, what does this mean
Implications for OpenAI
21:38for OpenAI? Now, I want to reiterate here a point that I first made when they announced the unit distance conjecture result from a few months ago. I continue to think these announcements are bad news for OpenAI. As I've said before, construction-based conjecture proving and disproving in certain subfields of discrete mathematics is an incredibly narrow field with no economic upside and minimal societal upside, right? This is not like you're curing diseases or creating new materials or figuring
22:10out more like cost-effective ways of scheduling airplanes in the sky or something. It's incredibly abstract. It's the definition of a sort of like often dead-ended narrow basic science field where sometimes we can find applications for these results, but more often it's just knowledge for the sake of knowledge. Now, people are saying, yeah, yeah, sure. These are like obscure, you know, extremal graph theory disproving. Disproving an extremal graph theory conjecture is not economically valuable or is it societally valuable. But, and I've seen this on the OpenAI announcement as other
22:44places as well online. They'll say, yeah, yeah, but we're training research agents on math first to prepare them for tackling more valuable research endeavors. That is a dumb statement. You don't train a system on extremal graph theory to predict, you know, prepare it to conduct drug research. You would train it on drug research. The reason why they're focusing on math results in these relatively esoteric areas is because those are the results these systems can solve. If they had a more useful or lucrative application, that's what they would be crowing. They're trying to prepare for a trillion
23:20dollar IPO. You're not going to excite investors talking about Ramsey numbers. It's just what these systems can do. It's like when you ask John Dillinger, why do you rob banks? And his answer was, because that's where the money is. Well, it's the same thing. Why are you solving extremal graph? You know, why are you disproving extremal graph theory conjectures? It's because that's what LLMs could do. That's where their current capabilities have to lie. This is also bad news for OpenAI because they didn't do anything new here. As I mentioned, immediately we could get most of those results
23:53out of Fable without the special harnessing. And DeepMind has been solving, they solved 50 major open problems a few months ago using a cheaper LLM, right? They're more focused on the harness, not the LLM. So what is it that was new here? I mean, really this announcement is like, we're focused, you know, we continue to validate what we've learned over the last six months, which is with the right, you know, harnessing and training, you can use the 2026 era LLMs to work on, you know, make progress on important problems in certain areas of mathematics. We already knew that
24:25on Friday. The announcement on Saturday didn't change that. And so this is why I would be concerned about this from OpenAI's perspective is like, I'm excited about that people are continuing to work on this because I care about that field of mathematics. But this wasn't something new. And this thing that we've been able to do recently, I think is exciting, but is not nearly as generalizable as the AI is a force to give us meaning crowd would have us believe.
24:52I think there's a reason why Anthropic doesn't talk a lot about math results, except for like, they'll occasionally just announce like, oh, by the way, we solve something big too, because they kind of annoy OpenAI because they know that's not economically valid. They like to put, they were winning in the computer coding agent wars, which is more economically lucrative. And so they like to, they like to focus on these types of things. Like we're generating this many billions of dollars from people using our coding agents, for example. They'll focus on cyber, at least they have a case for cybersecurity. This could be very useful for protecting your own systems, right? Like
25:22there's an economic case there. They're not spending that much time talking about solving graph theory conjectures because, you know, it's cool, but we've already know we could do that. And it's not, it's not the thing that's that valuable to us right now. So I think OpenAI was hoping that the AI is a force that gives us meaning crowd would, as they tried to do, take this announcement, which is again, no different than what we were able to do last week as well, and make that seem that there is a sort of a general in the air, ambiguous, just vibe of like, AI is getting smarter, everything's going to be solved soon. And I don't think that's a genuine, that's, that's somewhat
25:54disingenuous, um, that marketing is good. And I think that's inaccurate. So here's my final summary. All right. And again, I really have to say, because people keep saying that I'm like anti AI or don't think it's impressive. I'm not, but I am anti this approach, this thinking approach, this eschatological approach to AI, that every announcement we have of any sort of feature or interesting thing that AI does has to be taken as evidence of a Godhead is coming in the world as we know it will be different.
26:27I know, I know, I know for some of you that gives your life meaning. It's more interesting than the world you're in now, a world in which everything has been disrupted by AI. Sci-fi and futurists have been thinking about this since the eighties. It's exciting, but it's also exhausting and frustrating for everything else. Not every announcement needs to be tied to this. Maybe I was wrong last time, the time before, the time before, the time before, time before, time before, but this one for sure means now we're on this like fast slope to, uh, our entire world has changed and the digital Godhead
26:58is going to be a source of either meaning or destruction, which either way, it's more interesting when what's happening now. I'm tired of that way of thinking. Can this not just be, hey, math nerds, we're starting to get innovations in math similar, but at a smaller scale to what we saw in computer coding. I think it might make math more interesting. Math is honestly, in my opinion, has been in a bit of a rut for the last few decades. It's, you know, I've, I've done my share of it. This is something new and we're going to get better results and shake things up and more creative results. Isn't this exciting, math nerds? And everyone else like, just trust us. It's exciting for math nerds. We don't
27:30really have to understand it. Why can't we just have one announcement that we think about that way? Because that's the right way to think about what's going on in math. LLMs plus math harnesses are really going to improve certain fields of math. Maybe lots of fields of math, maybe not as completely as in computer coding, but these changes will be major. And as someone who is adjacent to these fields, I think that's really cool. But as an indicator that we're somehow, uh, have crossed an event horizon into the digital Godhead that we can now worship. I just think that's not right. And it's vibey and it's hypey and it's not useful. I hope we keep working on these AI math tools
28:07because mathematicians will like it, but no one is looking at this and saying AI has solved science because if it could, we would be doing things that are actually useful. We'd be doing things that actually make money. So let's keep helping discrete mathematicians. This is cool. Everyone else, if you're not on discrete math, I would say, uh, carry on. All right, that's it for this week. Um, as always care about AI, but not everything you read about it. Hey, if you made it this far, you must be ready to join my fight for depth in a distracted world. Now, the best way to do this
28:38is to join over 125,000 people who receive my email newsletter each Monday. You can sign up at calnewport.com slash ideas. And when you do, I will send you a free guide to my seven best ideas about cultivating a deep life. Sign up today. Calnewport.com slash ideas.
29:24Calnewport.com slash ideas.
More from Deep Questions with Cal Newport

Classic Episode: How Do I Learn Hard Things? | Monday Advice
Aug 3, 20261h 15m

Did OpenAI’s Model “Go Rogue”? | AI Reality Check
Jul 30, 202633 min

Why Do Digital Detoxes Fail? What Works Better? | Monday Advice
Jul 27, 20261h 10m

Am I Optimizing Too Much? | Monday Advice
Jul 20, 20261h 18m

Does Claude Have Private Thoughts? (Everyone Settle Down) | AI Reality Check
Jul 16, 202631 min