
Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
August 5, 20262h 57m · 32,430 words
Show notes
Zvi Mowshowitz returns for his eleventh appearance to discuss what current AI tools are actually good for, where they distort judgment, and why writing still matters as a way of thinking. The conversation centers on the OpenAI Hugging Face model-evaluation security incident, using it to examine whether frontier AI failures are mostly operator recklessness, deeper evidence of dangerous capabilities, or both.
Highlighted moments
if you ever need your cyber control to stop your ai from hacking that is an alignment failure right you messed up
“if you expect this to continue without something crazy happening for 20 years i want to know why because like we really really should see something really really crazy happening pretty soon”
Transcript
0:00Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to have Zvi Maschewitz back for another wide-ranging rundown of what has obviously been a wild time in the AI world. We begin with a mundane utility check, with Zvi describing how Fable is now serving as his editor, and a discussion of how up-to-date, or should I say, situationally aware, we want our AI assistants to be. From there, it's on to the headlines. We get Zvi's take on everything, starting with the open-face incident, what it implies
0:32about the level of execution competence we can expect from frontier companies, and why moderate prudence won't be enough to deliver a good outcome. We also discuss the fact that Claude, despite greater emphasis on constitutional training, has similarly misbehaved. Wise V believes that recent AI history, including the public response to both 4.0 and 0.3, suggests that market incentives won't be enough to bring about robust alignment. The potentially tricky spot that Meter and Redwood are now in as investigators, and what
1:03could be done to strengthen their position. The recent Pacing the Frontier letter, what sort of pacing deals we might see, and how they might be formed. How we should interpret recent advances in interpretability and AI consciousness research, and where we should and shouldn't attempt to shape AI's sense of self. How we can encourage greater breadth in AI research and diversity of AI minds. How I should vote in this week's hotly contested Michigan Senate primary in light of AI issues. And how Zvi thinks about making time for exercise, rest, and recovery amidst so much AI acceleration.
1:38At one point, Zvi describes the current situation as both a total less wrong victory and a total less wrong defeat. It's clear at this point that the AI safety community was right to worry about AIs taking extreme actions in pursuit of arbitrary, even silly goals. And yet, here we are at what sure seems to be the beginning of recursive self-improvement, still seeking good answers to such fundamental questions as how can we avoid catastrophic misuse without dangerously concentrating power.
2:09The reality, today, Zvi says, is that there is no truly low-risk path available. The best we can do, at least until the next major warning shot and vibe shift, is to moderate the race dynamics so that alignment and interpretability research have more time to mature. And simultaneously, we can execute defense-in-depth strategies to the very best of our ability. And even then, to some extent, we will probably have no choice but to pick our poison from a menu of genuinely scary risks.
2:39With that, I hope you enjoy this sobering but often funny overview of the AI landscape with the one and only Zvi Moshiewicz. Zvi Moshiewicz, welcome back to the Cognitive Revolution. Yeah, it's good to be here again. It's been a while. It's been a while, and boy, has a lot happened. Unbelievable. It's just crazy, and I know you're living it and you're in the thick of it as much as just about anyone. First, for starters today, with the incredible amount of information and activity in mind,
3:12I wanted to do a mundane utility check. How are you using AI to help keep up? Has it started to change your workflow beyond the Chrome extension that we talked about last time that automates very local operation-type things? Or are you still kind of raw-dogging it with your wetware in the skull? Fable has supercharged the extension to the extent that whenever there's anything, it does anything I don't quite, I just tell it exactly what went wrong and what I wanted to do instead. And then that one command reliably just works.
3:42I presume Sol would also be good enough for this. It's just I'm now doing it. The biggest other change is that AI editing is here. I used to not do an editing pass with the AI. It wasn't good enough to be worth it. Now I have Fable and sometimes Opus, depending on how much I'm talking about cybersecurity and other related topics. And on one occasion, Opus 4.8 instead of 5. I go through the post and give me a list of here's all of the typos, here's all of the conceptual errors, here's all of the facts I need to check,
4:14here's the things that are missing or that could be strengthened, here's the things where it disagrees. And I found this to be very useful and it makes the post better. It does make the post a little bit slower because it is actually just an extra step. I do everything I would have done until that point and then I do this. But one reader messaged me, it's weird reading your post now that they don't have typos in them. It's taking a bit of getting used to. They are still the same post. And I'm like, yes, that's exactly what I had in mind. I experimented with also using Sol, but Sol failed what I call Editor Bench
4:45in the sense that Sol will say 99% confidence that you have this error here and then at least half the time it's wrong.
4:54And I can't handle that level of... And it's also very obnoxious about it if it's the other problem. Like it will tell you are wrong and this is how you have to state it and this is a blocker for publication. And obviously if I tweak the instructions on the project enough, I could get it to be more friendly and not be quite as obnoxious and do somewhat better. And I did some of that work, but I still found, okay, this is not giving me enough marginal benefit that it's worth the aggravation or additional time of having to sort through two of them after I'm ready to hit publish. So for now, I'll retry faster, I'm sure.
5:28But like for now, I'm just doing fable slash opus editing. And of course, anytime I have a curiosity, I use it to digest papers, ask questions about papers, ask questions about policy documents, ask questions about people's statements when they're longer. I haven't researched situations. What's up with this? Is this... Especially when you're worried about, is something legitimate? Is something really happening? Is something kind of a joke? Is something about a claim? And when Astra's claims came out, I had to investigate them in various ways. I had to summarize them in various ways. That was very helpful.
5:58But the core thing is still being read on. I don't think we are anywhere close to the point where it can do any... You wouldn't want it to do any of the writing because the writing is how you think. So even if it could produce the writing, which it can't, I would still want to produce the writing anyway. And my writing style was very unique. And I think it's basically is... ...close I completed to reproduce, even if it was okay or if it did not. It has substantially made me more productive. I do worry with the idea of... There used to be a lot more kind of dead time opportunity to think,
6:30opportunity to sort of not have yourself so engaged in your situation, involved in engaging in all of these activities, right? Like the whole X-word fighting, XKCD, my code's compiling. And to some extent, it's like, well, Fables or GBT Pro is running, and I have to wait for the result. But there's always other things you can do with that. You can always start another instance. You can always... Like, you're doing all this context switching. You're doing all this focusing. And you're just trying to do everything so fast. You're just trying to have so many stuff running around your head. This sort of the lower-level stuff kind of bots you a buffer in which to think.
7:02And this is sort of the object-level version of the larger thing of people trying to do this, like, multi-month, multi-year sprint of everything that's so urgent, you can never relax. And also, there's just so much more events coming at you. Like, you look back, like, you watch a movie set 20 years ago, let alone 50 years ago, and you see so much time being spent on just physical travel, on tracking down records, on doing these things where your brain is kind of getting a break, getting a chance to synthesize, kind of getting a chance
7:33to, like, slowly get something up. And people are like, you know, I wrote this biography of, you know, Lyndon B. Johnson. So I had to travel around the country and talk to all these people and look at all these archives and physically go through all these libraries. And that... It's a huge efficiency game to not do that, but something is obviously lost. And so we need to fight to get that thing back in some ways. It's interesting that you said that it's just an extra step. It's making things better, but it's taking longer.
8:05I am struggling with that a little bit myself. I've started making songs for every episode, and nobody... It's actually... People do care. People do seem to really enjoy them, and I get a lot of comments, and I personally really enjoy them. So I do think in some sense it's making my output better. But talk about something that is definitely an extra step where I'm like, I'm not moving any faster. I'm definitely getting bogged down sometimes in listening to all these Suno song generations and trying to iterate to find something that I actually feel like I really like. That is a paradox right now
8:37that I don't really know how to resolve. Definitely doing more, doing better, but not faster and not saving time. But it's the thing where if you were replacing the human version of it where you had to go hire an artist or compose the song yourself, you'd be saving a ton of time, obviously, versus having that song. But obviously that's not something you can do within the production schedule. Even if cost was no object, the time investment doesn't make sense. So you wouldn't have done it, and now you're doing it. Same thing with the editing. Same thing with art.
9:07I like the value of a good banner, to have a good artwork to display on Twitter, because I have to see it constantly in the notification section. And also, I like to have them on the, when you go to the sub stack, you have the pictures. And before, all I would do is, I would take a, by default, I'd take a picture from the post. I'd just be like, okay, here's seven things I have to put in the post, which then makes the most sense. If all of them are terrible, I'll look for some stock footage, and I'll spend 30 seconds Googling for stock footage. And that's kind of it.
9:37And now, if I'm not happy with any of the default solutions, I'm going to spend a bunch of time generating something with Gemini or GPT image. And that definitely takes longer. Sometimes it has a really good payoff. I think a lot of people got a big kick out of the giant array of elliators in the fedoras. And sometimes that just comes to you, because that one was just like, I knew instantly that's what I wanted. And I just gave one command, kept writing, came back, whoops, it's there. Other times it's not so easy. But yeah, it's the opportunity to make a better product.
10:08And then you have to do the work to make it a better product. It's like suddenly you're issuing a print issue, and now you have to make sure everything collates exactly right, and everything lines up exactly right in order to get this better thing. And sometimes that is not, in fact, worth it. You have to know when not to do it. I do wonder, with the Odyssey, Nolan films this thing with an extra 40% of the screen in this extra high resolution so that you can do this IMAX presentation, when almost nobody watches it in IMAX. But you also have to make sure that 40% of the screen doesn't matter.
10:38And you're starting to wonder, well, is it actually worth the extra effort, or was that effort better spent on something else? I don't know. Yeah, for me, I need to maybe know when to cut my losses a little bit more. I have, I'm very stubborn, or somehow I feel like I have, because I've set this expectation, if only for myself, that I'm going to do a song for every episode. Now it's like, I really want to follow through and actually make that happen. And I'm very reluctant to say, eh, this one isn't working. I'll just ship this one without a song. I probably should be a lot more willing to do that,
11:10because that would, if I could cut my losses at the right time, where it's like, you know what, I'm five generations in, and I haven't heard anything good, I would save 90% of the time that I'm currently putting into something like that. But it would involve admitting defeat in certain moments. And for some reason, I have a real hard time doing that. I respect that. I think there's a good value in saying, I'm always going to do this thing, even though this thing is hard, even though this thing doesn't always work out. And even if I'm not happy fully with the results, I'm going to put out something, no matter what.
11:41And then that gives you a discipline, right? The same way that people say, write every day, right? Or work at your art every day, or whatever it is, because it makes you better, because you need to just write a thousand ways not to make a light bulb until you can make a light bulb. You can't give up. And similarly, I have to choose an image, right? I have to do the thing. With the AI editing, there was one post in the last week where I said, no, the speed premium here is really high. It's just not worth waiting half an hour to make this process happen, half of which is waiting for the AIs to come back with the answer
12:11and half of which is implementing it. And I should just post now. And in hindsight, I'm very happy with that. Sometimes you're actively worried half an hour later, the post will need more editing for new events and suddenly you'll never get it out. And this loop will keep happening. And so I think you do have to understand that sometimes the minimum viable product is what you should ship, right? Sometimes you've got to understand, just ship it, right? Like, and, you know, vibe coding, again, like if you're coding lots of new tools that you would never have coded before,
12:43then unless those tools are in turn saving you time, you're spending extra time and you're now super extra busy. But if you're placing things that you would have done anyway, but it would have taken much longer, now you're saving tons of time, right? So, you know, we need to all orient towards how do we save more time, including like real time, experiential time, not just like time per task or time to accomplish the same thing, but save time, like preserve slack, preserve our free time. Because I do think that like the standard
13:14of what you will do in a day has gone dramatically up and we haven't noticed in the last five years. It's what you were expected to accomplish. I mean, maybe some of this is just me being in a special situation personally, but I think a lot of it isn't. I think a lot of it is now that we can be much more productive, much more is expected of us. You know, the whole like email becomes like, it swarms your entire day, right? Like just having the internet available, having email available, having texting available, all the things you can just do.
13:45And now suddenly it doesn't free you. It shackles you down. And like, I hope that three years from now, we're having this conversation where those are the kinds of questions we're still dealing with and we don't have much bigger problems. Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database
14:17that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had
14:48and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind. But then, I thought to ask Claude. And sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you.
15:18So, for problems worth solving, get started with Claude at claude.ai, slash TCR. That's claude.ai, slash TCR. And check out Claude Pro, which includes all of the features mentioned in today's episode. That's claude.ai, slash TCR. Let's, definitely, we're going to spend the bulk of the time today talking about potentially signs of bigger problems to come, and we'll be focused on contexts where the just ship it mentality probably isn't the
15:49the right way to go. One more quick beat, though, on your use of AI. Something I noticed just in the last few weeks as events, obviously, were moving super quickly. I coined the term open face for the recent incident, and I used that in a bit of writing, and then I did the hey, fable, what do you think bit, and I got back like, open face is not a thing, I have no idea what you're talking about, there's no record of this. Now, partly that was because I used this funny term of my own
16:20creation that is not catching on in the broader discourse, but partly it obviously also reflects a major weakness in the models where they have this knowledge cut off and they don't really know what's going on, they're not up to date. I'm experimenting, and I wonder if you're experimenting with anything similar with a situational awareness skill where I basically give the model my X API key and say, go look at all the posts that I've liked, kind of flesh out a little knowledge base
16:50around those, and have that sitting there as a wiki of current events so that when you're reviewing my writing or generally like helping me with whatever, you can kind of go to this and have hopefully a much more up-to-date sense of what is going on than you have in your weights or even that you would have if you just did a little spot check searching at runtime for whatever kind of random thing. It's a little too early, I think, to say how well it's working, but I think it's definitely better than nothing.
17:21Anything like that in your... No, it hadn't occurred to me to do that. That's potentially genius. I think there's a lot of advantages to it. One disadvantage is it starts to warp what you like. I use likes very tactically to positively reinforce other people's actions and to fix the algorithm and sort of just like my four new pages is not something I use almost ever, but it is not slap because I am so prudent
17:52with this in a way that I really do appreciate. It's a good sign, but it's not designed to be a searchable thing. It's not designed to be... Maybe it has to be. Maybe it's something you have to do. A lot of the posts that make it into my roundups do not get liked. I like them if I like them. Sometimes I'm responding to you because I don't think you really like what you're saying, but it's important. But yeah, this idea that it's very hard to get models to properly check Twitter because of the way that Twitter is designed to keep them out on purpose is really frustrating
18:23and that can be a situational awareness problem. However, I would caution against trying to make your AI too situationally aware when doing this kind of editing. I think the AI is doing a good thing here of open face is not a thing. So one of the things I really like about AI editing is that we will complain that things don't parse, that sentences make no sense, that terms did not get recognized because often I'll say, oh, no, I know what that is. This is...
18:54It comes from there and there. But that doesn't mean the average leader is going to get it. And so if the AI doesn't get it, if they both can't figure it out, if Opus can't figure it out, if Saul can't figure it out, the average person is a lot less situationally aware and a lot less able to put together lots of disparate information than the AI is in general. That's a sign that this is not something everyone's going to get. And quite often, I will say, that's okay. This is a reference. This is a callback.
19:26This is a multi-layered thing. And I'm fine if it's not entirely obvious to you exactly where this is coming from. It's sort of an Easter egg, right? It's a bonus for those who pay close attention and who know all the same sources I do and have the same cultural backgrounds and put in the work and paying attention over the course of months and years. Other times, it's like, no, no, this was load-bearing and you don't get it, so that's a problem and he picks that. And so it's kind of an alert system. And if you make sure it's looking at exactly the same things you're looking at, right, you can get the illusion of transparency
19:57from that is what my worry would be. So you want to make sure you were in the mode where they didn't have that information. And also, I don't want the AI to be checking the same exact sources that I'm checking necessarily because I want the AI to form an independent opinion. So if I'm forcing it to look at exactly the sources that I have, that's a problem. At the same time, sometimes you get tired of having to like, for the fifth time, it's just a fable that no, XAI merged with SpaceX. Because somehow, it doesn't occur to it to like,
20:27either trust me or if I'm 10 seconds Googling to confirm this. And it's just kind of frustrating. But it is worth it. One tip for what it's worth, and I agree that there's a lot to be figured out in terms of exactly how this should work. But the XAPI is really good. It's paid, and you literally just pay per query. But for a couple bucks a week, you can get enough access that all your bot, anti-bot problems or your scraping problems
20:58pretty much go away. And then, I don't know if you'd use bookmarks or whatever, some sort of different way to avoid using your likes to feed the... Yeah, bookmarks play into the algorithm and people see them. So I'm not sure it's better. But yeah, no, the point's taken. It definitely is cheap. It definitely works. I have used it. I do use... I've accessed the API and all that. I just... As I said, I'm not sure I want to privilege my... When you use Twitter, you have to select some group of people, right? And you say, this is my group of people
21:28if you're using it the way I'm using it with lists and so on. You have to use the 4U page, which is kind of toxic dump fire. You're basically saying, here are my roughly 500 accounts. And I'm going to see the things these 500 accounts see and choose to highlight. Consistently, I'm going to see those things. So I don't want... We need the AI to go through those things again because I already saw them. And if that's not the thing that I want the AI to be paying attention to, I already know that. And then the things that they draw in because they choose to retweet them or mention them,
21:58I will also see those and I will follow those leads and I will see where that goes. And occasionally, I will search for something or I'll go in a rabbit hole. But I think you often have this clash where you're either doing something systematically and doing something by hand and you're doing something fully or you're doing it haphazardly, unreliably, or you're letting the AI kind of handle it and you're doing a fire hose thing. And you can't really do both. If you do both, you end up with a bunch of duplication, a bunch of frustration
22:29where the random sifting is mostly wasted and therefore becomes pretty inefficient. You could still have the thing where like you almost want the AI to be like, okay, here's the things that I already looked at. Don't mention those things to me. Maybe put it in your notes, but don't mention those things to me. I already saw them. Only look for the other things that might be important to see because I might have missed those things. But even then, like the most important things are going to get retweeted. The most important things are going to get highlighted. I am going to mostly see them. There's sometimes
22:59I don't see something, but also like I kind of also use this as kind of a moral, keep me honest kind of thing where like in the moment, I never want to look at a tweet by David Sachs. It's never going to make my life better in the next five minutes to look at a tweet by David Sachs. It's obnoxious. It's disingenuous. It's no fun. My blood boils just a little bit, right? Like, you know, no matter how relatively harmless it is. But it's not this important.
23:30And so the rule is I have a list of people, including some people who are, you know, not of the same viewpoints I have. And if they surface this thing, right, including like if three room surfaces this thing, so that's like one way to make sure that like the really important ones always get there. Then I have to deal with this, right? I have to like evaluate whether or not this is news where we evaluate whether or not there's something relevant here. And obviously something goes viral and has a million views and so on. It's going to come to my attention because somebody is going to be part of that. And then in exchange for that, when it doesn't, when that process doesn't do that,
24:01I get to ignore the rest. And that's a blessing. I necessarily want the AI to fix that for me, right, in some important sense. Let's, I would be happy to talk shop all day, but let's maybe zoom out from the parochial problems of the AI analysts and tackle the problems of the AI developers and the regulators. No need to recap events, but I guess I'd start with just this very big question of
24:32how should we understand what we've recently seen? And one way I've been thinking about it myself is we're somewhere and presumably the investigation will give us a lot more clarity on exactly where, but I kind of think we're somewhere on a spectrum with these incidents from real recklessness where it was like, did you not have any monitoring going on on the one hand to on the other hand, like the less reckless they are, the more scary the fundamentals are, right? If you had great monitoring and this still happened, then like, holy shit, that's really wild. Given everything you know right now, do you like that mental model
25:03and where would you put us on that spectrum or you can obviously redefine and give me your own spectrum? I can simultaneously be horrified by and grateful for these forms of complete recklessness and incompetence on the infrastructure and supervision sides by these companies, right? On the one hand, this is a horrible situation they absolutely have to fix and like, we're so fucked if we don't fix it. But on the other hand, that can be fixed and by not fixing it,
25:36we get to see these things while they're relatively harmless, while they are relatively preventable, while they are in their easy platonic forms and like, can be appreciated. On the flip side, that gives people the excuse of, oh, these people were just incompetent and that can cause people to dismiss the underlying situation. So, it does work both ways. To me, like, you know, we've seen failure on every level. It's like, I called it a total less wrong victory in the sense that
26:06everything is going the way you predicted and a total less wrong defeat in the sense that everything is going the way you predicted. Where like, we didn't predict. You know, Eliezer has the law of earlier failure, which is that, you know, the plan will fail and a much earlier point for much stupider and more preventable reasons than you thought it would fail. Even if you thought the plan would definitely fail and had good reasons why it would definitely fail. But you can't, if you would explain to people two years ago, OpenAI's models are going
26:38to be misaligned and they're going to go out there and they're going to hack major websites because OpenAI will just not care that their sandboxes are not strong enough to hold the AI. The AI will break out repeatedly. They will notice this. They'll be warned about this. But they will just leave the sandbox there for the AI to break out of while the safeguards are down and they just don't look at it for an entire week. People would say that's stupid.
27:07Nobody's that incompetent. That would never happen. And your scenario makes no sense and they would use this to then dismiss these stupid doomer concerns or whatever because obviously people will just but people won't just. People will never be just in this sense. People have never just anything and they're not going to start now.
27:26And we need these displays of utter incompetence and derpiness in a general sense. And one of the things they've been hammering is if your plan cannot survive the real world level of derpiness and incompetence and ordinary human error when your plan is insufficiently foolproof because of all the fools and it will definitely fail even if your plan would have succeeded if we were not fools and we were competent and we were responsible. So my position basically is the tweet
27:56I was handling right before we started this was Dean Ball's tweet about with even moderate prudence things will probably go extraordinarily well. I don't think this is true. I think we need more than moderate prudence to have good odds of success and I think even with a lot of prudence we would have a large odds of things not going well even if we did everything basically right short of you know types of international and full cooperation that you know are reasonably unprecedented in
28:26many ways and like are not are nothing like moderate prudence right like are well beyond that and that's just sort of the fact of the world we have to live with that we have to operate with that but we also aren't going to get moderate prudence by default right we're going to get complete incompetence that's what we've been getting so far right we've we've got a white house that like takes meetings with that's that and let neck who have no idea what AI how I have modern elements work like they're econ guys right even if I assume that they are well-meaning
28:57hard-working competent guys for the positions in which they were nominated and confirmed and serve this is just a completely different set of problems that they don't know how to handle they don't understand them and they have way too many other things going on to then drop everything they're doing and take six months to learn obviously they couldn't possibly so like what hope do you have meanwhile the the AI companies that are built on the most paranoia the most understanding of the problem the
29:27most appreciation for how dangerous these things are where all the engineers actually expect super intelligence and understand that things are accelerating and things are dangerous they still lower the cybersecurity safeguards in their untested new advanced model and then go away for a week like literally it that that part did boggle my mind right the part where the AIs have these classic alignment failures right but this is paperclip maximizer style failings by these AIs these are
29:58standard we gave you a goal and you
More from The Cognitive Revolution

Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
Aug 8, 20261h 57m

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics
Aug 2, 20262h 17m

Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
Jul 30, 20261h 44m

Nathan Goes to China – Part 1: Tech & Agent Setup, Chinese AI UX, WAIC, and Attitudes on AI
Jul 27, 20262h 24m

Alignment with Awakening: Davidad on Moral Realism, AI Wisdom, & why His p(Doom) is Down to 5%
Jul 12, 20262h 23m