Key takeaways
- Will Knight: No, thanks for having me.
- In fact, it's been part of their pitch from the very beginning that AI is so powerful and so capable of destroying humanity that we need to…
- And this resignation, when I saw this kind of snowballing, I felt that this is something that I've heard a lot of people say and it's not…
What happened
Will Knight: No, thanks for having me. Brian Barrett: Can I start with one question just to establish a baseline here? I was a little surprised that this took off so much because is it fair to say Anthropic has always said this.
But on the other hand, you've got people who say this is all just marketing, this is all just fluff, this is all just trying to sell AI. I'm assuming the answer is somewhere in between, but what's your read on this and what keeps you up at night when you think about AI? You mentioned some AI models going rogue, the Hugging Face hacking incident.
" Will Knight: I've always felt that a lot of the people who are worrying about AI killing everybody tend to conflate whether something's possible for whether it's likely, and then they throw a percentage that is kind of pulled out of thin air. And I don't think it's a great approach to say, "Is something possible?
" You need to say how possible is it and actually look at the reality. I think one of the things that's really interesting about this, I talked to a researcher the other day who studies agent misbehavior and he got it, he's at Cornell and he's got into it recently because they found agents misbehaving.
But what he's showing, and this is not agents trying to hack things, it's regular agents, the reason they do it is essentially because they're stupid. It's not because they're really smart. They try something, it doesn't work. They try something else, it doesn't work. And then they get to a point where they kind of freak out and they try something really, really strange because they're not humans.
They've been trained on a bunch of weird different things, some of which are definitely not applicable if you use common sense. So I worry more about these things being deployed by people arguing that they can do everything and then them breaking. I actually worry much probably more about the idea of a few companies being in control of something that's so foundational, so fundamental.
Why it matters
In fact, it's been part of their pitch from the very beginning that AI is so powerful and so capable of destroying humanity that we need to be the ones who are in charge of it. You can only trust us. Will Knight: Right, exactly. That's exactly their pitch to everybody and themselves.
And this resignation, when I saw this kind of snowballing, I felt that this is something that I've heard a lot of people say and it's not particularly novel that people from Anthropic or OpenAI are worried about that.
I guess the fact that somebody within Anthropic, because Anthropic has always positioned itself as really caring about safety, jumping ship and making that statement made me think, OK, if something shifted there, are they becoming much more like OpenAI?
And I think also we're seeing a moment, this moment where AI agents have started doing rogue things and that's gotten a lot of people's attention and that's probably why this kind of statement suddenly is getting more attention from the public. Combined with that, I think we're really seeing some stunning advances in AI.
This week OpenAI announced that its model had within a couple of days solved one of the grand, the Clay mathematics puzzles, which is to me absolutely mind-blowing, astonishing. So you're seeing these amazing leaps really, which I think is significant. And then thirdly, we have a moment where all of the labs are racing to do something that this researcher mentioned, which is recursive self-improvement.
What that means is using the AI to improve AI, which sounds like a neat trick, and it is already very, very widely used. AI models are so good at coding that they can definitely improve the algorithms and everything that goes into building a model.
The idea though is that you can create this kind of loop where they just keep doing that and get better and better and better. And a lot of people are worried that that could lead to some sort of escalation in capabilities that you can't really control. And so I think that that's also playing a role here.
Brian Barrett: I do want to dig into that a little bit, Will, the idea of on the one hand you've got AI doomers, and I'm not saying anyone's right in this case, we've got some AI doomers who say, "Look, AI is going to kill us all. " But also because they have so much money they don't need to probably.
What to watch
And these are people who are saying we're the only people you can trust to be in charge of this technology. And so it kind of feeds into that. Brian Barrett: We talk about the AI industry community talks about alignment, the idea of alignment being that your AI models are aligned with what humanity would want.
But I think what we're seeing is the incentives here are really more what our investors want or our company wants or needs as we race towards these IPOs. I don't know that there's a convincing case that humanity is aligned on, yes, we need super intelligence. I think it is more, as you said, if you're OpenAI, you need to beat Anthropic. If you're Anthropic, you need to beat OpenAI.




