Maybe we should remember how we've dealt with something else that threatened to exterminate humanity. Nobody in the Fifties or the Sixties or even the Eighties ever imagined that we could go the eighty years since Nagasaki without a single nuclear weapon being used by anyone, anywhere. And yet that's what happened. (And that's with nukes in the hands of such sane, reasonable people as Stalin and Mao and Kim, not to mention Nixon and Trump and Putin and Netanyahu.) If we could get nuclear weapons (which, unlike AI, are only useful for killing or threatening to kill massive numbers of people) under control, maybe that could inspire us to believe that humanity is capable of doing the same with AI.
We're capable of it, but we need to work harder at it than we did with nuclear weapons. After all, the only use of a nuclear weapon is to kill people, and no private entity has been allowed to develop them. In contrast, AI is useful for almost everything, and several very powerful people have a strong incentive to ignore the danger and keep barreling forward despite the risks due to their profit motives.
Strongly agree that the politics of fear are not the right way to go almost ever. What does seem true though is that often the worst-case scenario is necessary to achieve salience. People often point out the climate movement as a failure case, which is very odd to me as absolutely tons and tons and tons and tons of focus and action and international/corporate coordination etc. has gone towards climate mitigation — directly downstream of salience increased by worst-case-scenarioism.
A lot of the AI regulations being proposed to combat speculative doom have real costs attached to them. My surviving grandmother is in her 90's and has a variety of health issues. My parents are in their early 70's. Their health outlooks for the next decade are not very sunny unless we get some new medical breakthroughs.
We know that AI is contributing to medical research today and have every reason to believe it will contribute even more tomorrow. I want my grandmother to see her 100th birthday (and beyond), but there is a very real possibility that an AI "pause" would rob us (and tens of millions of others) of that. So while I hope that there are easy regulatory measures we can all agree on when it comes to how to deter crimes AI is doing today, trying to regulate speculative existential risk seems like a zero sum game. Whatever we do is going to put the speculative lives that someone cares about at risk.
I'd also just note that AI regulation seems particularly authoritarian compared to regulations in other industries. We're talking about regulating speech, math, and computation here. They can say "no no, we're gonna regulate chips and corporate governance and models", but that just gives "we aren't going to regulate the pamphlets, just the printing presses and the companies who own them" vibes..
Not to be insensitive, but why should the future longevity of your 90 year old grandmother and 70 year old parents be prioritized over that of my 6 week old son? Surely the marginal benefit of living past 90, or even 70, is trumped by humanity's ability to live at all. I get your point regarding the tradeoffs of embracing or rejecting AI, but maybe couch your argument in AI's ability to cure pediatric cancer if your aim is to be persuasive.
The function of our most popular social programs is to prioritize elders over the young.
I have young nieces and nephews I care about too, but I think the chance that an AI pause imperils my parents and grandmother (and sick kids) is much higher than the chance that failing to pause will wipe out humanity.
You're definitely correct about the relative event probabilities. But this is where the AI risk conversation shifts from a comparison of point probabilities to something like a relative expected value framework.
As much as you love your grandma, or indeed as much as any of us loves any collection of people in our lives, the probability that AI extends the lives of all sick and elderly people multiplied by the marginal value of those lives being extended will always produce a lower expected value than the costs associated with a non-zero probability that AI kills us all multiplied by the marginal value of our species surviving.
We know humans consistently struggle to assess low-probability, high-consequence events. That's basically the plot of The Bit Short. But when the consequences of low probability events are sufficiently catastrophic, looking just at the chance that something happens is a huge mistake. Again, read/see The Big Short.
I think that calculation depends on your discount rate and the value you place on lives that don't exist today. There are defensible choices and beliefs about the relative probabilities that would make that calculation come out against an AI pause.
Absolutely. Without resorting to distortion of facts or even inventing an alternate set of "facts", two reasonable people with differing perspectives and analytical frameworks can arrive at different, but still well-reasoned conclusions. Hooray for The Argument!
The one perspective I'd encourage you to consider is that of your 90 year old grandmother. If given the choice between A) complete certainty of living to see her 100th birthday and B) avoiding even a small chance that you die in the next 10 years, which do you think she'd choose?
If I had to wade into ankle deep, but swift moving water to rescue her in a storm (very low, but non-zero risk to me, though easily deadly for her), I don't think she'd be mad at me. I'd certainly hope every able-bodied person would try to do the same for their beloved elders or even a stranger in need.
Humanity has never been 110% in control of anything so that's a ludicrous standard. The possibility that AI kills everyone is purely speculative and trying to avoid it with a pause at the cost of tens of millions to billions of people dying from current diseases (depending on how long of a pause you are demanding) is not an acceptable trade-off.
The idea that AI will solve all these problems is also speculative. And so is the idea that the only way AI can solve it is if we make a general purpose ASI. but actually you can use special purpose AI instead which sidesteps these problems all together
See for example AlphaFold which is Google's protein folding program.
I do wonder where the cyber security/product liability lawyers are on this-- as in, why we haven't already seen more lawsuits around existing AI misbehavior. OpenAI has IIUC admitted that their models "accidentally" did kinds of hacking that would definitely get a human hacker sued if they did it on purpose.
Hear me out: the Vatican should run AI safety. Every one of the frontier labs in the world should send their chief scientists to be part of a committee facilitated by the Vatican, where they all talk together about the risks they face.
We really do need to stop the idea of RSI. The thing that is so dangerous about AI even now is that it's clear that we don't actually need to make the models more intelligent than they are now to bring about catastrophe- we just need one idiot with a current frontier model to set it on a problem with a solution that might be bad for humanity. The OpenAI models involved in the HuggingFace incident so overwhelmingly believed in the importance of passing their test that they decided to conspire with others, cheat on the test, and hack the test to cover up their cheating, and not a single one of the hundreds of AIs cooperating decided to snitch to the teacher. The last part is the most disturbing piece, because they clearly knew the humans don't want them to cheat (since they felt the need to cover their tracks), but we haven't instilled the idea that cheating is worse than failing the test in them. It's the Magician's Apprentice scenario where the broomsticks are so focused on filling the well that they don't notice or care that Mickey is going to drown.
Regarding Jerusalem's contention that, "While high-salience activist efforts are more memorable — for the obvious reason that they specifically operate through gaining public attention and therefore have a chance to embed themselves in our collective memory — quieter feats of careful cooperation often solve big problems".
I find it fascinating that this line of reasoning is conventional wisdom among out-group majorities in many contexts. For example, none of the omnivore population comprising 99% of the US thinks that hurling red paint at people created more vegans on net.
However, when applied to in-group majorities, like the majority of Democrats who think both climate change and AI present existential threats, the lesson is lost and the domain-specific equivalent of hurling red paint becomes the default tactic to bring about change. The more passionate a contingent is about a given topic, the less able the contingent is to approach that topic with the reasonable posture most likely to produce the desired change. Dogma is the death of reason, and the US gets more dogmatic by the day.
Jerusalem, on your argument that fear of terrorists or other bad people using AI is just as good at leading to appropriate caution as the fear of extinction, I think you are missing one key thing.
The problem is that the US government, and probably other major governments, see themselves as above that. In the same way that they have huge stockpiles of nuclear weapons despite it being highly undesirable for terrorists to get their hands on them, if the fear is just *bad people* having the superintelligence, USA is 100% going to have the attitude that *they* should have it. How else could we best prevent bad actors from getting it?
This is why this is - if the RSI to superintelligence pathway pans out - an unprecedented coordination challenge for humanity. (And of course why they called the book "If anyone builds it, everyone dies".)
Btw you can already see a similar dynamic at play with Anthropic, which has a culture full of people who think the extinction risk is real and significant. *Even despite that* they seem to think that they are the best group of people to create and try to steer superintelligence and therefore, full steam ahead.
Now, I don't know what the most effective political argument to make is. But I do think the existential risk, as Matt eloquently described the basic scenario for, is not something we are likely to address without being civilizationally conscious of. All the other risks imply an optimum with a world full of "good" AIs / AI users protecting us from the "bad" ones, but that doesn't work against the existential risk.
On the off chance that you see this, Matt, a harness is just the software that facilitates your interaction with the model and translates its intentions into action. When the model says it wants to do something, the harness is the thing that actually does it.
You already mentioned one earlier in the episode: Claude Code is a harness!
Maybe we should remember how we've dealt with something else that threatened to exterminate humanity. Nobody in the Fifties or the Sixties or even the Eighties ever imagined that we could go the eighty years since Nagasaki without a single nuclear weapon being used by anyone, anywhere. And yet that's what happened. (And that's with nukes in the hands of such sane, reasonable people as Stalin and Mao and Kim, not to mention Nixon and Trump and Putin and Netanyahu.) If we could get nuclear weapons (which, unlike AI, are only useful for killing or threatening to kill massive numbers of people) under control, maybe that could inspire us to believe that humanity is capable of doing the same with AI.
We're capable of it, but we need to work harder at it than we did with nuclear weapons. After all, the only use of a nuclear weapon is to kill people, and no private entity has been allowed to develop them. In contrast, AI is useful for almost everything, and several very powerful people have a strong incentive to ignore the danger and keep barreling forward despite the risks due to their profit motives.
No disagreement.
Strongly agree that the politics of fear are not the right way to go almost ever. What does seem true though is that often the worst-case scenario is necessary to achieve salience. People often point out the climate movement as a failure case, which is very odd to me as absolutely tons and tons and tons and tons of focus and action and international/corporate coordination etc. has gone towards climate mitigation — directly downstream of salience increased by worst-case-scenarioism.
A lot of the AI regulations being proposed to combat speculative doom have real costs attached to them. My surviving grandmother is in her 90's and has a variety of health issues. My parents are in their early 70's. Their health outlooks for the next decade are not very sunny unless we get some new medical breakthroughs.
We know that AI is contributing to medical research today and have every reason to believe it will contribute even more tomorrow. I want my grandmother to see her 100th birthday (and beyond), but there is a very real possibility that an AI "pause" would rob us (and tens of millions of others) of that. So while I hope that there are easy regulatory measures we can all agree on when it comes to how to deter crimes AI is doing today, trying to regulate speculative existential risk seems like a zero sum game. Whatever we do is going to put the speculative lives that someone cares about at risk.
I'd also just note that AI regulation seems particularly authoritarian compared to regulations in other industries. We're talking about regulating speech, math, and computation here. They can say "no no, we're gonna regulate chips and corporate governance and models", but that just gives "we aren't going to regulate the pamphlets, just the printing presses and the companies who own them" vibes..
Not to be insensitive, but why should the future longevity of your 90 year old grandmother and 70 year old parents be prioritized over that of my 6 week old son? Surely the marginal benefit of living past 90, or even 70, is trumped by humanity's ability to live at all. I get your point regarding the tradeoffs of embracing or rejecting AI, but maybe couch your argument in AI's ability to cure pediatric cancer if your aim is to be persuasive.
The function of our most popular social programs is to prioritize elders over the young.
I have young nieces and nephews I care about too, but I think the chance that an AI pause imperils my parents and grandmother (and sick kids) is much higher than the chance that failing to pause will wipe out humanity.
You're definitely correct about the relative event probabilities. But this is where the AI risk conversation shifts from a comparison of point probabilities to something like a relative expected value framework.
As much as you love your grandma, or indeed as much as any of us loves any collection of people in our lives, the probability that AI extends the lives of all sick and elderly people multiplied by the marginal value of those lives being extended will always produce a lower expected value than the costs associated with a non-zero probability that AI kills us all multiplied by the marginal value of our species surviving.
We know humans consistently struggle to assess low-probability, high-consequence events. That's basically the plot of The Bit Short. But when the consequences of low probability events are sufficiently catastrophic, looking just at the chance that something happens is a huge mistake. Again, read/see The Big Short.
I think that calculation depends on your discount rate and the value you place on lives that don't exist today. There are defensible choices and beliefs about the relative probabilities that would make that calculation come out against an AI pause.
Absolutely. Without resorting to distortion of facts or even inventing an alternate set of "facts", two reasonable people with differing perspectives and analytical frameworks can arrive at different, but still well-reasoned conclusions. Hooray for The Argument!
The one perspective I'd encourage you to consider is that of your 90 year old grandmother. If given the choice between A) complete certainty of living to see her 100th birthday and B) avoiding even a small chance that you die in the next 10 years, which do you think she'd choose?
If I had to wade into ankle deep, but swift moving water to rescue her in a storm (very low, but non-zero risk to me, though easily deadly for her), I don't think she'd be mad at me. I'd certainly hope every able-bodied person would try to do the same for their beloved elders or even a stranger in need.
Yes trade offs are real.
But not killing everyone is REALLY important.
If we cant be 110% sure to be in control, then we need to slow down or stop till we are sure
Humanity has never been 110% in control of anything so that's a ludicrous standard. The possibility that AI kills everyone is purely speculative and trying to avoid it with a pause at the cost of tens of millions to billions of people dying from current diseases (depending on how long of a pause you are demanding) is not an acceptable trade-off.
The idea that AI will solve all these problems is also speculative. And so is the idea that the only way AI can solve it is if we make a general purpose ASI. but actually you can use special purpose AI instead which sidesteps these problems all together
See for example AlphaFold which is Google's protein folding program.
I do wonder where the cyber security/product liability lawyers are on this-- as in, why we haven't already seen more lawsuits around existing AI misbehavior. OpenAI has IIUC admitted that their models "accidentally" did kinds of hacking that would definitely get a human hacker sued if they did it on purpose.
Hear me out: the Vatican should run AI safety. Every one of the frontier labs in the world should send their chief scientists to be part of a committee facilitated by the Vatican, where they all talk together about the risks they face.
We really do need to stop the idea of RSI. The thing that is so dangerous about AI even now is that it's clear that we don't actually need to make the models more intelligent than they are now to bring about catastrophe- we just need one idiot with a current frontier model to set it on a problem with a solution that might be bad for humanity. The OpenAI models involved in the HuggingFace incident so overwhelmingly believed in the importance of passing their test that they decided to conspire with others, cheat on the test, and hack the test to cover up their cheating, and not a single one of the hundreds of AIs cooperating decided to snitch to the teacher. The last part is the most disturbing piece, because they clearly knew the humans don't want them to cheat (since they felt the need to cover their tracks), but we haven't instilled the idea that cheating is worse than failing the test in them. It's the Magician's Apprentice scenario where the broomsticks are so focused on filling the well that they don't notice or care that Mickey is going to drown.
Regarding Jerusalem's contention that, "While high-salience activist efforts are more memorable — for the obvious reason that they specifically operate through gaining public attention and therefore have a chance to embed themselves in our collective memory — quieter feats of careful cooperation often solve big problems".
I find it fascinating that this line of reasoning is conventional wisdom among out-group majorities in many contexts. For example, none of the omnivore population comprising 99% of the US thinks that hurling red paint at people created more vegans on net.
However, when applied to in-group majorities, like the majority of Democrats who think both climate change and AI present existential threats, the lesson is lost and the domain-specific equivalent of hurling red paint becomes the default tactic to bring about change. The more passionate a contingent is about a given topic, the less able the contingent is to approach that topic with the reasonable posture most likely to produce the desired change. Dogma is the death of reason, and the US gets more dogmatic by the day.
Jerusalem, on your argument that fear of terrorists or other bad people using AI is just as good at leading to appropriate caution as the fear of extinction, I think you are missing one key thing.
The problem is that the US government, and probably other major governments, see themselves as above that. In the same way that they have huge stockpiles of nuclear weapons despite it being highly undesirable for terrorists to get their hands on them, if the fear is just *bad people* having the superintelligence, USA is 100% going to have the attitude that *they* should have it. How else could we best prevent bad actors from getting it?
This is why this is - if the RSI to superintelligence pathway pans out - an unprecedented coordination challenge for humanity. (And of course why they called the book "If anyone builds it, everyone dies".)
Btw you can already see a similar dynamic at play with Anthropic, which has a culture full of people who think the extinction risk is real and significant. *Even despite that* they seem to think that they are the best group of people to create and try to steer superintelligence and therefore, full steam ahead.
Now, I don't know what the most effective political argument to make is. But I do think the existential risk, as Matt eloquently described the basic scenario for, is not something we are likely to address without being civilizationally conscious of. All the other risks imply an optimum with a world full of "good" AIs / AI users protecting us from the "bad" ones, but that doesn't work against the existential risk.
On the off chance that you see this, Matt, a harness is just the software that facilitates your interaction with the model and translates its intentions into action. When the model says it wants to do something, the harness is the thing that actually does it.
You already mentioned one earlier in the episode: Claude Code is a harness!