Arvind Narayanan:AI 安全运动应做大帐篷还是小帐篷
A big-tent or small-tent AI safety movement?
Arvind Narayanan 提出 AI 安全运动面临大帐篷与小帐篷两种路径的选择。他既不认同生存风险迫在眉睫,也不认为相关警告是炒作或监管俘获,而是认为警告可能真诚但错误且适得其反。
In recent weeks, two narratives about AI safety have emerged: either AI existential risk is real and imminent, or AI leaders’ and whistleblowers’ claims to that effect are insincere — a “psyop” or hype or a twisted form of regulatory capture.
Few have considered the third possibility that x-risk warnings are sincere but simply wrong and counterproductive to AI safety. It’s as if everyone takes for granted that those raising the alarm are geniuses, and the only question is whether they’re the benevolent or the evil kind of genius.
We subscribe to neither narrative. In our writing on AI safety, we’ve consistently tried to take the arguments seriously but rebut the many claims we think are unsound while highlighting safety concerns that we think are real. This essay is a continuation of that agenda.
It is true that we systematically underinvest in resilience against catastrophic and systemic risks
Let’s start by steelmanning a subset of AI safety and Effective Altruist (EA) arguments. Something is deeply wrong in the world when it comes to how much we invest in defending against catastrophic risks, and the safety and EA communities are right about that. Take superintelligence out of the equation entirely and it remains true that evidence-grounded risk arguments have been made for years and not produced a commensurate policy response.
The most well known example is the failure to prepare for pandemics. We had years of warnings and quantitative estimates. A 2019 report convened by the WHO/World Bank — before COVID — began by saying there is a “very real threat ... of a respiratory pathogen killing 50 to 80 million people and wiping out nearly 5% of the world’s economy.” It estimated the cost of Adequate preparedness at a mere $1–2 per person per year for most countries.
Since the 2009 H1N1 pandemic, at least 11 high-level panels and commissions have made specific recommendations to improve global pandemic preparedness, and yet “the majority of recommendations were never implemented”. More remarkably, even after COVID, only a fraction of the need was funded. The same independent panel wrote in 2024 that “There is a cynical attraction to providing significant and immediate resources during a crisis, rather than proactively preventing one. Yet history shows this will not end well.” The root cause seems to be that voters reward spending on relief, not preparedness.
To be clear, the inadequacy of pandemic preparedness is by no means a uniquely Effective Altruist insight, let alone AI safety. But EA does deserve credit for championing the cause. It’s also true that EA’s consequentialist moral framework that centers expected-utility calculations is highly effective at revealing and quantifying our underpreparedness. The framework is highly controversial, but we don’t think there’s anything wrong with it or its application to areas like biorisk. We do, however, strongly disagree with its extensions to longtermism and AI existential risk. We also disagree with the tendency to reframe problems such as AI’s amplification of biorisk as “AI safety” problems to be solved through methods such as alignment. Our view is that AI is an amplifier of existing systemic risks.
The underpreparedness problem is also true of cybersecurity — actually, it’s worse. Unlike pandemics or finance, there is no culture of modeling systemic or catastrophic risks in the cybersecurity community. The economic approach we described in the previous essay has been successful in keeping the expected losses from everyday attacks at a manageable level, but markets are notoriously bad at appropriately pricing tail risks and externalities.
Insurance companies know this. Lloyd’s is pretty blunt about the fact that “losses have the potential to greatly exceed what the insurance market is able to absorb”, and does not insure against severe state-backed cyberattacks (that was in 2022).
Once again, this has little to do with AI’s recent capabilities. Five years ago, before AI was in the picture, when teaching undergraduate information security at Princeton, Arvind emphasized that cascading failures and catastrophic risks are possible, resulting in essentially the whole internet going down for an extended period, with potentially civilization-altering consequences.1
(Looking back at the speaker notes in my slides, I’m reminded that I apologized for presenting original research in an undergrad class that’s supposed to be about basics, but explained that this is because the near-total lack of attention to this critical topic in the security community left me no choice. — AN.)
In other words, Dario Amodei’s claim about agent swarms potentially being capable of “taking over the entire internet”, while worded somewhat hyperbolically, is not obviously outside the realm of possibility, and deserves study at the very least. The cybersecurity community has the right skillset to do that, but unfortunately they don’t seem to find this even worth thinking about.
There are also systemic risks that are diffuse and gradual rather than catastrophic. Consider what chatbots are doing to news traffic — clickthrough rates are far lower than traditional search, accelerating the decline of traditional journalism. The second-order societal effects of this shift won’t be apparent for quite a while. And that’s just one of dozens of types of creeping institutional decay. Atoosa Kasirzadeh introduced the concept of accumulative risk to describe this gradual erosion of societal resilience.2 Here, the psychology of inaction is the reverse of the case of pandemics: rather than being so rare that we are able to memory-hole the event, decay is so pervasive and gradual that we become inured to it.
In short, based entirely on well-grounded rather than speculative causal pathways, it is reasonable to believe that policymakers are drastically underinvesting in protecting society against AI risks, and it is hard not to empathize with those who want to ratchet up the rhetoric.
Of course, the question is whether more urgent public warnings actually lead to better policymaking.
The Coxon incident shows how the warped media environment elevates x-risk over more grounded concerns
This newsletter is largely an exercise in perspective-taking. Let’s take a moment to articulate the primary reason why people inside and outside the AI bubble tend to talk past each other.
For people in AI — whether or not they’re concerned about safety — it’s hard to understand how anything else in the world matters right now. As computer scientist Scott Aaronson put it in his essay The Age of Wonders and Terrors:
“Yes, there’s still enormous uncertainty about what the rest of our lives will look like, but as far as I can tell, there’s no longer any real uncertainty that it’ll all mostly revolve around AI, and the extent to which we succeed or fail at directing its power toward human flourishing.”
This essay made the rounds in the community because this perspective is widely shared. Even the two of us — best known for pushing back against some elements of this view — consider AI important enough that we decided long ago to devote our careers to it. The AI as Normal Technology position is that AI will “only” be as impactful as the industrial revolution.
People outside AI don’t understand the intensity of this view in the AI community — that it’s an AI world and we’re all just living in it.
But there’s a flip side — people in AI vastly overestimate how much the public cares about AI. In Gallup polls, the share of people who rate AI (or “advancement of computers/technology”) as the most important issue is about half a percent. Most of those people are likely thinking of relatively mundane aspects such as a datacenter being built in their community. The share of people outside AI who treat it as the defining issue of our time is incomprehensibly small.
The reason AI people fail to recognize this is that AI is a topic where salience and concern are decoupled from each other. There is high concern about AI (that is, many people think AI is important or express concern when specifically asked about it) but very low salience (that is, few people rank it as an important issue versus others).
Anyway, all this is mere throat-clearing to explain why, even though there are many catastrophic and systemic AI-amplified risks that we agree the world urgently needs to act on, they don’t tend to break through to public consciousness.
After the agent swarm hacks over the summer came to light, there was a smattering of headlines, but it was orders of magnitude below the level of media attention and public awareness that Coxon’s resignation and x-risk predictions would later draw.
The safety community expressed amazement and frustration that these concerns hadn’t broken through to the public. This put the whole community on edge. As observers of the community, we felt the level of anxiety in the community to be qualitatively different from any other moment in the past.
It is not hard to explain why AI cyberrisk didn’t get more headlines: AI is a low-salience topic, hacks are in the news every day, and this one didn’t even result in any known harm to ordinary people. It is far from obvious why everyday people should care.
In contrast, existential risk is visceral, and the Coxon resignation made AI safety finally break through. To be clear, x-risk warnings have been made many times before, but there were two differences this time. First, there was a tinderbox waiting to be lit, namely the safety community’s eagerness to talk to the media and amplify concerns once there was a hook in the form of a newsworthy event. This response was so full-throated and unanimous that it left people concluding that it must have been a co-ordinated effort, but we don’t think that was the case.
There’s another reason this time was different: the revelation that many executives have a high “p(doom)” and yet continue down the path of racing to build superintelligence, a choice that feels reckless and incomprehensible to the public. Of course, there is a complex set of rationalizations that AI leaders invoke to explain why they think they have no other choice, but that’s not what reaches people. Perhaps the bigger disconnect is that it is all underpinned by a utilitarian value system that the public doesn’t share.
Why the x-risk framing may be counterproductive for AI safety policy
To recap the argument so far, there are two clusters of beliefs about what AI safety means. Like any binary, this one leaves out some nuances, but we think it is accurate enough to be useful.
There are specific (potentially catastrophic but sub-existential) risks, predating AI but amplified by it, that we urgently need to defend against through prevention and resilience.
AI development itself is an existential risk — either because there are too many downstream risks to defend against once superintelligent AI is built, or because of unknown unknowns — and we either need to figure out how to make superintelligence safe (such as through alignment) or stop it from being built.
Critically, due to the information environment, the x-risk framing is the one that is driving public concern and much policy discourse, despite the fact that credible causal pathways haven’t been articulated (and attempts to do so are unconvincing to say the least).
So what’s going to happen next? One happy possibility is that some of the energy of the current moment gets channeled into building a broad set of defenses (perhaps accompanied by a slowdown in AI development, though that is much less important in our view).
But this path seems unlikely to us. The more likely outcome is that the dominance of the x-risk framing is actively counterproductive, for two main reasons.3
Polarization and partisanship. The doom framing has always been polarizing because many of us remain completely unconvinced by it, but now the polarization is starting to split along partisan lines. Partisanship is bad for safety policy because it leads to inaction, high uncertainty, and partisan policy reversals. Worse, when people realize that there is no quantitative model behind the “10% extinction risk” predictions, they will feel misled, resulting in even more pushback to safety policies, especially in a polarized environment where one side is already wary of experts making stuff up.
Further, causing the public to panic is counterproductive because it constrains the solution space toward those that are the most blunt and offer an emotional salve, which don’t tend to be the most effective.
Misdirection. We have repudiated the distraction argument against x-risk concerns — if we’re worried about job loss and safety, we can address both, and they don’t distract from each other. But this isn’t about distraction: the two views of safety represent two different diagnoses of the same problem. It seems likely that they will compete for resources. Let’s call it misdirection, to distinguish it from distraction.
In particular, a generalized existential threat admits exactly one policy — a ban. Framing the threat in terms of probabilities biases us even more in that direction. The statement “if we develop superintelligence there’s a 10% risk of extinction” makes it look like the only lever for intervention is stopping development. Unfortunately, a ban on superintelligence is definitionally ambiguous to the point of being incoherent.4 And AI safety is still not a model property. Other potential interventions such as pausing datacenter construction are exceedingly impotent as safety measures.
Most importantly, bans on development do nothing about the catastrophic and systemic risks from existing AI systems. The response we need to those risks is much less sexy, which in turn means that the movement one would build to address them is quite different from one targeted at superintelligence: a hundred disparate interventions to shore up defenses and resilience rather than concentrating effort and attention toward erecting one mammoth, flashy wall. This is why we worry that the two framings of AI safety are in competition. We hope we are wrong about this!
Conclusion: a big tent or small tent AI safety movement?
There are three main characteristics of what we call big-tent safety.
Beliefs. A big-tent movement would include members who are not concerned about x-risk (and, like us, may not consider superintelligence a threat). They may care about cyberrisk, biorisk and other specific catastrophic risks, various accumulative risks, and the need to build societal resilience.
Values and movement building. The social movement must be pluralist in terms of values and not just beliefs, recognizing that EA / safetyist moral frameworks are unpopular with the public. The public appeal should be based not on panic but articulating positive visions.
Policy. A big-tent movement would recognize that there is much common ground on policy despite divergence in beliefs (though these policies arguably fall short if you’re narrowly focused on x-risk). Two such clusters of policies have become clear: transparency and liability. Embedded evaluation is urgent, and organizations that build and use AI should be systematically accountable for the negative externalities they create. Big-tent safety policy focuses on picking such low-hanging fruit while building muscle for stronger interventions like bans (such as banning fully autonomous recursive self-improvement). But the latter should involve doing the hard work of convincing skeptics of their necessity rather than pursuing policy-by-panic.
In contrast, small-tent safety is narrower at the level of beliefs, values, and policy goals. It prizes ideological homogeneity (while claiming the contrary). This risks a purity spiral, gradually driving out big-tent safety people. Small-tent safety de-prioritizes spending on defenses against risks that are “only” catastrophic (!). Finally, such a movement doesn’t recognize the gap between itself and the public in terms of both values and the salience of AI safety. As a result, it vastly overplays its hand, risking backlash.
On a personal note, we’ve often wondered why our work has been sometimes welcomed and sometimes attacked by the AI safety community, and we think the big-tent versus small-tent distinction captures it well.
Let’s build a bigger tent.
We are grateful to Abi Olvera, Joshua Saxe, and Rohit Krishnan for feedback on a draft.
Why is the cybersecurity community allergic to discussion of catastrophic risk? We have a few guesses.
Historically, “Cyber Pearl Harbor” warnings have been issued one too many times and were generally unproductive (and even a distraction from low-severity / high-frequency risks).
We’ve never had an actual catastrophic cyber incident — during the heyday of worms in 2001-2004, the internet was much smaller and the economy depended much less on it. So it is harder to prepare for catastrophic cyber-risk given that we have zero data points on what a catastrophe actually looks like and what the recovery path would be.
The computer networking community does have a culture of measuring concentration, and this work helps explain why catastrophic attacks are possible — bringing down Cloudflare alone can disrupt tens of thousands of other websites and services. But there are two major caveats. First, the work stops there — it doesn’t extend to cascading failure, resilience / recovery, etc. Second, measuring concentration is not very actionable on its own. There are legitimate reasons for concentration, including that it improves average-case security. We don’t see the trend toward concentration as reversible. In fact, the additional risk that today’s AI poses will likely lead to an increase in concentration.
She discusses accumulative existential risk, in contrast to decisive existential risk, but the accumulative-versus-decisive distinction is extremely helpful even if we don’t think it’s existential.
These are downsides of the rhetoric itself. If x-riskers actually succeed in enacting a nonproliferation agenda, the consequences both for safety and growth will be incalculably bad, as we’ve written about. Fortunately open-weight models may have left the barn.
A ban on recursive self improvement without human oversight seems more tractable, and could make sense even if one doesn’t think RSI will lead to superintelligence.
来源:AI as Normal Technology(RSS) · normaltech.ai