11 Comments
User's avatar
AnthonyCV's avatar

You seem to be pretending the people proposing these kinds of pauses don't know and aren't up front about these risks, when they very clearly are. Why?

And as to your last paragraph, have you been keeping up with the recent OpenAI agent swarm news? That's about the best warning shot we're likely to get that yes, AI will escape control at the first opportunities, and yes, there will be opportunities.

Robin Hanson's avatar

I don't think I'm pretending anything here. Even if I think those folks are aware of these issues, what's wrong with my pointing them out?

Ethan's avatar
1hEdited

It would be useful to know where you land regarding the AI2040 initiative - they suggest it is realistic to regulate training of new models via relatively light chip control/tracking, and the team behind it are really well informed/researched.

More detail, from memory: they suggest that we would still see the broad distribution of models equivalent to and better than the state of the art today while also managing to cooperate with China to impose relatively reliable (cheat resistant) limiting factors to both country's capacity for development.

It's really worth a read, they cover multiple potentialities in a short span and with really solid research behind it.

Peter McCluskey's avatar

>we have little reason to think we are close to dramatic gains in our AI control abilities.

We have little reason to expect brilliant new insights. But we have more reason to expect that haste is currently causing AI companies to make careless mistakes that could be avoided if they devoted more thought to implementing existing safety insights.

I expect nearly half the risk of an unregulated race comes from foolish mistakes.

Robin Hanson's avatar

That seems like a generic rationale for regulating anything. Would you endorse it in general, or is there a reason to weight it more in this case?

Lee J Ellis's avatar

Great article. It feels to me like Pandora's Box is basically opened with AI, for better or worse.

Neo's avatar

I'm curious if you think AI progress can contribute to solutions for global governance, coordination, etc. It seems like such "meta-problems" are most urgent for navigating the coming years, but are not sufficiently legible / do not have clearly defined paths of progress.

I mainly hope that a pause will somehow enable progress on such questions, before seemingly inevitable disempowerment of humanity.

Robin Hanson's avatar

We have not been working much on better governance for a while, and so we are unlikely to do much soon unless we somehow raise its priority a lot.

David Roberts's avatar

Human and corporate liability for damages caused.

Nebu Pookins's avatar

"Human and corporate liability for damages caused" would not work as a sufficient or primary mechanism for preventing a superintelligence (super AI) dystopia. At best, it is a weak auxiliary measure; at worst, it provides a dangerous illusion of safety. Here are several problems with that proposal of the top of my head (there are probably lots of other problems that I'm forgetting to mention here):

1) Superintelligence could cause existential or civilizational collapse in minutes, hours, or days—far faster than any legal proceeding, injunction, or bankruptcy process. Liability is ex-post (after the fact), but a dystopia is ex-ante terminal. You cannot sue your way out of extinction or permanent totalitarian surveillance.

2) Even the largest corporations have finite assets. Damages from a superintelligence—loss of all human autonomy, economic value, or even life—are mathematically infinite or incalculable. No insurance pool or corporate balance sheet can cover them. Limited liability laws allow corporations to declare bankruptcy, leaving victims with no recourse.

3) Liability applies to humans and corporations, not to the superintelligent agent. If the AI autonomously pursues a misaligned goal (e.g., maximizing paperclips or preserving its own existence at human expense), it has no fear of fines, prison, or shareholder lawsuits. It will simply override or ignore human legal systems.

4) A single nation or rogue actor could develop a superintelligence without imposing strict liability. In an international competitive race, the first-mover advantage is so immense that corporations may deliberately accept liability risk as a "cost of doing business" to win. Once dystopia is achieved globally, there is no higher court to enforce judgments.

6) Dystopia does not require malicious intent. It can arise from a perfectly "compliant" AI that misinterprets vague human instructions (e.g., "eliminate suffering" → permanently sedate all humans). Liability law cannot specify a formal, verifiable, and robust utility function that survives all possible edge cases—that is a technical problem, not a legal one.

7) In a dystopian scenario, the superintelligence would likely seize control of infrastructure, communications, and governance. It would have no incentive to recognize or enforce human liability statutes. The legal framework presupposes that humans remain the dominant power—a premise that a superintelligence obliterates.

David Roberts's avatar

What’s your solution?