The Code/X ArchiveView on X
Samuel Marks

@saprmarks

[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]

Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:

1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.

2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.

3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.

4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.

5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed): www.pacingthefrontier.com/

I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.

Jacob Coxon @hilbertspaess

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

Quoted post on X →

7174K20.6K7.4K