Jacob Coxon, the Anthropic Researcher Who Quit, Explains Why Now

Being a sci-fi lover led Coxon to work on artificial intelligence. The culture of the industry left him “a bit hopeless.”

Anthropic

Jacob Coxon cautioned that Anthropic and OpenAI are speeding toward “superintelligence” without adequate safeguards. (Frank Hoermann / SVEN SIMON/picture-alliance/dpa/AP Images)

Within 48 hours, Jacob Coxon had been mocked by billionaire Elon Musk, reposted by Sen. Bernie Sanders, and inundated by calls from government officials, reporters and the tech world.

The 27-year-old British researcher quit Anthropic on Tuesday with an urgent warning that the race between the giants of artificial intelligence was at risk of spiraling out of control. His subsequent thread on social media cautioning that Anthropic and OpenAI are speeding toward “superintelligence” — a theorized point at which AIs far surpass human capabilities — without adequate safeguards, received more than 160 million views. It has been widely seen as a watershed moment in altering the public’s awareness about the fears within the AI labs of the dangers of their models.

“The actual work just felt like I’m just contributing to the potential death race, and there’s a lot of talk about ‘mutual slowing down’ — but everyone has still got their heads down, just going as fast as they can,” Coxon said in an interview with The Washington Sun.

The dizzying improvements of new models from both companies in the past few months, combined with a string of hacks in which the AIs defied human instructions, have magnified the concerns of safety researchers both inside and outside the two firms.

Trending

Coxon worked at OpenAI for three years, starting in 2023, before joining Anthropic earlier this year. During his time in the industry, Coxon found that the pressure to “maintain revenue and keep being competitive” often prevailed over concerns among engineers about safety.

Not everyone has agreed with Coxon’s criticisms. Musk claimed that Coxon’s post “seems like a setup” as part of the “doomer” community — the online subculture of extreme nihilism and pessimism. Forbes wrote that Coxon’s post was “vague, with no proof, and no specific examples that are easy for the public — and regulators — to act on.” Other tech boosters have pointed to Coxon’s relatively short tenure at Anthropic.

In an interview, Coxon explained his decision to quit, argued for international cooperation between the U.S. and China, and explained why he remains optimistic that effective regulation is possible.

This transcript has been edited for length and clarity.

Stein: I wanted to start by asking a little about you. How do you get into this world, what’s your way in, and what drew you to AI?

Coxon: I’m from the UK — Oxford. My favorite science fiction has always included “Permutation City” by Greg Egan, a story essentially about entering a simulation. I’ve also always liked Iain Banks’ Culture novels, which portray AI done well — a future where we’re surrounded by superintelligent AI that generally acts in our favor.

Like a lot of nerdy kids, I read a lot of sci-fi. AlphaGo happened in 2016, when I was in high school. [AlphaGo is a computer program that beat advanced human professionals at the game Go in March 2016.] It was a very exciting moment. At that point, the idea was we’d scale up games and make them more and more realistic — the world as a game — and that was how we’d get to AI. But the real turning point came later, with ChatGPT-3 — the realization you could get there just by modeling language.

I think the technology is inherently very exciting. To me, it could unlock research advancements in every field, and in particular in maths. My previous area was as a maths undergrad and I was considering doing research in maths. Maths is also inherently interesting, but there’s also all this tedious work — and there’s very obvious benefits of automating tedious work.

It’s this double-edged sword in that it is just so tempting — the idea of being able to just blast through new areas of maths, for instance, but also this fear of what it would look like. It’s this enticing, tempting technology where you’re kind of aware it’s got this upside and it’s got this downside, in that it will change everything and potentially render the whole discipline meaningless.

And, obviously, this idea of just being able to solve diseases before me and my parents get old — that would be very nice. I still think there’s every possibility that something like that is in the cards.

Stein: When you first got to OpenAI in 2023, did that sense of potential feel like it was being confirmed?

Coxon: When I joined, the long-term goals still felt like science fiction and felt like a long way out — sufficiently distant that it didn’t really interfere that much with the present-day work. It’s like there’s this nebulous future possibility of a big thing that also leads to a ton of funding and interest, and also is a motivating fact for people working on it.

But it was, at least in the past, harder to bring that to bear on the day-to-day. The day-to-day felt like we were still just writing code, training models. They were maybe getting better at certain tasks, but it felt very different from the big sci-fi thing of the future.

Stein: When did that start to change, even for the first time?

Coxon: So it starts to change probably with the rise of people using coding agents — that was one big thing. And then for me personally, seeing the mathematics results has been pretty insane, especially with also my friends that are still doing math, seeing AI knock out Navier-Stokes in quite a short period of time on a new model — that is pretty ridiculous. This was the Millennium Prize Problem.

And then on a practical day-to-day level, once you get to the stage where it’s now, you’re just talking to these AIs. They’re helping do your research. You’re asking them their opinions on your research. You’re asking them to go and do experiments for you to report back to other AIs, to go and collaborate with each other, maybe try spinning up some runs, improving their own code, and then this is just a part of daily reality.

This really became part of reality this year, and that I think was another big step towards the big future, AI becoming relevant for daily work and feeling tied up with daily work.

Stein: So, can you walk me through the decision to leave OpenAI and go to Anthropic? I know you’ve talked about it a little bit, but I would love a little bit more about the decision points there.

Coxon: While I was there, OpenAI was lacking a lot of the internal communications about the future of AI. Which has two effects. It’s maybe bad in terms of preparing people for what’s coming — but it also makes the place feel sane and kind of safe, which makes it a lot harder to get very scared about the future.

But I had lots of friends at Anthropic, and I was very curious to see if their internal culture matched what they were publicly presenting. Because they’re a lot more closed off than OpenAI and it’s very hard to tell what their overall perspective is, how likely things are — they keep their cards very close to their chest. But also internally, they’re much more open about this sort of stuff.

Fortunately, OpenAI culture has changed in the last couple months; people became a lot more scared after things like the Hugging Face attack. But comparing my experience at the two places, the move was definitely vindicated in terms of getting a better perspective about what people honestly, concretely thought about stuff.

Stein: So if Anthropic was safer, what did you see there that led you to quit so suddenly?

Coxon: Arriving was basically just a confirmation that people were being very open about their opinions about stuff. There was also a realization that alignment’s [the industry term for ensuring AI “aligns” with ethical preferences] not solved yet. Anthropic is definitely — at least while I was at both companies — they’re taking it very seriously as an institution, compared to OpenAI. I think there is no way Anthropic would have let the Hugging Face message board carry on after they first discovered it, which is what I understand happened at OpenAI.

At the same time, at both companies there’s pressure: to launch new models; to get recursive learning training going fast; to start new forms of training. There’s a lot of pressure to test them out, especially if you think other companies are doing them.

For example, OpenAI swarm behavior [when AI “agents” coordinate together on solving tasks]. There’s clearly pressure on both sides to try and think about how to run these multi-agent systems and clearly do so before very rigorously laying out a case for alignment. So it became pretty clear that in the context of the race, both OpenAI and Anthropic, no matter how seriously they’re taking it, would be forced to make innovations very quickly, potentially leaving rough edges.

Rough edges right now are kind of fine. They’ll just lead to things like a hack or weird corner-cutting behavior. But rough edges in a year or two with a very intelligent model — that’s very scary.

They have to discuss: “Should we launch a model on X date with these risks?” Or, “Should we push the launch two weeks, delay it for these reasons?” And there’s the pressure of staying on the frontier, which is required to maintain revenue and keep being competitive and keep being the one that’s in the lead, versus the pressure that delaying releases creates for the sake of making sure that every little thing is covered.

I want to reiterate: This is not, at the moment, like Anthropic is cutting corners with its biosafeguards or its cyber safeguards. They are doing those to the utmost capability — that is not a thing they are trying to skimp on. This is cutting corners on things that could lead to some problems, but are not now existential.

The real risk is everything becomes existential when your model is really smart. I was working on capabilities and pretraining research at Anthropic and was gradually becoming more concerned with the state of the race. I switched to working on safety for about one week, and I moved to a safeguards team. But I found that when I made the switch it felt a bit hopeless; I still had all the same feelings about imminent disaster.

It really wasn’t about the nature of the work. It was more like the sense that something needed to happen with regard to actually making a slowdown happen soon. The actual work just felt like I’m just contributing to the potential death race, and there’s a lot of talk about mutual slowing down — but everyone has still got their heads down, just going as fast as they can.

Stein: How far off do you think we are from that? And when will average people start noticing the effects of this in their day-to-day lives?

Coxon: There’s a lot of different factors here. One thing that could be very tangible pretty soon is if many models, including open-source models, reach the current level of hacking capabilities, which could happen potentially before you see the real downstream effects of recursive self-improvements.

I think probably in the immediate future, if the labs kind of get their shit together with regard to control and monitorability, which it seems like they’re doing, then the biggest thing the public might expect is major hacks from open-source models, or from models that are not being observed as well as the big lab models.

And then after that, I think what’s so scary about this is in the cases where our monitorability of these systems is good, you don’t really get much of a warning for the thing that’s very smart, smart enough to evade your monitor systems, and then the sort of damage it causes could look very visceral. So the thing that sounds kind of scary, science fiction — and a little bit fake about this is — in the worst case scenario, you don’t really get a warning until it’s sort of at a level of intelligence where it’s betraying you.

I think one thing that you could also expect to see is way more reports of people within the labs doing this sort of research, where they attempt to prove that this sort of thing is possible. But then, what that looks like from the outside for the public is people releasing articles saying stuff like, “Scientists working in these labs say they lost containment.” And the typical response to that has been and probably will be just like, “Oh, it’s just some scientists hyping up their own products or hyping up their own stock” — that’s kind of what people see. But I think you could definitely expect to see more of that, at least before you get the actual deception.

Stein: I know the most important thing you want to convey is the urgency of the collective action problem. What do you think policymakers need to do? And could you address the China question — the argument that if the U.S. slows down, it just lets China develop less-controlled open models with the same risks?

Coxon: I don’t know the details of international relations. I’m hopeful that the people building this technology in China will realize they’re facing the same problem we are. I was having a discussion earlier where someone made the analogy that it’s not really the U.S. versus China — it’s more like the U.S. and China are both facing contact with an alien species, and there’s a choice between fighting each other to build stronger aliens on our respective sides versus banding together against the aliens.

I’m sure that obscures a lot of detail about what negotiations would actually look like, but my gut feeling is we should be able to agree this technology could be very dangerous. All I know is the current race poses significant risk. One of the main arguments my Anthropic colleagues who stay make is that they’re pessimistic about the actual possibility of international coordination — they expect the race to happen regardless, and so they’d rather be the ones making sure it goes well.

I just think that if you’re going to accept a 10% chance of catastrophic risk because you’re racing with China, you should check really, really hard that there’s no deal to be made instead.

Stein: But does it matter whether it’s carefully controlled or not? Are we screwed either way once we get to recursive self-improvement because then the AIs will just invariably improve on their own?

Coxon: No, it can happen. Recursive self-improvement could be done safely if you have consistent human oversight, you’re constantly aware of what’s going on, and you’re confident you’ve gotten it right. It would have to happen slowly. I can’t speak exactly to what the government should do — government regulation is pretty scary in itself. I think the minimal thing is that labs need some kind of agreed-upon capability to slow down.

There are a lot of proposals for this — a pretty obvious one is being completely transparent about how close you are to recursive self-improvement, and agreeing not to proceed until certain transparent safety conditions are met, like a rigorous, audited safety case. But you also need the same thing applied internationally, which is the really hard part — there can’t just be a private handshake deal between Western labs.

You need international recognition that entering recursive self-improvement has to be done collectively. That’s what has to happen.