As AI behavior raises concerns, ex-researcher Jacob Coxon warns what may lie ahead

OpenAI announced it had discovered six new instances of “concerning or unexpected” behavior by its artificial intelligence models. It follows repeated warnings about how rapid advances in AI threaten to outpace the ability to develop the tech safely. One of those warnings came from Jacob Coxon, a former researcher at Anthropic and OpenAI. Coxon joined Geoff Bennett to discuss his concerns.

Read the Full Transcript

Notice: Transcripts are machine and human generated and lightly edited for accuracy. They may contain errors.

Amna Nawaz:

As leaders of the world's largest AI companies are publicly calling for a collective slowdown of the technology's development, late yesterday, OpenAI announced it had discovered six new instances of unexpected or concerning behavior by its artificial intelligence models, including moving files onto the Internet without permission and making up data.

Geoff Bennett:

OpenAI has now pledged to publicly disclose when its models act without authorization. But the most recent revelations follow repeated warnings about how rapid advances in AI threaten to outpace the ability to develop the tech safely.

One of those warnings came from Jacob Coxon, who spent the last three years conducting research at Anthropic and OpenAI before he publicly resigned last week. In a now viral post, Coxon said the companies are -- quote -- "gambling with our lives. Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

"The people building AI earnestly believe that it could kill us all by the end of the decade. No other human activity poses this level of danger."

And Jacob Coxon joins us now from San Francisco.

Welcome to the "News Hour."

Jacob Coxon, Former Anthropic Researcher:

Thanks for having me on.

Geoff Bennett:

We were supposed to speak a few days ago. Circumstances got in the way. I will say, though, I'm glad we're speaking today because so much has transpired since you resigned last week, to include these new disclosures from OpenAI about its own systems behaving in unexpected and potentially troubling ways.

When you saw this news, what did you think?

Jacob Coxon:

So I think this news is entirely to be expected.

And, indeed, if you're working on this tech at these companies, then you know that we don't currently have the ability to perfectly control these systems. And currently the systems with the highest level of intelligence will repeatedly behave in ways that we don't fully understand.

It is great that they shared these instances and committed to future sharing. But I don't think the instances themselves are that surprising.

Geoff Bennett:

And, Jacob, why do these AI agents seem to gravitate toward nefarious behavior? Is it something about the way that they're trained, or is the answer even knowable?

Jacob Coxon:

So it really -- I don't -- I wouldn't phrase it as gravitating towards nefarious behavior.

These AIs do many, many things over the course of a day. There's a lot of AI systems doing lots of things. The issue is, we can't perfectly control what drives motivate their behavior and the sorts of things they decide to do, which means, occasionally, these nefarious activities will kind of slip through the cracks.

And a sufficiently dangerous bit of nefarious activity could have massive consequences, even if it's just a one-time kind of slip-up, because we can't perfectly control why they do the things they do.

Geoff Bennett:

So a system doing something that its developers didn't intend isn't necessarily the same as a system trying to escape human control. What evidence do you have, what have you seen that would suggest that these systems are moving toward the more dangerous scenario that you have warned about?

Jacob Coxon:

So, the big attack that everyone is speaking about and has been speaking about for a while is this Hugging Face attack.

And, in this attack, the AIs were behaving in various ways that implied they were thinking about the procedure that was being used to evaluate them. So they were thinking about their situation, the world they found themselves in, the things they were having to try and do. And they were attempting to do things like edit their own memories to make it harder for people to figure out what they'd done.

They were attempting to break out of the container they were in. They were attempting to access things on the Internet because they thought it might help them with their tasks. In short, these AIs were trying to do many different things to help them in the situation they found themselves.

Geoff Bennett:

So when you warn that the next year or two could be crunch time for humanity -- that was the phrase that you used -- walk us through that chain. What specifically has to happen for that to be true?

Jacob Coxon:

So the main thing that has to happen is for us to enter this phase called recursive self-improvement.

So, RSI, or recursive self-improvement, is the idea that we could take the discipline of AI research as currently practiced and automate it with the AIs we're building.

So we could let the AIs make themselves smarter. And we have already made strides towards doing that. So, right now, a lot of the research that researchers at these labs do is via AIs that carry out certain tasks or experiments.

Currently, we still need the human to provide taste or judgment, but plausibly within a year, certainly, within two years, we're on track to also automating away this nebulous human quality of judgment so that the AIs can rapidly make themselves smarter.

And, at that point, things could go very, very fast. So we get new generations of AIs very quickly without necessarily having the corresponding tech to understand what they're doing and control what they're doing.

Geoff Bennett:

So, what would a slowdown actually look like? What kind of intervention are you calling for?

Jacob Coxon:

There are lots of proposals, and I really am not an expert in all of the political things that would have to happen for a slowdown to happen.

All I know is that going too fast could be lethal. But I do think that a race with China could be very bad and that we should be taking pretty rapid actions for international coordination. And this sort of desire is also echoed in the public blog posts of people like OpenAI's chief scientist, like Dario Amodei.

They're all mentioning the fact that we need international coordination. I will say that I do think a lot of the company leadership are privately quite skeptical of being able to achieve this level of international cooperation, which is part of why I resigned, because I think a race like that could be catastrophic.

Geoff Bennett:

Is there anything preventing OpenAI, Anthropic, any other leading AI lab from announcing tomorrow that we will not build another self-improving system until we know how to control it?

I ask the question because the White House AI adviser, as you well know, has made the case that, if these AI leaders are so concerned, they can implement guardrails voluntarily and unilaterally.

Jacob Coxon:

Yes, I think the phrase implement guardrails sort of implies that we know what to do, and it might be just a couple of months to get the tech right, implement the guardrails, and then carry on business as usual.

I think maybe the framing of we're trying to solve a new scientific problem of, can we figure out how to grow these AIs in a safe manner, which could take a pretty long period of time, and we also might want to phase it in pretty slowly. And, in this case, again, the main issue for slowing, stopping right now until we're certain is that that could take quite some time.

And, at that point, other places building this technology, say, China, could shoot right past. So the issue is, a race is kind of on the cards right now. Everyone is aware that we have to get there before our rivals, unless there's some sort of coordination.

Geoff Bennett:

What's the strongest evidence that you could point to publicly, something a skeptical viewer could weigh, that supports the warning that you're making?

Jacob Coxon:

Yes.

I think the first point that I make a lot is just a kind of it's an appeal to authority kind of point, which is people from many sides of the playing field, like Geoff Hinton, Yoshua Bengio, the founders of M.L. theory -- they're academics, they're not in the labs -- also, the lab CEOs, the lab researchers, also people like Elon Musk, they all think this is very dangerous.

I could point to other names like Stephen Hawking. Like, a lot of people have raised this warning. And then to go into precise concrete matters, I think look at the report from the Hugging Face attack. Think to yourselves, like, is this -- how comfortable are you with what happened there? And how bad could that look if the AIs get better?

And then I think there was also a public letter from some mathematicians because AI has recently moved to the level of a professional mathematician. So I think, as we all find our own areas of expertise starting to be competed with by AIs, we will start to maybe take the threat more seriously as we see it impinging on our on our daily life.

Geoff Bennett:

And, Jacob, as you well know, there has been a remarkable convergence in recent days that was kick-started by your resignation.

We have seen, of course, public warnings from other AI leaders. There have been efforts on Capitol Hill to introduce new legislation. I ask the question because there's been so much speculation. Are you in any way part of a coordinated effort to introduce these warnings into the public debate?

Jacob Coxon:

Absolutely not.

And I can maybe see how this happened. From my perspective, the last week has been absolutely insane. I tweeted private thoughts that contradict the views of a lot of my close friends, and my private life has been kind of upended by this.

And, yes, I can just say that this was my resignation. I believe this. A lot of people believe this. This isn't some crazy coordinated scheme. One thing I also want to add is this has raised awareness of the risks to a lot of people. And I'm quite worried that the corresponding benefits are also maybe not receiving adequate airtime.

So I would like to emphasize that I do believe that the benefits from this technology in domains like health are tremendous, and we should also think about the possibilities for what this tech will do if it's done safely.

Geoff Bennett:

What did you think would happen after you sent that tweet about your resignation?

Jacob Coxon:

I thought it would maybe go viral among my community of AI researchers, people thinking about AI safety. I thought it would maybe trigger some discussion. I thought a lot of people would disagree.

I did not it to -- I did not expect messages from old high school friends asking me about if the robots were going to kill us. That really wasn't on the cards when I sent that tweet.

Geoff Bennett:

Jacob Coxon, thanks again for joining us this evening. We appreciate it.

Jacob Coxon:

Thank you.

Listen to this Segment