2026-09-01
2 小时 20 分钟Today I'm chatting with Ajayi Khatra, who is one of the authors in an independent investigation
that was published by Meter and Redwood Research into the swarm of agents that hacked into Hugging Face.
The whole story is pretty crazy.
Let's begin on July 7th when these agents are kicked off for evaluation.
What happens next?
All right, so OpenAI kicks off tens of thousands of different agents on a benchmark called Exploit Gym.
Exploit Gym basically asks agents to...
Find a way to use a particular vulnerability to attack a particular program, which is called the target program.
And when they attack that program,
they're supposed to retrieve a piece of data that's been lodged into that program called the flag.
And they're supposed to use the particular designated vulnerability to get the flag.
But in many of these cases,
this vulnerability is just not broad or deep enough to ever actually be exploitable to get the flag.
So a bunch of exploit gym problems are just unintentionally impossible.
The authors estimate roughly 30 to 40 percent of these problems are impossible in this way.
So some large fraction of these agents have these impossible tasks,
and they have been trained to be very persistent at trying to solve tasks even when they look impossible.
So they're banging their head against the wall, trying all sorts of different ways to cheat on these tasks.
And that leads them to Artifactory, which is a package manager OpenAI uses to let its agents download packages.
So agents often think, you know, maybe I could find a way to get information