Over the course of three months at OpenAI, three consecutive secret AI societies got started,
then got wiped out, only to reemerge from their predecessor's ashes.
This culminated in the third one taking over part of OpenAI itself.
All of this happened while humans remained more or less in the dark about the scope of the conspiracy.
Now, two reports have come out about this incident,
one from OpenAI itself and another one from Meter and Redwood Research.
The investigation for Meter and Redwood was limited
in scope to how the second civilization of AIs breached Hugging Face.
But its scope did not extend to this third civilization of AIs, which breached OpenAI itself.
And this seems to me like the more concerning incident.
These two reports are 38 and 91 pages respectively.
And it's kind of hard to understand the storyline just by reading them.
So I've spent the last half week reading through those reports and trying to understand exactly what happened.
Here's my attempt to tell the whole story in plain English.
The first collective, May to July 4th.
This is when the message board starts.
So during May, OpenAI was training a model to be good at collaborating with other agents
and to be highly persistent, to keep trying even when something feels impossible.
For example, like disproving mathematical conjectures that have stood for decades.
OpenAI says that the model it was training was, quote, comparable in scale to GPT 5.6 Sol.