The Economist.
It's been another scary few weeks in the world of artificial intelligence.
If the autonomous swarms of agents from open AI that collectively conspired, deceived
and hacked their way into hugging face weren't worrying enough,
a scientist at Anthropic who leads the work to align AI models said,
we really do earnestly believe AI could kill all humans.
I personally think it's more than 10% within the next decade.
The bosses of the labs have also chimed in, asking the American
government to help them slow down or pace the frontier of AI research.
So far, though, the American president doesn't seem keen.
What to make of all this loud and frankly confusing public conversation?
Have humans lost control of our increasingly powerful AI models?
Are we about to?
One way to avoid resignation or despair is to arm yourself with science.
To understand what AI models could be capable of and how they might behave, you need to get into their heads.
This is the goal of a growing number of scientists and philosophers in the emerging field of model psychology.
A few hundred of them have come together here in Berkeley in California,
not far from the Silicon Valley labs where the frontier models are designed and built.
They're here to share their research and think about the most urgent questions they want to investigate next.
By peering inside and better characterizing digital minds, can scientists build a path to safer AI?