2026-09-17
1 小时 20 分钟Today, I'm chatting with Noam Brown, who is a researcher at OpenAI.
He was one of the foundational contributors to what became O1 and the reasoning models,
and now he's working on multi-agent systems.
Speaking of which, you guys announced last week that you solved one of the million price problems
with a system of 10,000 different AI agents that spent 130 billion tokens over 88 hours.
One of the reasons I'm interested in talking
to you is I think you were in the first people maybe two or three years ago who was...
Thinking about how the reasoning models would allow us to see into the future,
because if you scale up inference compute,
you can see what the base capabilities of the models will be a few years in the future.
And I feel like you're in a similar position now to help us understand what future capabilities will look like,
given the enormous scaling of agent sizes that we can do right now.
So the way I think about it...
When you plot the performance of these reasoning models with test time compute on the x-axis and performance
on basically any reasoning benchmark on the y-axis, you see a very clear pattern
where the longer these models take to think about their answer, the better they do.
And this is like a very natural thing.
It's the same thing with people.
If you're taking the SATs, you have five minutes to go through the entire exam, you're not going to do very well.
If you have five hours, you're probably going to do a lot better.