2026-06-26
19 分钟So here's the big research bet that all the labs are making.
They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments,
then we will have basically built AGI because this kind of training will have created a kind of problem solving agent,
the kind of thing that can make progress on opening
the tasks for weeks on end in the face of errors and mistakes and ambiguity.
And the people who are optimistic about this vision will say that all these things
that we talk about as the fundamental deficits in the current training paradigm, for example,
the data inefficiency of these models or the fact that they lack continual learning,
these things can just be steamrolled if we just scale training more.
And the same way that all the fundamental research problems in natural language processing collapsed
when we just threw enough compute into LLMs.
So in the previous essay, I talked about how these models are one one millionth of sample efficient as humans.
And the people who are in favor of the current training paradigm will say, look,
that might be true, but this is only true during training.
And training is this one time cost that is amortized across billions of sessions that a model will experience.
And what really matters is how smart and general and sample efficient the model is during a session.
And this has clearly been improving as we've been doing more RL training.
AI agents are able to solve more and more ambitious problems over longer and longer time spans.
Anybody who has used these models for coding knows that.
Similarly, people would say, look, continue learning.