The next big breakthrough will be AIs learning on the job

下一个重大突破将是人工智能在工作中学习。

Dwarkesh Podcast

2026-06-26

19 分钟
PDF

单集简介 ...

Read it here. Thanks to Mercury for sponsoring this essay. Mercury has automated basically my entire bill pay process for my business. I just give contractors a dedicated email address, and when they send an invoice, Mercury automatically creates a draft payment for me to review. I no longer have to hunt through my inbox for invoices or deal with messy spreadsheets to track my bills. Mercury handles it all. Learn more at mercury.com Timestamps: (00:00:00) – The big research bet the labs are making (00:02:12) – Grindability is just as important as verifiability (00:06:10) – Will RLVR alone generalize? (00:08:41) – Getting the learning back to the weights (00:15:22) – Dreaming (00:17:23) – What 2027 looks like Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
更多

单集文稿 ...

  • So here's the big research bet that all the labs are making.

  • They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments,

  • then we will have basically built AGI because this kind of training will have created a kind of problem solving agent,

  • the kind of thing that can make progress on opening

  • the tasks for weeks on end in the face of errors and mistakes and ambiguity.

  • And the people who are optimistic about this vision will say that all these things

  • that we talk about as the fundamental deficits in the current training paradigm, for example,

  • the data inefficiency of these models or the fact that they lack continual learning,

  • these things can just be steamrolled if we just scale training more.

  • And the same way that all the fundamental research problems in natural language processing collapsed

  • when we just threw enough compute into LLMs.

  • So in the previous essay, I talked about how these models are one one millionth of sample efficient as humans.

  • And the people who are in favor of the current training paradigm will say, look,

  • that might be true, but this is only true during training.

  • And training is this one time cost that is amortized across billions of sessions that a model will experience.

  • And what really matters is how smart and general and sample efficient the model is during a session.

  • And this has clearly been improving as we've been doing more RL training.

  • AI agents are able to solve more and more ambitious problems over longer and longer time spans.

  • Anybody who has used these models for coding knows that.

  • Similarly, people would say, look, continue learning.