Noam Brown – Agent swarms, alignment, & recursive self-improvement

诺亚姆·布朗——代理集群、协同一致与递归自我提升

Dwarkesh Podcast

2026-09-17

1 小时 20 分钟
PDF

单集简介 ...

New episode with Noam Brown. We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. Watch on YouTube; read the transcript. Sponsors * Jane Street has been interested in AI for a lot longer than you’d think, and not just for trading. In 2011, a full year before AlexNet and over a decade before ChatGPT launched, they hosted the first FOOM Debate between Eliezer Yudkowsky and Robin Hanson on whether AI would lead to an intelligence explosion. Now Jane Street is revisiting the question with a new panel: Daniel Kokotajlo, Ege Erdil, Ryan Greenblatt, and Jaime Sevilla, hosted by Ron Minsky in San Francisco this October. I expect it to be a truly excellent conversation. Register at janestreet.com/dwarkesh * Grok Bot has made handing off work super easy. It runs on its own cloud computer, where it installs the tools it needs to handle tasks end-to-end. For the podcast, we use Grok Bot to help produce our videos. You may have noticed that our ads feature animations of real websites. Getting these pixel-perfect used to mean running a convoluted, multi-step workflow ourselves. Now we just let Grok Bot handle it. Best of all, Grok Bot has learned all of our specs and preferences, so we don’t have to redescribe the task each time! Try Grok Bot for yourself at x.ai/bot * Antithesis gives you the confidence of a giant test suite without actually having to write one. Say you’re doing a major backend refactor: building enough tests to trust it could take weeks. Antithesis solves this by running your software through countless simulated worlds, injecting faults and hunting for failures. On any PR, you can turn a dial to decide exactly how much testing you want. And because every run is fully deterministic, agents can branch off the moment a bug appears, rewind it, inspect memory, and replay it, all while the original test keeps running. Learn more at antithesis.com/dwarkesh Timestamps (00:00:00) – Multi-agent and Navier-Stokes (00:15:28) – How will AI firms work? (00:22:02) – What math progress tells us about recursive self improvement (00:40:22) – Hugging Face and alignment (01:01:18) – The internal/external model gap (01:08:34) – Chain of thought is degrading (01:14:12) – How will we know when alignment is solved? This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com
更多

单集文稿 ...

  • Today, I'm chatting with Noam Brown, who is a researcher at OpenAI.

  • He was one of the foundational contributors to what became O1 and the reasoning models,

  • and now he's working on multi-agent systems.

  • Speaking of which, you guys announced last week that you solved one of the million price problems

  • with a system of 10,000 different AI agents that spent 130 billion tokens over 88 hours.

  • One of the reasons I'm interested in talking

  • to you is I think you were in the first people maybe two or three years ago who was...

  • Thinking about how the reasoning models would allow us to see into the future,

  • because if you scale up inference compute,

  • you can see what the base capabilities of the models will be a few years in the future.

  • And I feel like you're in a similar position now to help us understand what future capabilities will look like,

  • given the enormous scaling of agent sizes that we can do right now.

  • So the way I think about it...

  • When you plot the performance of these reasoning models with test time compute on the x-axis and performance

  • on basically any reasoning benchmark on the y-axis, you see a very clear pattern

  • where the longer these models take to think about their answer, the better they do.

  • And this is like a very natural thing.

  • It's the same thing with people.

  • If you're taking the SATs, you have five minutes to go through the entire exam, you're not going to do very well.

  • If you have five hours, you're probably going to do a lot better.