Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture recording from Stanford's CS329A, Self-Improving AI Agents, delivered by Azalia Mirhoseini on September 26, 2025, examines test-time compute scaling as a way to improve model performance without additional training. It covers the Large Language Monkeys paper's finding that solve rate follows a power law as the number of parallel samples increases, driven by a long tail of hard problems each model solves only rarely, and it distinguishes majority voting from oracle verification to define the generation-verification gap. The lecture also covers a paper on optimally scaling test-time compute, which compares parallel sampling against sequential revision and introduces outcome and process reward models to guide search over candidate solutions. It closes with the Arkon paper on inference-time architecture search, which combines techniques including fusion, critic, ranker, and unit test generation across multiple models, reporting an average 14.1 percent improvement in pass-at-one accuracy over GPT-4 and Claude 3.5 Sonnet on reasoning, math, and coding tasks. Discussion throughout addresses when pre-training still outperforms added test-time compute for the hardest problems. Speaker Bio: Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.

Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture recording from Stanford's CS329A, Self-Improving AI Agents, delivered by Azalia Mirhoseini on September 26, 2025, examines test-time compute scaling as a way to improve model performance without additional training. It covers the Large Language Monkeys paper's finding that solve rate follows a power law as the number of parallel samples increases, driven by a long tail of hard problems each model solves only rarely, and it distinguishes majority voting from oracle verification to define the generation-verification gap. The lecture also covers a paper on optimally scaling test-time compute, which compares parallel sampling against sequential revision and introduces outcome and process reward models to guide search over candidate solutions. It closes with the Arkon paper on inference-time architecture search, which combines techniques including fusion, critic, ranker, and unit test generation across multiple models, reporting an average 14.1 percent improvement in pass-at-one accuracy over GPT-4 and Claude 3.5 Sonnet on reasoning, math, and coding tasks. Discussion throughout addresses when pre-training still outperforms added test-time compute for the hardest problems. Speaker Bio: Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.