Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture recording from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on September 29, 2025, traces the evolution of verification methods for large language model outputs across four research papers. It covers OpenAI's "Training Verifiers to Solve Math Word Problems," which introduced the GSM8K dataset and outcome-based reward models, and "Let's Verify Step by Step," which compares outcome-supervised and process-supervised reward models using the PRM800K dataset of human-labeled reasoning steps. The lecture also covers Math-Shepherd, which automates step-level annotation without human labels, and Weaver, a Stanford paper that combines ensembles of weak verifiers, including reward models and LLM judges, to close the generation-verification gap. Topics include majority voting and self-consistency baselines, credit assignment in process versus outcome supervision, reward hacking, and using trained verifiers as reward signals for reinforcement learning fine-tuning. Speaker Bio: Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.
Want to dive deeper? This curriculum is covered in the following online courses:
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents
- Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html
Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/
Azalia Mirhoseini
Assistant Professor of Computer Science, Stanford University
View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA
Video Summary:
This lecture recording from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on September 29, 2025, traces the evolution of verification methods for large language model outputs across four research papers. It covers OpenAI's "Training Verifiers to Solve Math Word Problems," which introduced the GSM8K dataset and outcome-based reward models, and "Let's Verify Step by Step," which compares outcome-supervised and process-supervised reward models using the PRM800K dataset of human-labeled reasoning steps. The lecture also covers Math-Shepherd, which automates step-level annotation without human labels, and Weaver, a Stanford paper that combines ensembles of weak verifiers, including reward models and LLM judges, to close the generation-verification gap. Topics include majority voting and self-consistency baselines, credit assignment in process versus outcome supervision, reward hacking, and using trained verifiers as reward signals for reinforcement learning fine-tuning.
Speaker Bio:
Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.