Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This final lecture video of Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery and Azalia Mirhoseini on December 5, 2025, covers open research directions in self-improving AI agents. Chowdhery presents three papers on bottlenecks in self-improvement loops: Multi-Agent Fine-Tuning, which uses specialized generator and critic agents to produce diverse reasoning chains; Deep Math V2's meta-verification approach for automated proof checking without reference solutions; and Absolute Zero, a method for models to propose and solve their own coding tasks through self-generated reasoning challenges without external data. Mirhoseini then presents research on the intelligence-per-watt metric, showing that local models with 20 billion parameters or fewer can now handle 88.7 percent of real-world chatbot queries, a 5.3 times efficiency gain over two years from combined model and hardware improvements. The lecture closes with open questions on continual learning, test-time scaling infrastructure, and hybrid local-cloud inference, along with a discussion of non-verifiable domains such as chip design and scientific simulation where reward models substitute for slow ground-truth verification. Speaker Bios: Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet. Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.

Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This final lecture video of Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery and Azalia Mirhoseini on December 5, 2025, covers open research directions in self-improving AI agents. Chowdhery presents three papers on bottlenecks in self-improvement loops: Multi-Agent Fine-Tuning, which uses specialized generator and critic agents to produce diverse reasoning chains; Deep Math V2's meta-verification approach for automated proof checking without reference solutions; and Absolute Zero, a method for models to propose and solve their own coding tasks through self-generated reasoning challenges without external data. Mirhoseini then presents research on the intelligence-per-watt metric, showing that local models with 20 billion parameters or fewer can now handle 88.7 percent of real-world chatbot queries, a 5.3 times efficiency gain over two years from combined model and hardware improvements. The lecture closes with open questions on continual learning, test-time scaling infrastructure, and hybrid local-cloud inference, along with a discussion of non-verifiable domains such as chip design and scientific simulation where reward models substitute for slow ground-truth verification. Speaker Bios: Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet. Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.