Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Video Summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 10, 2025, covers train-time scaling and scaling reinforcement learning through three papers. STaR, the Self-Taught Reasoner, bootstraps reasoning chains through rationalization and filtering by answer correctness. DeepSeekMath introduces Group Relative Policy Optimization, a memory-efficient alternative to PPO, and shows gains from training on curated math data. DAPO addresses entropy collapse and training instability in reinforcement learning on long chain-of-thought reasoning through techniques including asymmetric clipping and dynamic sampling. Using the AIME math benchmark, the lecture traces how these methods let smaller models match the accuracy of much larger systems, and it closes with open questions on why majority-at-K accuracy improves while pass-at-K does not. Speaker Bio: Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet.
Want to dive deeper? This curriculum is covered in the following online courses:
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents
- Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html
Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/
View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA
Aakanksha Chowdhery
Adjunct Professor of Computer Science, Stanford University
Video Summary:
This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 10, 2025, covers train-time scaling and scaling reinforcement learning through three papers. STaR, the Self-Taught Reasoner, bootstraps reasoning chains through rationalization and filtering by answer correctness. DeepSeekMath introduces Group Relative Policy Optimization, a memory-efficient alternative to PPO, and shows gains from training on curated math data. DAPO addresses entropy collapse and training instability in reinforcement learning on long chain-of-thought reasoning through techniques including asymmetric clipping and dynamic sampling. Using the AIME math benchmark, the lecture traces how these methods let smaller models match the accuracy of much larger systems, and it closes with open questions on why majority-at-K accuracy improves while pass-at-K does not.
Speaker Bio:
Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet.