Want to dive deeper? This curriculum is covered in the following online courses: - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents - Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on November 17, 2025, covers methods for evaluating AI agents on long-horizon and economically valuable tasks. It reviews METR's time-horizon methodology, which measures the task duration models complete at 50 and 80 percent reliability across the HCAST, SWE-bench, and RE-Bench suites, showing that model capability has roughly doubled every seven months, from seconds for GPT-2 in 2019 to nearly an hour for Claude 3.7 Sonnet in 2025. It also covers GDPval, OpenAI's benchmark that scores model output against work from industry professionals across 44 occupations and nine sectors, where win rates rose from 12.4 percent for GPT-4o to 47.6 percent for Claude Opus 4.1. Stanford's Deep Scholar Bench, also discussed, tests whether models can generate literature review sections for academic papers by scoring knowledge synthesis, retrieval quality, and citation verifiability. The lecture closes by detailing recurring agent failure modes, including poor planning, incorrect tool selection, premature task abandonment, and repetitive action loops. Speaker Bio: Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet.
Want to dive deeper? This curriculum is covered in the following online courses:
- XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents
- Agentic AI professional education program: https://learn.stanford.edu/agentic-ai-2026.html
Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/
View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA
Video Summary:
This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on November 17, 2025, covers methods for evaluating AI agents on long-horizon and economically valuable tasks. It reviews METR's time-horizon methodology, which measures the task duration models complete at 50 and 80 percent reliability across the HCAST, SWE-bench, and RE-Bench suites, showing that model capability has roughly doubled every seven months, from seconds for GPT-2 in 2019 to nearly an hour for Claude 3.7 Sonnet in 2025. It also covers GDPval, OpenAI's benchmark that scores model output against work from industry professionals across 44 occupations and nine sectors, where win rates rose from 12.4 percent for GPT-4o to 47.6 percent for Claude Opus 4.1. Stanford's Deep Scholar Bench, also discussed, tests whether models can generate literature review sections for academic papers by scoring knowledge synthesis, retrieval quality, and citation verifiability. The lecture closes by detailing recurring agent failure modes, including poor planning, incorrect tool selection, premature task abandonment, and repetitive action loops.
Speaker Bio:
Aakanksha Chowdhery
Adjunct Professor of Computer Science, Stanford University
Aakanksha is pushing the frontier of agentic LLMs by leveraging RL techniques to enable autonomous self-improving agents, especially in software engineering at the startup Reflection AI. At Stanford, she is co-teaching CS329A (Self-Improving AI agents) in Fall/Winter 2025 and is the Program Chair for MLSys 2026. Before this, she was the technical Lead of 540B PaLM model and lead researcher in Gemini at Google in pre-training, scaling, and finetuning of Large Language Models. She was also a core contributor in PaLM-E, MedPaLM, and Pathways project at Google. Prior to joining Google, She was technical lead for several interdisciplinary research initiatives at Microsoft Research and Princeton University across machine learning and distributed systems. She completed my PhD in Electrical Engineering from Stanford University and was awarded the Paul Baran Marconi Young Scholar Award for the outstanding scientific contributions of her dissertation in the field of communications and Internet.