Azalia Mirhoseini
Assistant Professor of Computer Science
Bio
Azalia Mirhoseini is an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a research lab focused on developing scalable and self-improving AI systems. She is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on Gemini and Claude, among other projects. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading frontier AI models; AlphaChip, a reinforcement learning approach for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.
Program Affiliations
-
Stanford SystemX Alliance
2025-26 Courses
- Self Improving AI Agents
CS 329A (Aut) -
Independent Studies (12)
- Advanced Reading and Research
CS 499 (Aut, Win, Spr, Sum) - Advanced Reading and Research
CS 499P (Aut, Win, Spr, Sum) - Curricular Practical Training
CS 390A (Spr, Sum) - Curricular Practical Training
CS 390B (Sum) - Independent Project
CS 399 (Aut, Win, Spr, Sum) - Independent Project
CS 399P (Aut, Win, Spr, Sum) - Independent Work
CS 199 (Aut, Win, Spr, Sum) - Independent Work
CS 199P (Aut, Win, Spr, Sum) - Master's Research
CME 291 (Aut, Win, Spr, Sum) - Senior Project
CS 191 (Aut, Win, Spr) - Supervised Undergraduate Research
CS 195 (Aut, Win, Spr, Sum) - Writing Intensive Senior Research Project
CS 191W (Aut, Win, Spr)
- Advanced Reading and Research
-
Prior Year Courses
2024-25 Courses
- Self Improving AI Agents
CS 329A (Win) - Systems for Machine Learning
CS 229S (Aut)
2023-24 Courses
- Systems for Machine Learning
CS 229S (Aut)
- Self Improving AI Agents
Stanford Advisees
-
Doctoral Dissertation Reader (AC)
Avanika Narayan, Qizheng Zhang -
Orals Evaluator
Avanika Narayan, Zhiqiang Xie -
Doctoral Dissertation Advisor (AC)
Jon Saad-Falcon -
Master's Program Advisor
Armaan Abraham, Gia Ancone, Sarah Barragan, Gabriel Bo, William Briger, Jett Carruth, Manat Kaur, Brandon Liu, Nora Menon, Sid Potti, Arnuv Tandon, Andy Wang, Berk Yalcinkaya, Daniel Yang, Shirley Yu -
Doctoral Dissertation Co-Advisor (AC)
Hermann Kumbong, Jacky Kwok, Yuzhen Mao -
Doctoral (Program)
Bradley Brown, Simon Guo, Jordan Juravsky, Anne Ouyang, Jon Saad-Falcon
All Publications
-
Training of physical neural networks.
Nature
2025; 645 (8079): 53-61
Abstract
Physical neural networks (PNNs) are a class of neural-like networks that make use of analogue physical systems to perform computations. Although at present confined to small-scale laboratory demonstrations, PNNs could one day transform how artificial intelligence (AI) calculations are performed. Could we train AI models many orders of magnitude larger than present ones? Could we perform model inference locally and privately on edge devices? Research over the past few years has shown that the answer to these questions is probably "yes, with enough research". Because PNNs can make use of analogue physical computations more directly, flexibly and opportunistically than traditional computing hardware, they could change what is possible and practical for AI systems. To do this, however, will require notable progress, rethinking both how AI models work and how they are trained-primarily by considering the problems through the constraints of the underlying hardware physics. To train PNNs, backpropagation-based and backpropagation-free approaches are now being explored. These methods have various trade-offs and, so far, no method has been shown to scale to large models with the same performance as the backpropagation algorithm widely used in deep learning today. However, this challenge has been rapidly changing and a diverse ecosystem of training techniques provides clues for how PNNs may one day be used to create both more efficient and larger-scale realizations of present-scale AI models.
View details for DOI 10.1038/s41586-025-09384-2
View details for PubMedID 40903603
View details for PubMedCentralID 8791835
-
An Architecture Search Framework for Inference-Time Techniques
edited by Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 52475-52507
View details for Web of Science ID 001693162000292
-
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
edited by Lim, J., Song, S., Park, H. W.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 3200-3217
View details for Web of Science ID 001676333500151
-
Think, Prune, Train: Can Small Models Teach Themselves to Reason?
IEEE COMPUTER SOC. 2025: 195-203
View details for DOI 10.1109/ICLAD65226.2025.00012
View details for Web of Science ID 001562425500026
-
How Do Large Language Monkeys Get Their Power (Laws)?
edited by Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 53132-53176
View details for Web of Science ID 001693164800017
-
KernelBench: Can LLMs Write Efficient GPU Kernels?
edited by Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 47356-47415
View details for Web of Science ID 001693162000094
-
Addendum: A graph placement methodology for fast chip design.
Nature
2024
View details for DOI 10.1038/s41586-024-08032-5
View details for PubMedID 39327495
-
Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification
edited by Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S.
NEURAL INFORMATION PROCESSING SYSTEMS (NIPS). 2023
View details for Web of Science ID 001220818804023
-
A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators
edited by Falsafi, B., Ferdman, M., Lu, S., Weinisch, T.
ASSOC COMPUTING MACHINERY. 2022: 27-42
View details for DOI 10.1145/3503222.3507767
View details for Web of Science ID 000810486300003
-
A graph placement methodology for fast chip design.
Nature
2021; 594 (7862): 207-212
Abstract
Chip floorplanning is the engineering task of designing the physical layout of a computer chip. Despite five decades of research1, chip floorplanning has defied automation, requiring months of intense effort by physical design engineers to produce manufacturable layouts. Here we present a deepreinforcementlearning approach to chip floorplanning. In under six hours, our method automatically generates chip floorplans that are superior or comparable to those produced by humans in all key metrics, including power consumption, performance and chip area. To achieve this, we pose chip floorplanning as a reinforcementlearning problem, and develop an edge-based graph convolutional neural network architecture capable of learning rich and transferable representations of the chip. As a result, our method utilizes past experience to become better and faster at solving new instances of the problem, allowing chip design to be performed by artificial agents with more experience than any human designer. Our method was used to design the next generation of Google's artificial intelligence (AI) accelerators, and has the potential to save thousands of hours of human effort for each new generation. Finally, we believe that more powerful AI-designed hardware will fuel advances in AI, creating a symbiotic relationship between the two fields.
View details for DOI 10.1038/s41586-021-03544-w
View details for PubMedID 34108699