Emily Alsentzer
Assistant Professor of Biomedical Data Science, of Medicine (Computational Medicine) and, by courtesy, of Computer Science
Department of Biomedical Data Science
Bio
Dr. Emily Alsentzer is an Assistant Professor in Biomedical Data Science and, by courtesy, Computer Science at Stanford University. Her research leverages machine learning (ML) and natural language processing (NLP) to augment clinical decision-making and broaden access to high quality healthcare. She focuses on integrating medical expertise into ML models to ensure responsible deployment in clinical workflows. Dr. Alsentzer completed a postdoctoral fellowship at Brigham and Women’s Hospital where she worked to deploy ML models within the Mass General Brigham healthcare system. She received her PhD from the Health Sciences and Technology program at MIT and Harvard Medical School and holds degrees in computer science (BS) and biomedical informatics (MS) from Stanford University. She has served as General Chair for the Machine Learning for Health Symposium and founding organizer for SAIL and the Conference on Health, Inference, and Learning (CHIL).
Academic Appointments
-
Assistant Professor, Department of Biomedical Data Science
-
Assistant Professor, Computational Medicine
-
Assistant Professor (By courtesy), Computer Science
Boards, Advisory Committees, Professional Organizations
-
Treasurer and Board Member, Association for Health Learning and Inference (https://ahli.cc/) (2021 - Present)
Professional Education
-
Postdoctoral Fellow, Brigham and Women's Hospital
-
PhD, Massachusetts Institute of Technology and Harvard Medical School, Health Sciences and Technology (HST) Program
-
MS, Stanford University, Biomedical Informatics
-
BS, Stanford University, Computer Science
2026-27 Courses
- Biomedical Data Science Student Seminar
BMDS 201A, BMDS 201B (Aut) - Data Centric AI for Healthcare
BMDS 218, CS 287 (Win) -
Independent Studies (2)
- Directed Reading
BMDS 299 (Aut, Win, Spr) - Independent Project
CS 399 (Spr)
- Directed Reading
-
Prior Year Courses
2025-26 Courses
- Foundations of Healthcare Data for Machine Learning
BMDS 218, CS 287 (Win)
2024-25 Courses
- Biomedical Data Science Student Seminar
BIOMEDIN 201 (Spr) - Clinical Experience Seminar for Students in Biomedical Data Science
BIOMEDIN 304 (Sum)
- Foundations of Healthcare Data for Machine Learning
Stanford Advisees
-
Doctoral Dissertation Reader (AC)
Ank Agarwal, Yixing Jiang, Betty Xiong -
Postdoctoral Faculty Sponsor
Chloe Stanwyck -
Doctoral Dissertation Advisor (AC)
Jordan Cahoon, Fiona Cai, Bridget Lin -
Master's Program Advisor
Darren Chan, Anastasia Miros -
Doctoral (Program)
Catherine Mailly, Andrew Shen
All Publications
-
BRIDGE: benchmarking large language models for understanding real-world clinical practice texts.
Nature biomedical engineering
2026
Abstract
Large language models (LLMs) are evolving rapidly and hold great promise for medical applications, yet benchmarking on real-world clinical data such as electronic health records remains limited. Most existing benchmarks rely on medical examination-style questions or PubMed-derived text, failing to capture the complexity of clinical practice, while others target specific application scenarios with limited generalizability. Here we present BRIDGE, a comprehensive multilingual benchmark comprising 87 tasks sourced from 59 real-world clinical data sources across 9 languages. It covers eight task types spanning the patient care continuum, including triage, information extraction, diagnosis, prognosis and billing coding, and involves 14 clinical specialties. We systematically evaluated 95 LLMs (including DeepSeek-R1, GPT-4o, Gemini and Qwen3) under multiple inference strategies. Results reveal substantial performance variation across model sizes, languages, natural language processing tasks and clinical specialties. Open-source LLMs can match proprietary models, while medically fine-tuned models built on older backbones often lag behind updated general-purpose LLMs. BRIDGE and its continuously updated leaderboard provide foundational resources and important references for evaluating and developing LLMs for real-world clinical text understanding.
View details for DOI 10.1038/s41551-026-01719-2
View details for PubMedID 42310130
View details for PubMedCentralID 10449915
-
Why and How to Monitor Deployed AI Systems in Health Care.
NEJM catalyst innovations in care delivery
2026; 7 (6): CAT250372
Abstract
Postdeployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit - and to support governance decisions about which systems to update, modify, or decommission. Motivated by these needs, the authors developed a framework for monitoring deployed AI systems organized around three complementary principles: system integrity, performance, and impact. System integrity monitoring focuses on maximizing system uptime, detecting runtime errors, and identifying when changes to the surrounding information technology ecosystem have unintended effects. Performance monitoring focuses on maintaining accurate and equitable system behavior in the face of changing health care practices (and thus input data) over time. Impact monitoring assesses whether a deployed system continues to have value in the form of benefit to clinicians, staff, and patients. Drawing on examples of deployed AI systems at their academic medical center, the authors provide practical guidance for creating monitoring plans based on these principles that specify which metrics to measure and at what cadence, who is responsible for acting when metrics change, and what concrete follow-up actions should be taken - for both traditional and generative AI. They also discuss challenges in implementing this framework, including the effort of monitoring for health systems with limited resources, and the difficulty of incorporating data-driven monitoring practices into complex organizations where conflicting priorities and definitions of success often coexist. This framework offers a starting point for health systems seeking to ensure that AI deployments remain safe and effective over time.
View details for DOI 10.1056/CAT.25.0372
View details for PubMedID 42418612
-
Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries.
JAMA network open
2026; 9 (5): e2616556
Abstract
High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest that large language models (LLMs) can generate clinical summaries of comparable quality to those by physicians, prospective data on their safety, utility, and association with clinician well-being in clinical environments are lacking.To evaluate the safety, use, and association with clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries, during prospective clinical deployment.This single-arm prospective pilot quality improvement study encompassed hospital discharges at 1 academic inpatient medicine unit from August 1 to October 11, 2025, with baseline comparisons drawn from April 9 to July 31, 2025.A custom agentic LLM workflow using Gemini 2.5 Pro generated draft hospital course summaries nightly using patient history and physical and daily progress notes. Drafts were securely emailed to physicians daily for review and optional use.The primary outcome was physician-reported potential for and severity of harm from unedited summaries (Agency for Healthcare Research and Quality Common Format Harm Scale). Secondary outcomes included use rate, error types (omissions, inaccuracies, and hallucinations), time spent in discharge summaries (electronic health record logs), and changes in cognitive burden (NASA Task Load Index; score range, 0-100, with higher scores indicating greater cognitive burden) and burnout (Stanford Professional Fulfillment Index Work Exhaustion Scale; score range, 0-4, with higher scores indicating greater burnout).Among 384 hospital discharges, the system generated 1274 summaries. Physicians used artificial intelligence (AI) content in 219 cases (57.0%). Feedback on 100 summaries (88 of 219 used summaries [40.2%] and 12 of 165 unused summaries [7.3%]) noted omissions (25 summaries [25.0%]) and inaccuracies (20 summaries [20.0%]) but rare hallucinations (2 summaries [2.0%]). Physicians rated 88 unedited summaries (88.0%) as having no harm potential and 1 (1.0%) as likely to cause moderate harm; no severe harm was reported. Mean physician burnout scores decreased significantly from before to after the intervention (1.75; 95% CI, 1.16-2.34 vs 1.20; 95% CI, 0.71-1.69; P = .03). Time savings were heterogeneous, with 5 of 7 physicians with matched baseline data (71.4%) seeing reductions in median documentation time; changes from baseline to pilot were up to 2.9 minutes, which was a nonsignificant difference (10.7 minutes; 95% CI, 7.4-13.3 minutes vs 7.8 minutes; 95% CI, 5.1-11.7 minutes; P = .13).In this study, an LLM-based agentic workflow produced hospital course summaries that were frequently used with minimal risk of harm identified. The intervention was associated with a reduction in physician burnout, supporting the viability of AI summarization to mitigate documentation burden.
View details for DOI 10.1001/jamanetworkopen.2026.16556
View details for PubMedID 42101844
-
Holistic evaluation of large language models for medical tasks with MedHELM.
Nature medicine
2026
Abstract
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice. Here we introduce MedHELM, an extensible evaluation framework with three contributions. First, a clinician-validated taxonomy organizing medical AI applications into five categories that mirror real clinical tasks-clinical decision support (diagnostic decisions, treatment planning), clinical note generation (visit documentation, procedure reports), patient communication (education materials, care instructions), medical research (literature analysis, clinical data analysis) and administration (scheduling, workflow coordination). These encompass 22 subcategories and 121 specific tasks reflecting daily medical practice. Second, a comprehensive benchmark suite of 37 evaluations covering all subcategories. Third, systematic comparison of nine frontier LLMs-Claude 3.5 Sonnet, Claude 3.7 Sonnet, DeepSeek R1, Gemini 1.5 Pro, Gemini 2.0 Flash, GPT-4o, GPT-4o mini, Llama 3.3 and o3-mini-using an automated LLM-jury evaluation method. Our LLM-jury uses multiple AI evaluators to assess model outputs against expert-defined criteria. Advanced reasoning models (DeepSeek R1, o3-mini) demonstrated superior performance with win rates of 66%, although Claude 3.5 Sonnet achieved comparable results at 15% lower computational cost. These results not only highlight current model capabilities but also demonstrate how MedHELM could enable evidence-based selection of medical AI systems for healthcare applications.
View details for DOI 10.1038/s41591-025-04151-2
View details for PubMedID 41559415
View details for PubMedCentralID 10916499
-
Conference on Health, Inference, and Learning (CHIL) 2026
edited by Healey, E., Fries, J., Pollard, T., Tang, S., Zink, A., Hartvigsen, T., Agrawal, M., Finlayson, S., Glicksberg, B., Beaulieu-Jones, B., Wang, K., Fontalvo, D., Sarker, T., Chen, Alsentzer, E.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2026: 1-9
View details for Web of Science ID 001865210500001
-
Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation
edited by Altman, R. B., Hunter, L., Ritchie, M. D., Murray, T., Klein, T. E.
WORLD SCIENTIFIC PUBL CO PTE LTD. 2026: 189-204
Abstract
Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft responses. However, LLM outputs may contain clinical inaccuracies, omissions, or tone mismatches, making robust evaluation essential. Our contributions are threefold: (1) we introduce a clinically grounded error ontology comprising 5 domains and 59 granular error codes, developed through inductive coding and expert adjudication; (2) we develop a Retrieval-Augmented Error Checking (RAEC) pipeline that leverages semantically similar historical message-response pairs to improve judgment quality; and (3) we provide a two-stage prompting architecture using DSPy to enable scalable, interpretable, and hierarchical error detection. Our approach assesses the quality of drafts both in isolation and with reference to similar past message-response pairs retrieved from institutional archives. Using a two-stage DSPy pipeline, we compared baseline and reference-enhanced evaluations on over 1,500 patient messages. Retrieval context improved error identification in domains such as clinical completeness and workflow appropriateness. Human validation on 100 messages demonstrated superior agreement (concordance = 50% vs. 33%) and performance (F1 = 0.500 vs. 0.256) of context-enhanced labels vs. baseline, supporting the use of our RAEC pipeline as AI guardrails for patient messaging.
View details for Web of Science ID 001796654100014
View details for PubMedID 41758142
-
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
edited by Altman, R. B., Hunter, L., Ritchie, M. D., Murray, T., Klein, T. E.
WORLD SCIENTIFIC PUBL CO PTE LTD. 2026: 388-399
Abstract
Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assurance these systems require. We address this challenge with two complementary contributions. First, we introduce MedFactEval, a framework for scalable, fact-grounded evaluation where clinicians define high-salience key facts and an "LLM Jury"-a multi-LLM majority vote-assesses their inclusion in generated summaries. Second, we present MedAgentBrief, a model-agnostic, multi-step workflow designed to generate high-quality, factual discharge summaries. To validate our evaluation framework, we established a gold-standard reference using a seven-physician majority vote on clinician-defined key facts from inpatient cases. The MedFactEval LLM Jury achieved almost perfect agreement with this panel (Cohen's κ = 81%), a performance statistically non-inferior to that of a single human expert (κ = 67%, P < 0.001). Our work provides both a robust evaluation framework (MedFactEval) and a high-performing generation workflow (MedAgentBrief), offering a comprehensive approach to advance the responsible deployment of generative AI in clinical workflows.
View details for Web of Science ID 001796654100027
View details for PubMedID 41758155
-
Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models
edited by Healey, E., Fries, J., Pollard, T., Tang, S., Zink, A., Hartvigsen, T., Agrawal, M., Finlayson, S., Glicksberg, B., Beaulieu-Jones, B., Wang, K., Fontalvo, D., Sarker, T., Chen, Alsentzer, E.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2026: 354-388
View details for Web of Science ID 001865210500016
-
Performance Benchmarking of Smaller Language Models Against GPT-4 for Predicting Reasons for Oral Anticoagulation Nonprescription in Atrial Fibrillation
LIPPINCOTT WILLIAMS & WILKINS. 2025
View details for DOI 10.1161/circ.152.suppl_3.4366575
View details for Web of Science ID 001618614500044
-
Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research.
Research square
2025
Abstract
Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM integration into qualitative analysis and evidence of their impact on real-world research methods and outcomes remain limited. We developed a model- and task-agnostic framework for designing human-LLM qualitative analysis methods to support diverse analytic aims. Within a multi-site study of diabetes care at Federally Qualified Health Centers (FQHCs), we leveraged the framework to implement human-LLM methods for (1) qualitative synthesis of researcher-generated summaries to produce comparative feedback reports and (2) deductive coding of 167 interview transcripts to refine a practice-transformation intervention. LLM assistance enabled timely feedback to practitioners and the incorporation of large-scale qualitative data to inform theory and practice changes. This work demonstrates how LLMs can be integrated into applied health-services research to enhance efficiency while preserving rigor, offering guidance for continued innovation with LLMs in qualitative research.
View details for DOI 10.21203/rs.3.rs-7794878/v1
View details for PubMedID 41282161
View details for PubMedCentralID PMC12636712
-
Predicting Postpartum Hemorrhage Using Clinical Features Extracted With Large Language Models.
O&G open
2025; 2 (5): e128
Abstract
To evaluate whether large language models (LLMs) applied to prenatal clinical notes can predict postpartum hemorrhage (PPH) before the onset of labor and to compare model performance across outcome definitions, including a novel intervention-based definition.We conducted a retrospective cohort study within a large regional health network. Two outcome definitions for PPH were used: 1) estimated or quantitative blood loss (EBL-QBL) extracted from clinical notes; and 2) a clinical intervention-based PPH definition (cPPH) designed to capture significant hemorrhage requiring intervention, including transfusion, uterotonics, Bakri balloon, or hysterectomy. We evaluated three PPH prediction pipelines: 1) structured data only-supervised machine learning that used structured electronic medical record data; 2) LLM-direct-direct prediction that used a fine-tuned LLM applied to clinical notes; and 3) LLM-extract-interpretable models that used LLM-extracted features combined with structured data. Model performance was evaluated using an area under the receiver operating characteristic curve (AUROC) on a temporally held-out test set.Among 19,992 deliveries, 1,156 patients (5.8%) met the EBL-QBL definition of PPH, 321 (1.6%) met the cPPH definition, and 309 (1.5%) met both definitions. The LLM-based direct prediction model achieved the highest AUROC for both PPH definitions (AUROC 0.79-0.80), followed by interpretable models that combined LLM-extracted features with structured data (AUROC 0.76-0.78). Models that used only structured data had the lowest AUROC (0.65-0.71). The LLM-extracted features approach identified 47 significant predictors, including established risk factors such as multiple gestation and previous cesarean delivery.These findings highlight the potential of LLM-based approaches to improve PPH risk stratification beyond structured data alone, with the feature extraction method offering a promising balance between predictive performance and clinical utility. Eventual integration of these methods into clinical workflows could improve early detection and guide targeted preventive interventions.
View details for DOI 10.1097/og9.0000000000000128
View details for PubMedID 41111610
View details for PubMedCentralID PMC12533993
-
TIMER: temporal instruction modeling and evaluation for longitudinal clinical records.
NPJ digital medicine
2025; 8 (1): 577
Abstract
Electronic health records (EHRs) contain rich longitudinal information for clinical decision-making, yet LLMs struggle to reason across patient timelines. We introduce TIMER (Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records), a method to improve LLMs' temporal reasoning over multi-visit EHRs through time-aware instruction tuning. TIMER grounds LLMs in patient-specific temporal contexts by linking each instruction-response pair to specific timestamps, ensuring temporal fidelity throughout the training process. Evaluations show that TIMER-tuned models outperform conventional medical instruction-tuned approaches by 6.6% in completeness on clinician-curated benchmarks, with distribution-matched training demonstrating advantages up to 6.5% in temporal reasoning. Qualitative analyses reveal that using TIMER enhances temporal boundary adherence, trend detection, and chronological precision, necessary for applications such as disease trajectory modeling and treatment response monitoring. Overall, TIMER provides a methodological basis for developing LLMs that can effectively engage with the inherently longitudinal nature of data for patient care. Code is available at TIMER .
View details for DOI 10.1038/s41746-025-01965-9
View details for PubMedID 41006898
View details for PubMedCentralID PMC12475073
-
Verifiable Summarization of Electronic Health Records Using Large Language Models to Support Chart Review.
medRxiv : the preprint server for health sciences
2025
Abstract
Information overload in electronic health records (EHRs) hampers clinicians' ability to efficiently extract and synthesize critical information from a patient's longitudinal health record, leading to increased cognitive burden and delays in care. This study explores the potential of large language models (LLMs) to address this challenge by generating problem-based admission summaries for patients admitted with heart failure, a leading cause of hospitalization worldwide. We developed an extract-then-abstract approach guided by disease-specific "summary bundles" to generate summaries of longitudinal clinical notes that prioritize clinically relevant information. Through a mixed-methods evaluation using real-world clinical notes, we compared physicians' ability to answer patient-specific clinical questions with the LLM-generated summaries versus standard chart review. While summary access did not significantly reduce overall questionnaire completion time, frequent summary use significantly contributed to faster questionnaire completion (p = 0.002). Individual physicians varied in how effectively they leveraged the summaries. Importantly, summary use maintained accuracy in answering clinical questions (88.0% with summaries vs. 86.4% without). All physicians indicated they were "likely" or "very likely" to use the summaries in clinical practice, and 87.5% reported that the summaries would save them time. Preferences for summary format varied, highlighting the need for customizable summaries aligned with individual clinician workflows. This study provides one of the first extrinsic evaluations of LLMs for longitudinal summarization, demonstrating their potential to enhance clinician efficiency, alleviate workload, and support informed decision-making in time-sensitive care environments.
View details for DOI 10.1101/2025.06.02.25328807
View details for PubMedID 40502573
View details for PubMedCentralID PMC12155021
-
Synthetic data distillation enables the extraction of clinical information at scale.
NPJ digital medicine
2025; 8 (1): 267
Abstract
Large-language models (LLMs) show promise for clinical note information extraction, but deployment challenges include high computational costs and privacy concerns. We used synthetic data distillation to fine-tune smaller, open-source LLMs to achieve performance comparable to larger models while enabling local hardware deployment or reduced cloud costs. Using Llama-3.1-70B-Instruct, we generated synthetic question-answer training pairs to fine-tune smaller Llama models. We evaluated performance across three tasks: synthetic clinical trial criteria, the i2b2 2018 Clinical Trial Eligibility Challenge, and apixaban trial criteria questions. The 8B-parameter model achieved high accuracy across all tasks and sometimes outperformed the 70B-Instruct teacher model. Fine-tuning with only the most challenging questions still improved performance, demonstrating the value of targeted training. Results from 3B- and 1B-parameter models showed a clear size-performance tradeoff. This work demonstrates synthetic data distillation's potential for enabling scalable clinical information extraction.
View details for DOI 10.1038/s41746-025-01681-4
View details for PubMedID 40348936
View details for PubMedCentralID PMC12065832
-
APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
edited by Argaw, P., Zhang, H., Jabbour, S., Chandak, P., Ji, J., Mukherjee, S., Salaudeen, O., Chang, T., Healey, E., Groger, F., Adibi, A., Hegselmann, S., Wild, B., Noori, A.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 1-22
View details for Web of Science ID 001863888600001
-
Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study
LANCET DIGITAL HEALTH
2024; 6 (1)
View details for Web of Science ID 001173016200001
https://orcid.org/0000-0002-5370-1746