Bio


Timothy is an MD/PhD student studying cancer biology and biomedical informatics at the Stanford University School of Medicine. He is a joint member of Kara Davis's laboratory in the Department of Pediatrics and Garry Nolan's Laboratory in the Department of Pathology.

As a biomedical data scientist, Timothy's research focuses on the application of machine learning to single-cell data analysis in the context of pediatric leukemia. Through the use of emerging, high-throughout single-cell technologies such as mass cytometry and sequence-based cytometry, Timothy's research is designed to build predictive models of patient outcomes - such as relapse or minimal residual disease (MRD) - at the point of diagnosis. To do so, he uses a variety of computational tools including generalized linear models, clustering, and deep learning. In addition, his work prioritizes constructing easy-to-use, highly-reproducible data analysis pipelines that can be shared as open-source tools for the scientific community.

Outside of science, Timothy has a longstanding interest in human rights and social justice work among members of the lesbian, gay, bisexual, transgender, and queer (LGBTQ+) community. He currently serves as the resident data scientist for the Medical Student Pride Alliance (MSPA), a 501(c)(3) non-profit organization that advocates for diversity, equity, and inclusion for LGBTQ+ medicals students in medical schools across the United States. As a data scientist at MSPA, Timothy analyzes and visualizes data to guide MSPA's strategic decision-making as well as for academic publication. He also advises and mentors other student members of MSPA performing data analysis in Python and R.

In recognition of his accomplishments, Timothy has received several institutional and national award for both research and advocacy. These include a National Research Service Award (NRSA) from the National Cancer Institute, a Junior Leadership Award from the Building the Next Generation of Academic Physicians (BNGAP) LGBT Workforce, Stanford Medicine’s Integrated Strategic Plan Star Award, and a Point Foundation Scholarship.

Honors & Awards


  • Point Foundation Graduate Student Scholarship, Point Foundation (2020)
  • Ruth L. Kirschstein Pre-doctoral National Research Service Award, National Institutes of Health (National Cancer Institute) (2019)
  • Community Impact Award, Stanford University (2019)
  • Integrated Strategic Plan Star Award, Stanford Medicine (2019)
  • Junior Leadership Award, Building the Next Generation of Academic Physicians (BNGAP) LGBT Workforce (2019)
  • Award for Excellence in Promotion of Diversity and Societal Citizenship, Stanford University School of Medicine (2018)

Membership Organizations


  • LGBT-Meds, Financial Officer (former)
  • SUMMA: Stanford University Minority Medical Alliance, Chair (former)
  • Medical Student Pride Alliance (MSPA), Assistant Director for Data Analytics

Education & Certifications


  • Doctor of Philosophy, Stanford University, CANBI-PHD (2024)
  • Bachelor of Arts, Princeton University, Psychology (2014)
  • Master of Science, Stanford University, BMDS-MS (2024)
  • B.A., Princeton University, Psychology and Neuroscience (2014)

All Publications


  • Automating clinical history extraction for flow cytometry panel selection using an EHR-integrated large language model. Cytometry. Part B, Clinical cytometry Rojansky, R., Keyes, T., Oak, J. 2026

    Abstract

    Flow cytometry immunophenotyping is essential for diagnosing hematologic malignancies, but accurate antibody panel selection depends on clinical context that is often fragmented across the electronic health record (EHR). At our institution, clinical laboratory scientists (CLS) use a standardized decision-tree algorithm incorporating specimen characteristics, laboratory values, and prior diagnoses to select panels. This process requires manual review and synthesis of longitudinal clinical documentation and remains time-intensive. We evaluated whether ChatEHR-an institutionally developed, EHR-integrated large language model (LLM) platform-could automate clinical history extraction and support flow cytometry panel selection across 100 cases. Performance was compared with the current CLS-driven workflow using hematopathologist-reviewed panel selections as the reference standard. ChatEHR achieved comparable accuracy to CLS in extracting and categorizing prior hematologic diagnoses (78% vs. 78%). Average processing time was 20.3 s for ChatEHR versus 41 s for manual review, corresponding to an estimated annual reduction of approximately 120 staff hours. Direct panel selection accuracy was lower for ChatEHR (47% vs. 78%), primarily due to errors in deterministic protocol execution and structured data interpretation. Major error categories included incorrect decision-tree logic (39%), diagnostic conflation (22%), failure to retrieve or apply clinical history (22%), laboratory value misinterpretation (9%), and hallucination of nonexistent panels (7%). ChatEHR demonstrated strong performance in extracting and classifying clinical history but was less reliable for deterministic protocol execution and structured data interpretation. Integration of LLM-based clinical history extraction with rules-based laboratory algorithms provides a practical hybrid workflow that preserves diagnostic reliability while reducing manual chart review burden. This approach enables scalable automation while maintaining appropriate hematopathologist oversight in clinical flow cytometry workflows.

    View details for DOI 10.1002/cyto.b.70051

    View details for PubMedID 42427213

  • Challenges in AI Based Tumor Board Case Summarization and Recommendations. Research square Yim, W. W., Damm, H., Pakull, T. M., Preston, S., Keyes, T., Ellis-Caleo, T. J., Sun, Z., Yetisgen, M., Codella, N., Wei, M., Bekheet, F., Neal, J. W., Shah, N., Eryılmaz, B., Nensa, F., Livingstone, E., Friedrich, C. M., Lodde, G. 2026

    Abstract

    Tumor boards, recurring meetings at hospital institutions, assemble multiple cancer specialties (e.g. medical oncology, pathology) to discuss oncological care for patients. In this work, we formally describe the tasks of tumor board case summarization, options generation, and meeting outcomes prediction as AI problems. We study datasets from 4+ medical institutions, providing the performance of available state-of-the-art large language models (LLMs) across DeepSeek, GPT, Gemini, Qwen families, with human-expert evaluations. In total, medical oncologists created reference texts and human judgments across tasks totaling ~10k and ~13k respectively. Results revealed clinicians rated LLM-generated case summaries 3.57-4.59 out of 5.0 across institutions, but struggled with the recommendations generation task, averaging 2.0-3.6. Clinical alignment studies revealed modest correlations of 0.2 for case summarization subtasks, but higher for recommendations at 0.6 for LLM-as-Judge metrics. Our proposed TBFact metrics were shown to be competitive with LLM-as-Judge metrics in longer free text settings, suggesting a promising direction towards explainable metrics. This work, the largest study with expert grading with both granular and direct assessments on multiple axes (e.g. completeness, factual accuracy), revealed surprising nonlinear relationships between granular criterion-level scores and overall direct assessments. Finally, our study revealed experts frequently assigned high ratings to responses that differed substantially from their own reference texts, challenging the core assumptions underlying reference-based automatic evaluation metrics.

    View details for DOI 10.21203/rs.3.rs-9916397/v1

    View details for PubMedID 42370249

    View details for PubMedCentralID PMC13308393

  • Why and How to Monitor Deployed AI Systems in Health Care. NEJM catalyst innovations in care delivery Keyes, T., Callahan, A., Pandya, A. S., Ambers, N., Banda, J. M., Fuentes, M., Lugtu, C., Masariya, P., Nallan, S., O'Brien, C., Wang, T., Alsentzer, E., Chen, J. H., Dash, D., Eisenberg, M. A., Garcia, P., Kotecha, N., Revri, A., Pfeffer, M. A., Shah, N. H., Jain, S. S. 2026; 7 (6): CAT250372

    Abstract

    Postdeployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit - and to support governance decisions about which systems to update, modify, or decommission. Motivated by these needs, the authors developed a framework for monitoring deployed AI systems organized around three complementary principles: system integrity, performance, and impact. System integrity monitoring focuses on maximizing system uptime, detecting runtime errors, and identifying when changes to the surrounding information technology ecosystem have unintended effects. Performance monitoring focuses on maintaining accurate and equitable system behavior in the face of changing health care practices (and thus input data) over time. Impact monitoring assesses whether a deployed system continues to have value in the form of benefit to clinicians, staff, and patients. Drawing on examples of deployed AI systems at their academic medical center, the authors provide practical guidance for creating monitoring plans based on these principles that specify which metrics to measure and at what cadence, who is responsible for acting when metrics change, and what concrete follow-up actions should be taken - for both traditional and generative AI. They also discuss challenges in implementing this framework, including the effort of monitoring for health systems with limited resources, and the difficulty of incorporating data-driven monitoring practices into complex organizations where conflicting priorities and definitions of success often coexist. This framework offers a starting point for health systems seeking to ensure that AI deployments remain safe and effective over time.

    View details for DOI 10.1056/CAT.25.0372

    View details for PubMedID 42418612

  • Why and How to Monitor Deployed AI Systems in Health Care NEJM CATALYST INNOVATIONS IN CARE DELIVERY Keyes, T., Callahan, A., Pandya, A. S., Ambers, N., Banda, J. M., Fuentes, M., Lugtu, C., Masariya, P., Nallan, S., O'Brien, C., Wang, T., Alsentzer, E., Chen, J. H., Dash, D., Eisenberg, M. A., Garcia, P., Kotecha, N., Revri, A., Pfeffer, M. A., Shah, N. H., Jain, S. S. 2026; 7 (6)
  • Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries. JAMA network open Grolleau, F., Liang, A. S., Keyes, T., Ma, S. P., Lew, T., Huynh, T. R., Steele, N., Chung, P., Qin, P., Chandra, G., Wang, S. F., Mullen, E., Carpenter, L., Hoppenfeld, M., Morrin, M., Kyerematen, B. A., Ambers, N., Kotecha, N., Alsentzer, E., Hom, J., Shah, N. H., Schulman, K., Chen, J. H. 2026; 9 (5): e2616556

    Abstract

    High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest that large language models (LLMs) can generate clinical summaries of comparable quality to those by physicians, prospective data on their safety, utility, and association with clinician well-being in clinical environments are lacking.To evaluate the safety, use, and association with clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries, during prospective clinical deployment.This single-arm prospective pilot quality improvement study encompassed hospital discharges at 1 academic inpatient medicine unit from August 1 to October 11, 2025, with baseline comparisons drawn from April 9 to July 31, 2025.A custom agentic LLM workflow using Gemini 2.5 Pro generated draft hospital course summaries nightly using patient history and physical and daily progress notes. Drafts were securely emailed to physicians daily for review and optional use.The primary outcome was physician-reported potential for and severity of harm from unedited summaries (Agency for Healthcare Research and Quality Common Format Harm Scale). Secondary outcomes included use rate, error types (omissions, inaccuracies, and hallucinations), time spent in discharge summaries (electronic health record logs), and changes in cognitive burden (NASA Task Load Index; score range, 0-100, with higher scores indicating greater cognitive burden) and burnout (Stanford Professional Fulfillment Index Work Exhaustion Scale; score range, 0-4, with higher scores indicating greater burnout).Among 384 hospital discharges, the system generated 1274 summaries. Physicians used artificial intelligence (AI) content in 219 cases (57.0%). Feedback on 100 summaries (88 of 219 used summaries [40.2%] and 12 of 165 unused summaries [7.3%]) noted omissions (25 summaries [25.0%]) and inaccuracies (20 summaries [20.0%]) but rare hallucinations (2 summaries [2.0%]). Physicians rated 88 unedited summaries (88.0%) as having no harm potential and 1 (1.0%) as likely to cause moderate harm; no severe harm was reported. Mean physician burnout scores decreased significantly from before to after the intervention (1.75; 95% CI, 1.16-2.34 vs 1.20; 95% CI, 0.71-1.69; P = .03). Time savings were heterogeneous, with 5 of 7 physicians with matched baseline data (71.4%) seeing reductions in median documentation time; changes from baseline to pilot were up to 2.9 minutes, which was a nonsignificant difference (10.7 minutes; 95% CI, 7.4-13.3 minutes vs 7.8 minutes; 95% CI, 5.1-11.7 minutes; P = .13).In this study, an LLM-based agentic workflow produced hospital course summaries that were frequently used with minimal risk of harm identified. The intervention was associated with a reduction in physician burnout, supporting the viability of AI summarization to mitigate documentation burden.

    View details for DOI 10.1001/jamanetworkopen.2026.16556

    View details for PubMedID 42101844

  • Use of a large language model integrated within the electronic medical record for the evaluation of surgical site infections - Northern California, 2025. Infection control and hospital epidemiology Miranti, E., Keyes, T., Ayala, A., Ambers, N., Newman, G., de Leon, E., Viana-Cardenas, E. P., Tariq, W., Sampson, M., Salinas, J. L. 2026: 1-3

    Abstract

    Our study evaluated a large language model (gpt-4o-mini) for surgical site infection (SSI) adjudication, achieving 100% sensitivity but 69.4% specificity. While reducing the manual screening workload by 66%, the agent generated many false positives, underscoring the need for refined models to improve specificity without compromising accuracy.

    View details for DOI 10.1017/ice.2026.10432

    View details for PubMedID 41972262

  • MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing Grolleau, F., Alsentzer, E., Keyes, T., Chung, P., Swaminathan, A., Aali, A., Hom, J., Huynh, T., Lew, T., Liang, A., Chu, W., Steele, N., Lin, C., Yang, J., Black, K., Ma, S., Haredasht, F. N., Shah, N. H., Schulman, K., Chen, J. H. 2026; 31: 388-399

    Abstract

    Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assurance these systems require. We address this challenge with two complementary contributions. First, we introduce MedFactEval, a framework for scalable, fact-grounded evaluation where clinicians define high-salience key facts and an "LLM Jury"-a multi-LLM majority vote-assesses their inclusion in generated summaries. Second, we present MedAgentBrief, a model-agnostic, multi-step workflow designed to generate high-quality, factual discharge summaries. To validate our evaluation framework, we established a gold-standard reference using a seven-physician majority vote on clinician-defined key facts from inpatient cases. The MedFactEval LLM Jury achieved almost perfect agreement with this panel (Cohen's κ = 81%), a performance statistically non-inferior to that of a single human expert (κ = 67%, P < 0.001). Our work provides both a robust evaluation framework (MedFactEval) and a high-performing generation workflow (MedAgentBrief), offering a comprehensive approach to advance the responsible deployment of generative AI in clinical workflows.

    View details for DOI 10.1142/9789819824755_0027

    View details for PubMedID 41758155

  • DHODH as a Targetable Metabolic Achilles' Heel for chemo-resistant B-ALL. Blood Liu, Y., Jiang, H., Liu, J., Stuani, L., Merchant, M. J., Jager, A., Koladiya, A., Chang, T. C., Domizi, P., Sarno, J., Wang, A., Keyes, T., Jedoui, D., Meng, J., Hartmann, F., Hou, R., Fries, C., Pirillo, C., Gao, Q., Iacobucci, I., Bendall, S. C., Huang, M., Lacayo, N. J., Sakamoto, K. M., Mullighan, C. G., Loh, M. L., Yu, J., Yang, J. J., Ye, J., Davis, K. L. 2026

    Abstract

    Relapse remains a major barrier to survival in B-cell acute lymphoblastic leukemia (B-ALL). Both activation of B-cell signaling pathways and increased glucose consumption have been linked to chemo-resistance and relapse risk. Here, we connect these observations, showing that B-ALL cells with active signaling, marked by high phosphorylated ribosomal protein S6 (pS6+), are glucose dependent. Isotope tracing confirms that pS6+ cells are highly glycolytic and rely on glucose for de novo nucleotide synthesis. Uridine, but not other purines or pyrimidines, rescues pS6+ cells from glucose deprivation, highlighting uridine as essential for survival. Active mTOR signaling in pS6+ cells drives de novo pyrimidine synthesis by activating CAD (Carbamoyl phosphate synthetase 2, Aspartate transcarbamylase, and Dihydroorotase), which catalyzes the first steps of de novo pyrimidine synthesis. Inhibiting signaling abolishes glucose dependency and CAD phosphorylation. Primary pS6+ cells express high levels of pyrimidine synthesis proteins, including dihydroorotate dehydrogenase (DHODH), the rate-limiting enzyme in pyrimidine synthesis. Increased DHODH expression correlates with relapse and poor event-free survival. Most B-ALL molecular subtypes exhibit DHODH activity. BAY-2402234, a DHODH inhibitor, effectively kills pS6+ cells in vitro, with IC50 values correlating with pS6 signaling strength across 14 B-ALL patient-derived xenografts (PDX). In vivo, DHODH inhibition prolongs survival and reduces leukemia burden in pS6+ B-ALL models. These findings link active signaling to pyrimidine dependency and relapse risk, highlighting DHODH inhibition as a promising therapeutic strategy for chemo-resistant B-ALL.

    View details for DOI 10.1182/blood.2025029264

    View details for PubMedID 41576347

  • Holistic evaluation of large language models for medical tasks with MedHELM. Nature medicine Bedi, S., Cui, H., Fuentes, M., Unell, A., Wornow, M., Banda, J. M., Kotecha, N., Keyes, T., Mai, Y., Oez, M., Qiu, H., Jain, S., Schettini, L., Kashyap, M., Fries, J. A., Swaminathan, A., Chung, P., Haredasht, F. N., Lopez, I., Aali, A., Tse, G., Nayak, A., Vedak, S., Jain, S. S., Patel, B., Fayanju, O., Shah, S., Goh, E., Yao, D. H., Soetikno, B., Reis, E., Gatidis, S., Divi, V., Capasso, R., Saralkar, R., Chiang, C. C., Jindal, J., Pham, T., Ghoddusi, F., Lin, S., Chiou, A. S., Hong, H. J., Roy, M., Gensheimer, M. F., Patel, H., Schulman, K., Dash, D., Char, D., Downing, L., Grolleau, F., Black, K., Mieso, B., Zahedivash, A., Yim, W. W., Sharma, H., Lee, T., Kirsch, H., Lee, J., Ambers, N., Lugtu, C., Sharma, A., Mawji, B., Alekseyev, A., Zhou, V., Kakkar, V., Helzer, J., Revri, A., Bannett, Y., Daneshjou, R., Chen, J., Alsentzer, E., Morse, K., Ravi, N., Aghaeepour, N., Kennedy, V., Chaudhari, A., Wang, T., Koyejo, S., Lungren, M. P., Horvitz, E., Liang, P., Pfeffer, M. A., Shah, N. H. 2026

    Abstract

    While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice. Here we introduce MedHELM, an extensible evaluation framework with three contributions. First, a clinician-validated taxonomy organizing medical AI applications into five categories that mirror real clinical tasks-clinical decision support (diagnostic decisions, treatment planning), clinical note generation (visit documentation, procedure reports), patient communication (education materials, care instructions), medical research (literature analysis, clinical data analysis) and administration (scheduling, workflow coordination). These encompass 22 subcategories and 121 specific tasks reflecting daily medical practice. Second, a comprehensive benchmark suite of 37 evaluations covering all subcategories. Third, systematic comparison of nine frontier LLMs-Claude 3.5 Sonnet, Claude 3.7 Sonnet, DeepSeek R1, Gemini 1.5 Pro, Gemini 2.0 Flash, GPT-4o, GPT-4o mini, Llama 3.3 and o3-mini-using an automated LLM-jury evaluation method. Our LLM-jury uses multiple AI evaluators to assess model outputs against expert-defined criteria. Advanced reasoning models (DeepSeek R1, o3-mini) demonstrated superior performance with win rates of 66%, although Claude 3.5 Sonnet achieved comparable results at 15% lower computational cost. These results not only highlight current model capabilities but also demonstrate how MedHELM could enable evidence-based selection of medical AI systems for healthcare applications.

    View details for DOI 10.1038/s41591-025-04151-2

    View details for PubMedID 41559415

    View details for PubMedCentralID 10916499

  • MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries Grolleau, F., Alsentzer, E., Keyes, T., Chung, P., Swaminathan, A., Aali, A., Hom, J., Tridu Huynh, Lew, T., Liang, A., Chu, W., Steele, N., Lin, C., Yang, J., Black, K., Ma, S., Haredasht, F. N., Shah, N. H., Schulman, K., Chen, J. H. edited by Altman, R. B., Hunter, L., Ritchie, M. D., Murray, T., Klein, T. E. WORLD SCIENTIFIC PUBL CO PTE LTD. 2026: 388-399

    Abstract

    Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assurance these systems require. We address this challenge with two complementary contributions. First, we introduce MedFactEval, a framework for scalable, fact-grounded evaluation where clinicians define high-salience key facts and an "LLM Jury"-a multi-LLM majority vote-assesses their inclusion in generated summaries. Second, we present MedAgentBrief, a model-agnostic, multi-step workflow designed to generate high-quality, factual discharge summaries. To validate our evaluation framework, we established a gold-standard reference using a seven-physician majority vote on clinician-defined key facts from inpatient cases. The MedFactEval LLM Jury achieved almost perfect agreement with this panel (Cohen's κ = 81%), a performance statistically non-inferior to that of a single human expert (κ = 67%, P < 0.001). Our work provides both a robust evaluation framework (MedFactEval) and a high-performing generation workflow (MedAgentBrief), offering a comprehensive approach to advance the responsible deployment of generative AI in clinical workflows.

    View details for Web of Science ID 001796654100027

    View details for PubMedID 41758155

  • Using secure artificial intelligence agents integrated within the electronic medical record for the evaluation of blood culture appropriateness-Northern California, 2025. Infection control and hospital epidemiology Rodriguez-Nava, G., Keyes, T., Ambers, N., Miranti, E., Viana-Cardenas, E. P., Tariq, W., Sampson, M. M., Salinas, J. L. 2025: 1-4

    Abstract

    We evaluated large language model (LLM)-based agents integrated with the electronic medical record to assess blood culture appropriateness. While sensitivity was high, specificity remained low. Performance was shaped by prompt phrasing, sycophantic behavior, and semantic triggers, reflecting both the potential and limitations of LLMs in real-world clinical decision support.

    View details for DOI 10.1017/ice.2025.10349

    View details for PubMedID 41216923

  • Target Product Profile to Evaluate the Clinical Utility, Financial Impact, and Ethical Implications of an AI-Based HCM Detection Model Parsa, S., Keyes, T., Dash, D., Mello, M., Salisbury, H., Callahan, A., Goto, S., Salerno, M., Parikh, V., Mahaffey, K., Ashley, E., Shah, N., Jain, S. LIPPINCOTT WILLIAMS & WILKINS. 2025
  • Multi-omic analysis of early minimal residual disease identifies functional and transcriptomic signatures of treatment resistance in pediatric B-cell acute lymphoblastic leukemia Scribano, R., Delgado, A., Domizi, P., Buracchi, C., Jager, A., Keyes, T., Bugarin, C., Gaipa, G., Davis, K., Sarno, J. ELSEVIER. 2025: 334-335
  • Annotation-free discovery of disease-relevant cells in single-cell datasets. Science advances Craig, E., Keyes, T. J., Sarno, J., D'Silva, J. P., Domizi, P., Zaslavsky, M., Tsai, A., Glass, D., Nolan, G. P., Hastie, T., Tibshirani, R., Davis, K. L. 2025; 11 (35): eadv5019

    Abstract

    In single-cell datasets, patient labels indicating disease status (e.g., "sick" or "not sick") are typically available, but individual cell labels indicating which of a patient's cells are associated with their disease state are generally unknown. To address this, we introduce mixture modeling for multiple-instance learning (MMIL), an expectation-maximization approach that trains cell-level binary classifiers using only patient-level labels. Applied to primary samples from patients with acute leukemia, MMIL accurately separates leukemia from nonleukemia baseline cells, including rare minimal residual disease (MRD) cells; generalizes across tissues and treatment time points; and identifies biologically relevant features with accuracy approaching that of a hematopathologist. MMIL can also incorporate cell labels when they are available, creating a robust framework for leveraging both labeled and unlabeled cells. MMIL provides a flexible modeling framework for cell classification, especially in scenarios with unknown gold-standard cell labels.

    View details for DOI 10.1126/sciadv.adv5019

    View details for PubMedID 40864714

  • Uridine Metabolism as a Targetable Metabolic Achilles' Heel for chemo-resistant B-ALL. bioRxiv : the preprint server for biology Liu, Y., Jiang, H., Liu, J., Stuani, L., Merchant, M., Jager, A., Koladiya, A., Chang, T. C., Domizi, P., Sarno, J., Keyes, T., Jedoui, D., Wang, A., Meng, J., Hartmann, F., Bendall, S. C., Huang, M., Lacayo, N. J., Sakamoto, K. M., Mullighan, C. G., Loh, M., Yu, J., Yang, J., Ye, J., Davis, K. L. 2025

    Abstract

    Relapse continues to limit survival for patients with B-cell acute lymphoblastic leukemia (B-ALL). Previous studies have independently implicated activation of B-cell developmental signaling pathways and increased glucose consumption with chemo-resistance and relapse risk. Here, we connect these observations, demonstrating that B-ALL cells with active signaling, defined by high expression of phosphorylated ribosomal protein S6 ("pS6+ cells"), are metabolically unique and glucose dependent. Isotope tracing and metabolic flux analysis confirm that pS6+ cells are highly glycolytic and notably sensitive to glucose deprivation, relying on glucose for de novo nucleotide synthesis. Uridine, but not purine or pyrimidine, rescues pS6+ cells from glucose deprivation, highlighting uridine is essential for their survival. Active signaling in pS6+ cells drives uridine production through activating phosphorylation of carbamoyl phosphate synthetase (CAD), the enzyme catalyzing the initial steps of uridine synthesis. Inhibition of signaling abolishes glucose dependency and CAD phosphorylation in pS6+ cells. Primary pS6+ cells demonstrate high expression of uridine synthesis proteins, including dihydroorotate dehydrogenase (DHODH), the rate-limiting catalyst of de novo uridine synthesis. Gene expression demonstrates that increased expression of DHODH is associated with relapse and inferior event-free survival after chemotherapy. Further, the majority of B-ALL genomic subtypes demonstrate activity of DHODH. Inhibiting DHODH using BAY2402232 effectively kills pS6+ cells in vitro, with its IC50 correlated with the strength of pS6 signaling across 14 B-ALL cell lines and patient-derived xenografts (PDX). In vivo DHODH inhibition prolongs survival and decreases leukemia burden in pS6+ B-ALL cell line and PDX models. These findings link active signaling to uridine dependency in B-ALL cells and an associated risk of relapse. Targeting uridine synthesis through DHODH inhibition offers a promising therapeutic strategy for chemo-resistant B-ALL as a novel therapeutic approach for resistant disease.

    View details for DOI 10.1101/2025.01.27.635108

    View details for PubMedID 39975156

    View details for PubMedCentralID PMC11838334

  • The tidyomics ecosystem: enhancing omic data analyses. Nature methods Hutchison, W. J., Keyes, T. J., Crowell, H. L., Serizay, J., Soneson, C., Davis, E. S., Sato, N., Moses, L., Tarlinton, B., Nahid, A. A., Kosmac, M., Clayssen, Q., Yuan, V., Mu, W., Park, J. E., Mamede, I., Ryu, M. H., Axisa, P. P., Paiz, P., Poon, C. L., Tang, M., Gottardo, R., Morgan, M., Lee, S., Lawrence, M., Hicks, S. C., Nolan, G. P., Davis, K. L., Papenfuss, A. T., Love, M. I., Mangiola, S. 2024

    Abstract

    The growth of omic data presents evolving challenges in data manipulation, analysis and integration. Addressing these challenges, Bioconductor provides an extensive community-driven biological data analysis platform. Meanwhile, tidy R programming offers a revolutionary data organization and manipulation standard. Here we present the tidyomics software ecosystem, bridging Bioconductor to the tidy R paradigm. This ecosystem aims to streamline omic analysis, ease learning and encourage cross-disciplinary collaborations. We demonstrate the effectiveness of tidyomics by analyzing 7.5 million peripheral blood mononuclear cells from the Human Cell Atlas, spanning six data frameworks and ten analysis tools.

    View details for DOI 10.1038/s41592-024-02299-2

    View details for PubMedID 38877315

    View details for PubMedCentralID 545600

  • Thetidyomicsecosystem: Enhancing omic data analyses. bioRxiv : the preprint server for biology Hutchison, W. J., Keyes, T. J., tidyomics Consortium, Crowell, H. L., Serizay, J., Soneson, C., Davis, E. S., Sato, N., Moses, L., Tarlinton, B., Nahid, A. A., Kosmac, M., Clayssen, Q., Yuan, V., Mu, W., Park, J., Mamede, I., Ryu, M. H., Axisa, P., Paiz, P., Poon, C., Tang, M., Gottardo, R., Morgan, M., Lee, S., Lawrence, M., Hicks, S. C., Nolan, G. P., Davis, K. L., Papenfuss, A. T., Love, M. I., Mangiola, S. 2024

    Abstract

    The growth of omic data presents evolving challenges in data manipulation, analysis, and integration. Addressing these challenges, Bioconductor1 provides an extensive community-driven biological data analysis platform. Meanwhile, tidy R programming2 offers a revolutionary standard for data organisation and manipulation. Here, we present the tidyomics software ecosystem, bridging Bioconductor to the tidy R paradigm. This ecosystem aims to streamline omic analysis, ease learning, and encourage cross-disciplinary collaborations. We demonstrate the effectiveness of tidyomics by analysing 7.5 million peripheral blood mononuclear cells from the Human Cell Atlas3, spanning six data frameworks and ten analysis tools.

    View details for DOI 10.1101/2023.09.10.557072

    View details for PubMedID 38826347

  • Sociodemographic factors and research experience impact MD-PhD program acceptance JCI INSIGHT Williams, D., Christophers, B., Keyes, T., Kumar, R., Granovetter, M. C., Adigun, A., Olivera, J., Pura-Bryant, J., Smith, C., Okafor, C., Shibre, M., Daye, D., Akabas, M. H. 2024; 9 (3)

    Abstract

    The 2014 NIH Physician-Scientist Workforce Working Group predicted a future shortage of physician-scientists. Subsequent studies have highlighted disparities in MD-PhD admissions based on race, income, and education. Our analysis of data from the Association of American Medical Colleges covering 2014-2021 (15,156 applicants and 6,840 acceptees) revealed that acceptance into US MD-PhD programs correlates with research experience, family income, and research publications. The number of research experiences associated with parental education and family income. Applicants were more likely to be accepted with a family income greater than $50,000 or with one or more publications or presentations. Applicants were less likely to be accepted if they had parents without a graduate degree, were Black/African American, were first-generation college students, or were reapplicants, irrespective of the number of research experiences, publications, or presentations. These findings underscore an admissions bias that favors candidates from affluent and highly educated families, while disadvantaging underrepresented minorities.

    View details for DOI 10.1172/jci.insight.176146

    View details for Web of Science ID 001161909900001

    View details for PubMedID 38329127

    View details for PubMedCentralID PMC10967469

  • Teaching LGBTQ+ Health, a Web-Based Faculty Development Course: Program Evaluation Study Using the RE-AIM Framework. JMIR medical education Gisondi, M. A., Keyes, T., Zucker, S., Bumgardner, D. 2023; 9: e47777

    Abstract

    Many health professions faculty members lack training on fundamental lesbian, gay, bisexual, transgender, and queer (LGBTQ+) health topics. Faculty development is needed to address knowledge gaps, improve teaching, and prepare students to competently care for the growing LGBTQ+ population.We conducted a program evaluation of the massive open online course Teaching LGBTQ+ Health: A Faculty Development Course for Health Professions Educators from the Stanford School of Medicine. Our goal was to understand participant demographics, impact, and ongoing maintenance needs to inform decisions about updating the course.We evaluated the course for the period from March 27, 2021, to February 24, 2023, guided by the RE-AIM (Reach, Effectiveness, Adoption, Implementation, and Maintenance) framework. We assessed impact using participation numbers, evidence of learning, and likelihood of practice change. Data included participant demographics, performance on a pre- and postcourse quiz, open-text entries throughout the course, continuing medical education (CME) credits awarded, and CME course evaluations. We analyzed demographics using descriptive statistics and pre- and postcourse quiz scores using a paired 2-tailed t test. We conducted a qualitative thematic analysis of open-text responses to prompts within the course and CME evaluation questions.Results were reported using the 5 framework domains. Regarding Reach, 1782 learners participated in the course, and 1516 (85.07%) accessed it through a main course website. Of the different types of participants, most were physicians (423/1516, 27.9%) and from outside the sponsoring institution and target audience (1452/1516, 95.78%). Regarding Effectiveness, the median change in test scores for the 38.1% (679/1782) of participants who completed both the pre- and postcourse tests was 3 out of 10 points, or a 30% improvement (P<.001). Themes identified from CME evaluations included LGBTQ+ health as a distinct domain, inclusivity in practices, and teaching LGBTQ+ health strategies. A minority of participants (237/1782, 13.3%) earned CME credits. Regarding Adoption, themes identified among responses to prompts in the course included LGBTQ+ health concepts and instructional strategies. Most participants strongly agreed with numerous positive statements about the course content, presentation, and likelihood of practice change. Regarding Implementation, the course cost US $57,000 to build and was intramurally funded through grants and subsidies. The course faculty spent an estimated 600 hours on the project, and educational technologists spent another 712 hours. Regarding Maintenance, much of the course is evergreen, and ongoing oversight and quality assurance require minimal faculty time. New content will likely include modules on transgender health and gender-affirming care.Teaching LGBTQ+ Health improved participants' knowledge of fundamental queer health topics. Overall participation has been modest to date. Most participants indicated an intention to change clinical or teaching practices. Maintenance costs are minimal. The web-based course will continue to be offered, and new content will likely be added.

    View details for DOI 10.2196/47777

    View details for PubMedID 37477962

  • Single-cell technologies uncover intra-tumor heterogeneity in childhood cancers. Seminars in immunopathology Lo, Y., Liu, Y., Kammersgaard, M., Koladiya, A., Keyes, T. J., Davis, K. L. 2023

    Abstract

    Childhood cancer is the second leading cause of death in children aged 1 to 14. Although survival rates have vastly improved over the past 40years, cancer resistance and relapse remain a significant challenge. Advances in single-cell technologies enable dissection of tumors to unprecedented resolution. This facilitates unraveling the heterogeneity of childhood cancers to identify cell subtypes that are prone to treatment resistance. The rapid accumulation of single-cell data from different modalities necessitates the development of novel computational approaches for processing, visualizing, and analyzing single-cell data. Here, we review single-cell approaches utilized or under development in the context of childhood cancers. We review computational methods for analyzing single-cell data and discuss best practices for their application. Finally, we review the impact of several studies of childhood tumors analyzed with these approaches and future directions to implement single-cell studies into translational cancer research in pediatric oncology.

    View details for DOI 10.1007/s00281-022-00981-1

    View details for PubMedID 36625902

  • tidytof: a user-friendly framework for scalable and reproducible high-dimensional cytometry data analysis. Bioinformatics advances Keyes, T. J., Koladiya, A., Lo, Y., Nolan, G. P., Davis, K. L. 2023; 3 (1): vbad071

    Abstract

    Summary: While many algorithms for analyzing high-dimensional cytometry data have now been developed, the software implementations of these algorithms remain highly customized-this means that exploring a dataset requires users to learn unique, often poorly interoperable package syntaxes for each step of data processing. To solve this problem, we developed {tidytof}, an open-source R package for analyzing high-dimensional cytometry data using the increasingly popular 'tidy data' interface.Availability and implementation: {tidytof} is available at https://github.com/keyes-timothy/tidytof and is released under the MIT license. It is supported on Linux, MS Windows and MacOS. Additional documentation is available at the package website (https://keyes-timothy.github.io/tidytof/).Supplementary information: Supplementary data are available at Bioinformatics Advances online.

    View details for DOI 10.1093/bioadv/vbad071

    View details for PubMedID 37351311

  • Improved Relapse Prediction in Pediatric Acute Myeloid Leukemia By Deconvolving Lineage-Specific and CancerSpecific Features in Single-Cell Data Keyes, T., Jager, A., Krueger, M., Plevritis, S., Tibshirani, R., Aplenc, R., Nolan, G. P., Redell, M. S., Davis, K. L. AMER SOC HEMATOLOGY. 2022: 6288-6289
  • CytofIn enables integrated analysis of public mass cytometry datasets using generalized anchors. Nature communications Lo, Y., Keyes, T. J., Jager, A., Sarno, J., Domizi, P., Majeti, R., Sakamoto, K. M., Lacayo, N., Mullighan, C. G., Waters, J., Sahaf, B., Bendall, S. C., Davis, K. L. 2022; 13 (1): 934

    Abstract

    The increasing use of mass cytometry for analyzing clinical samples offers the possibility to perform comparative analyses across public datasets. However, challenges in batch normalization and data integration limit the comparison of datasets not intended to be analyzed together. Here, we present a data integration strategy, CytofIn, using generalized anchors to integrate mass cytometry datasets from the public domain. We show that low-variance controls, such as healthy samples and stable channels, are inherently homogeneous, robust against stimulation, and can serve as generalized anchors for batch correction. Single-cell quantification comparing mass cytometry data from 989 leukemia files pre- and post normalization with CytofIn demonstrates effective batch correction while recapitulating the gold-standard bead normalization. CytofIn integration of public cancer datasets enabled the comparison of immune features across histologies and treatments. We demonstrate the ability to integrate public datasets without necessitating identical control samples or bead standards for fast and robust analysis using CytofIn.

    View details for DOI 10.1038/s41467-022-28484-5

    View details for PubMedID 35177627

  • Documenting Social Media Engagement as Scholarship: A New Model for Assessing Academic Accomplishment for the Health Professions. Journal of medical Internet research Acquaviva, K. D., Mugele, J., Abadilla, N., Adamson, T., Bernstein, S. L., Bhayani, R. K., Büchi, A. E., Burbage, D., Carroll, C. L., Davis, S. P., Dhawan, N., English, K., Grier, J. T., Gurney, M. K., Hahn, E. S., Haq, H., Huang, B., Jain, S., Jun, J., Kerr, W. T., Keyes, T., Kirby, A. R., Leary, M., Marr, M., Major, A., Meisel, J. V., Petersen, E. A., Raguan, B., Rhodes, A., Rupert, D. D., Sam-Agudu, N. A., Saul, N., Shah, J. R., Sheldon, L. K., Sinclair, C. T., Spencer, K., Strand, N. H., Streed, C. G., Trudell, A. M. 2020; 22 (12): e25070

    Abstract

    The traditional model of promotion and tenure in the health professions relies heavily on formal scholarship through teaching, research, and service. Institutions consider how much weight to give activities in each of these areas and determine a threshold for advancement. With the emergence of social media, scholars can engage wider audiences in creative ways and have a broader impact. Conventional metrics like the h-index do not account for social media impact. Social media engagement is poorly represented in most curricula vitae (CV) and therefore is undervalued in promotion and tenure reviews.The objective was to develop crowdsourced guidelines for documenting social media scholarship. These guidelines aimed to provide a structure for documenting a scholar's general impact on social media, as well as methods of documenting individual social media contributions exemplifying innovation, education, mentorship, advocacy, and dissemination.To create unifying guidelines, we created a crowdsourced process that capitalized on the strengths of social media and generated a case example of successful use of the medium for academic collaboration. The primary author created a draft of the guidelines and then sought input from users on Twitter via a publicly accessible Google Document. There was no limitation on who could provide input and the work was done in a democratic, collaborative fashion. Contributors edited the draft over a period of 1 week (September 12-18, 2020). The primary and secondary authors then revised the draft to make it more concise. The guidelines and manuscript were then distributed to the contributors for edits and adopted by the group. All contributors were given the opportunity to serve as coauthors on the publication and were told upfront that authorship would depend on whether they were able to document the ways in which they met the 4 International Committee of Medical Journal Editors authorship criteria.We developed 2 sets of guidelines: Guidelines for Listing All Social Media Scholarship Under Public Scholarship (in Research/Scholarship Section of CV) and Guidelines for Listing Social Media Scholarship Under Research, Teaching, and Service Sections of CV. Institutions can choose which set fits their existing CV format.With more uniformity, scholars can better represent the full scope and impact of their work. These guidelines are not intended to dictate how individual institutions should weigh social media contributions within promotion and tenure cases. Instead, by providing an initial set of guidelines, we hope to provide scholars and their institutions with a common format and language to document social media scholarship.

    View details for DOI 10.2196/25070

    View details for PubMedID 33263554

  • Progressive B Cell Loss in Revertant X-SCID. Journal of clinical immunology Lin, C. H., Kuehn, H. S., Thauland, T. J., Lee, C. M., De Ravin, S. S., Malech, H. L., Keyes, T. J., Jager, A., Davis, K. L., Garcia-Lloret, M. I., Rosenzweig, S. D., Butte, M. J. 2020

    Abstract

    We report the case of a patient with X-linked severe combined immunodeficiency (X-SCID) who survived for over 20years without hematopoietic stem cell transplantation (HSCT) because of a somatic reversionmutation. An important feature of this rare case included the strategy to validate the pathogenicity of a variant of the IL2RG gene when the T and B cell lineages comprised only revertant cells. We studied the X-inactivation of sorted T cells from the mother to show that the pathogenic variant was indeed the cause of his SCID. One interesting feature was a progressive loss of B cells over 20years. CyTOF (cytometry time of flight) analysis of bone marrow offered a potential explanation of the B cell failure, with expansions of progenitor populations that suggest a developmental block. Another interesting feature was that the patient bore extensive granulomatous disease and skin cancers that contained T cells, despite severe T cell lymphopenia in the blood. Finally, the patient had a few hundred T cells on presentation but his TCRs comprised a very limited repertoire, supporting the important conclusion that repertoire size trumps numbers of T cells.

    View details for DOI 10.1007/s10875-020-00825-3

    View details for PubMedID 32681206

  • A Cancer Biologist's Primer on Machine Learning Applications in High-Dimensional Cytometry. Cytometry. Part A : the journal of the International Society for Analytical Cytology Keyes, T. J., Domizi, P., Lo, Y., Nolan, G. P., Davis, K. L. 2020

    Abstract

    The application of machine learning and artificial intelligence to high-dimensional cytometry data sets has increasingly become a staple of bioinformatic data analysis over the past decade. This is especially true in the field of cancer biology, where protocols for collecting multiparameter single-cell data in a high-throughput fashion are rapidly developed. As the use of machine learning methodology in cytometry becomes increasingly common, there is a need for cancer biologists to understand the basic theory and applications of a variety of algorithmic tools for analyzing and interpreting cytometry data. We introduce the reader to several keystone machine learning-based analytic approaches with an emphasis on defining key terms and introducing a conceptual framework for making translational or clinically relevant discoveries. The target audience consists of cancer cell biologists and physician-scientists interested in applying these tools to their own data, but who may have limited training in bioinformatics. © 2020 International Society for Advancement of Cytometry.

    View details for DOI 10.1002/cyto.a.24158

    View details for PubMedID 32602650

  • Medical Student Pride Alliance: The first national LGBTQ+ medical student affinity organisation. Medical education Goetz, T. G., Zucker, S., Keyes, T., Gisondi, M. 2020; 54 (5): 471-472

    View details for DOI 10.1111/medu.14112

    View details for PubMedID 32242963

  • Student Education About Pre-exposure Prophylaxis (PrEP) Varies Between Regions of the United States. Journal of general internal medicine Bunting, S. R., Garber, S. S., Goldstein, R. H., Ritchie, T. D., Batteson, T. J., Keyes, T. J. 2020

    Abstract

    Daily, oral pre-exposure prophylaxis (PrEP) is an effective and safe prevention strategy for people at risk for HIV. However, prescription of PrEP has been limited for patients at the highest risk. Disparities in PrEP prescription are pronounced among racial and gender minority patients. A significant body of literature indicates that practicing healthcare providers have little awareness and knowledge of PrEP. Very little work has investigated the education about PrEP among health professionals in training.The objective of this study was to compare health professions students' awareness of PrEP and education about PrEP between regions of the US, and to determine if correlations between regional HIV incidence and PrEP use were present.Survey study.A cross-sectional sample of health professions students (N = 1859) representing future prescribers (MD, DO, PA), pharmacists, and nurses in the US.Overall, 83.4% of students were aware of PrEP, but only 62.2% of fourth-year students indicated they had been taught about PrEP at any time during their training. Education about PrEP was most comprehensive in the Northeastern US, the area with the highest PrEP to need ratio (4.7). In all regions, transgender patients and heterosexual men and women were least likely to be presented in education as PrEP candidates, and men who have sex with men were the most frequently presented.There are marked differences in education regarding PrEP both between academic programs and regions of the USA.

    View details for DOI 10.1007/s11606-020-05736-y

    View details for PubMedID 32080792

  • Navigating Controversy: A Critical Element of Medical Education ACADEMIC MEDICINE Jia, J. L., Kamceva, M., Keyes, T. J. 2018; 93 (12): 1750
  • Navigating Controversy: A Critical Element of Medical Education. Academic medicine : journal of the Association of American Medical Colleges Jia, J. L., Kamceva, M. n., Keyes, T. J. 2018; 93 (12): 1750

    View details for PubMedID 30489297

  • Structural and functional features of central nervous system lymphatic vessels NATURE Louveau, A., Smirnov, I., Keyes, T. J., Eccles, J. D., Rouhani, S. J., Peske, J. D., Derecki, N. C., Castle, D., Mandell, J. W., Lee, K. S., Harris, T. H., Kipnis, J. 2015; 523 (7560): 337-?

    Abstract

    One of the characteristics of the central nervous system is the lack of a classical lymphatic drainage system. Although it is now accepted that the central nervous system undergoes constant immune surveillance that takes place within the meningeal compartment, the mechanisms governing the entrance and exit of immune cells from the central nervous system remain poorly understood. In searching for T-cell gateways into and out of the meninges, we discovered functional lymphatic vessels lining the dural sinuses. These structures express all of the molecular hallmarks of lymphatic endothelial cells, are able to carry both fluid and immune cells from the cerebrospinal fluid, and are connected to the deep cervical lymph nodes. The unique location of these vessels may have impeded their discovery to date, thereby contributing to the long-held concept of the absence of lymphatic vasculature in the central nervous system. The discovery of the central nervous system lymphatic system may call for a reassessment of basic assumptions in neuroimmunology and sheds new light on the aetiology of neuroinflammatory and neurodegenerative diseases associated with immune system dysfunction.

    View details for DOI 10.1038/nature14432

    View details for Web of Science ID 000357950900040

    View details for PubMedID 26030524

    View details for PubMedCentralID PMC4506234