Selin Everett
MD Student with Scholarly Concentration in Informatics & Data-Driven Medicine / Surgery, expected graduation Spring 2029
Stanford Student Employee, Technology & Digital Solutions
All Publications
-
Response to Letter to the Editor re: "A contemporary analysis of single- and multi-stage hypospadias repair in the United States".
Journal of pediatric urology
2026: 105967
View details for DOI 10.1016/j.jpurol.2026.105967
View details for PubMedID 42086426
-
A contemporary analysis of single- and multi-stage hypospadias repair in the United States.
Journal of pediatric urology
2026; 22 (4): 105898
Abstract
Hypospadias is a common congenital anomaly seen in newborn males. The phenotypic spectrum is wide, and complication rates are high, particularly in severe cases. Multi-stage approaches to surgical management have gained traction over recent years. However, a comprehensive understanding of national practice patterns is limited.To perform a contemporary analysis of hypospadias repair patterns and complication rates in the United States, with a focus on surgical staging.The Pediatric Health Information System (PHIS) database was queried to create a cohort of pediatric patients under 5 years of age who underwent single- or multi-stage hypospadias repair between January 1, 2016 and June 30, 2025. ICD-10 and CPT codes were used for cohort creation. Patients with inconsistent coding were excluded. Patient demographics and features of hypospadias phenotype were explored. Comparisons between single- and multi-stage patients, longitudinal trends in staging, and factors impacting complication rates were analyzed.We identified 25,989 patients at 45 children's hospitals. Overall, 96.5% of patients underwent single-stage repair at a median age of 9.3 months, while 3.5% underwent multi-stage repair, starting at 15.4 months. For proximal cases, 35.5% were managed in a multi-stage fashion. On longitudinal analysis, rates of multi-stage repair for proximal hypospadias have significantly increased over time (p = 0.003). Compared to single-stage patients, multi-stage patients were more likely to be non-White, receive care in the Northeast or Midwest, have a proximal meatus and associated chordee, and undergo grafting or complex scrotoplasty. On multivariate analysis of multi-stage patients, increasing age was a significant predictor of complications (p < 0.0001), while Midwest region (p = 0.003) and non-Hispanic Black (p = 0.02) and Hispanic race-ethnicity (p = 0.01) were protective. Meatal location was not a significant factor impacting complications after multi-stage repair (p = 0.28); however, in single-stage patients, complication rates did increase with a more proximal meatus (p < 0.0001).Use of multi-stage repairs for proximal hypospadias has increased over the past decade in the United States. Single- and multi-stage patients differ in terms of hypospadias phenotype, demographics, and factors associated with complication rates.
View details for DOI 10.1016/j.jpurol.2026.105898
View details for PubMedID 41996974
-
From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis.
NPJ digital medicine
2026
Abstract
Early studies of large language models (LLMs) in clinical settings have largely treated artificial intelligence (AI) as a tool rather than an active collaborator. As LLMs demonstrate expert-level diagnostic performance, the focus shifts from whether AI can offer valuable suggestions to how it integrates into physicians' diagnostic workflows. We conducted a randomized controlled trial (n = 70 clinicians) to assess a custom system designed for collaborative diagnostic reasoning. The design involved independent diagnostic assessments by the clinician and AI, followed by an AI-generated synthesis integrating both perspectives, highlighting agreements, disagreements, and offering commentary. We evaluated two collaborative workflows: AI as first opinion (preceding clinician) and AI as second opinion (following clinician). Both improved clinician diagnostic accuracy over conventional resources, (85% and 82% vs. 75%). Performance was comparable across workflows and not statistically different from AI-alone accuracy (90%), highlighting the potential of collaborative AI to complement clinician expertise. Qualitative analyses illustrate how workflow design shapes human-AI interaction. C: NCT06911645.
View details for DOI 10.1038/s41746-026-02545-1
View details for PubMedID 41851268
-
Automated Evaluation of Large Language Model Response Concordance with Human Specialist Responses on Physician-to-Physician eConsult Cases.
Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
2026; 31: 372-387
Abstract
Specialist consults in primary care and inpatient settings typically address complex clinical questions beyond standard guidelines. eConsults have been developed as a way for specialist physicians to review cases asynchronously and provide clinical answers without a formal patient encounter. Meanwhile, large language models (LLMs) have approached human-level performance on structured clinical tasks, but their real-world effectiveness requires evaluation, which is bottlenecked by time-intensive manual physician review. To address this, we evaluate two automated methods: LLM-as-judge and a decompose-thenverify framework that breaks down AI answers into verifiable claims against human eConsult responses. Using 40 real-world physician-to-physician eConsults, we compared AI-generated responses to human answers using both physician raters and automated tools. LLM-as-judge outperformed decompose-then-verify, achieving human-level concordance assessment with F1-score of 0.89 (95% CI: 0.750, 0.960) and Cohen's kappa of 0.75 (95% CI 0.47,0.90) , comparable to physician inter-rater agreement κ = 0.69-0.90 (95% CI 0.43-1.0).
View details for DOI 10.1142/9789819824755_0026
View details for PubMedID 41758154
-
Automated Evaluation of Large Language Model Response Concordance with Human Specialist Responses on Physician-to-Physician eConsult Cases
edited by Altman, R. B., Hunter, L., Ritchie, M. D., Murray, T., Klein, T. E.
WORLD SCIENTIFIC PUBL CO PTE LTD. 2026: 372-387
Abstract
Specialist consults in primary care and inpatient settings typically address complex clinical questions beyond standard guidelines. eConsults have been developed as a way for specialist physicians to review cases asynchronously and provide clinical answers without a formal patient encounter. Meanwhile, large language models (LLMs) have approached human-level performance on structured clinical tasks, but their real-world effectiveness requires evaluation, which is bottlenecked by time-intensive manual physician review. To address this, we evaluate two automated methods: LLM-as-judge and a decompose-thenverify framework that breaks down AI answers into verifiable claims against human eConsult responses. Using 40 real-world physician-to-physician eConsults, we compared AI-generated responses to human answers using both physician raters and automated tools. LLM-as-judge outperformed decompose-then-verify, achieving human-level concordance assessment with F1-score of 0.89 (95% CI: 0.750, 0.960) and Cohen's kappa of 0.75 (95% CI 0.47,0.90) , comparable to physician inter-rater agreement κ = 0.69-0.90 (95% CI 0.43-1.0).
View details for Web of Science ID 001796654100026
View details for PubMedID 41758154
-
Genetic disorders and associated morbidity, mortality, and congenital anomalies in preterm infants born at less than 34 weeks of gestation.
BMC pediatrics
2025
Abstract
Genetic disorders are recognized as key contributors to morbidity, mortality, and congenital anomalies in term infants. However, the rates of diagnosis and association with morbidity, mortality, and congenital anomalies in preterm infants are poorly characterized. We sought to determine rates of diagnosis of genetic disorders in preterm infants and to define the association of genetic disorders with morbidity, mortality, and congenital anomalies.This was a multicenter observational cohort study conducted in neonatal intensive care units in the Pediatrix Clinical Data Warehouse. Infants born from 23 to 0/7 to 33 and 6/7 weeks of gestation, admitted to 374 U.S. community and academic neonatal intensive care units from 2000 to 2020 were included. Infants transferred after birth or prior to discharge were excluded. We analyzed diagnosis of genetic disorders; predischarge morbidity (including acute kidney injury, bronchopulmonary dysplasia, necrotizing enterocolitis, sepsis, shock, severe retinopathy, and intracranial hemorrhage); mortality; and presence of congenital anomalies.Among 323,770 early preterm infants analyzed, 4,196 (1.3%) were diagnosed with one of twenty genetic disorders. Single gene disorders were identified in 2,250 (0.7%) infants, copy number variants in 88 (0.03%) infants, and aneuploidies in 1,885 (0.6%) infants. Morbidity, mortality, and congenital anomalies occurred in 1,319 (31.4%), 566 (13.5%), and 1,041 (24.8%) infants with genetic disorders compared to 77,957 (24.5%), 15,240 (4.7%), and 9,455 (3.0%) infants without genetic disorders. Common aneuploidies accounted for most of these associations. However, morbidity, mortality, and congenital anomalies were also significantly more common in early preterm infants with single gene disorders and pathogenic copy number variants. We did not detect meaningful differences in diagnostic rates of genetic disorders over the study period.1.3% of early preterm infants were diagnosed with genetic disorders. Genetic disorders were strongly associated with morbidity, mortality, and congenital anomalies. Clinicians should strongly consider genetic evaluation in early preterm infants with morbidity, mortality, or congenital anomalies. Prospective research is needed to determine the true prevalence of genetic disorders in this high-risk population.
View details for DOI 10.1186/s12887-025-06373-2
View details for PubMedID 41331421
-
From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis.
medRxiv : the preprint server for health sciences
2025
Abstract
Early studies of large language models (LLMs) in clinical settings have largely treated artificial intelligence (AI) as a tool rather than an active collaborator. As LLMs now demonstrate expert-level diagnostic performance, the focus shifts from whether AI can offer valuable suggestions to how it can be effectively integrated into physicians' diagnostic workflows. We conducted a randomized controlled trial (n=70 clinicians) to evaluate the value of employing a custom GPT system designed to engage collaboratively with clinicians on diagnostic reasoning challenges. The collaborative design began with independent diagnostic assessments from both the clinician and the AI. These were then combined in an AI-generated synthesis that integrated the two perspectives, highlighting points of agreement and disagreement and offering commentary on each. We evaluated two workflow variants: one where the AI provided an initial opinion (AI-first), and another where it followed the clinician's assessment (AI-second). Clinicians using either collaborative workflow outperformed those using traditional tools, achieving average accuracies of 85% (AI-first) and 82% (AI-second), compared to 75% with traditional resources (p < 0.0004 and p < 0.00001; mean differences = 9.8% and 6.8%; 95% CI = 4.6%-15% and 4.0%-9.6%). Performance did not differ significantly between workflows or from the AI-alone score of 90%. These results underscore the value of collaborative AI systems that complement clinician expertise and foster effective coordination between human and machine reasoning in diagnostic decision-making.
View details for DOI 10.1101/2025.06.07.25329176
View details for PubMedID 40502554
-
Staged Hybrid Treatment of a Large Aneurysmal Pulmonary Sequestration With Thoracic Endovascular Aortic Repair Followed by Lobectomy
Annals of Thoracic Surgery Short Reports
2025
View details for DOI 10.1016/j.atssr.2025.09.036