All Publications


  • Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving. AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science Shidara, K., Prem, P., Kim, J., Podlasek, A., Liu, F., Alaa, A., Bernardo, D. 2026; 2026: 428-437

    Abstract

    Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC: physician average accuracy was 66% [95% CI: 56-76%] and the top performing model was Claude with 75% accuracy [95% CI: 74-76%]. On questions most commonly missed by physicians, the top 5 performing models answered 50% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.

    View details for PubMedID 42317846

    View details for PubMedCentralID PMC13274283

  • Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving. AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science Shidara, K., Prem, P., Kim, J., Podlasek, A., Liu, F., Alaa, A., Bernardo, D. 2026; 2026: 428-437

    Abstract

    Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC: physician average accuracy was 66% [95% CI: 56-76%] and the top performing model was Claude with 75% accuracy [95% CI: 74-76%]. On questions most commonly missed by physicians, the top 5 performing models answered 50% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.

    View details for PubMedID 42317846

  • Limitations of large language models in clinical problem-solving arising from inflexible reasoning. Scientific reports Kim, J., Podlasek, A., Shidara, K., Liu, F., Alaa, A., Bernardo, D. 2025; 15 (1): 39426

    Abstract

    Large Language Models (LLMs) have attained human-level accuracy on medical question-answer (QA) benchmarks. However, their limitations in navigating clinical scenarios requiring flexible reasoning have recently been shown, raising concerns about the robustness and generalizability of LLM reasoning across diverse, real-world medical tasks. To probe potential LLM failure modes in clinical problem-solving, we present the medical abstraction and reasoning corpus (mARC-QA). mARC-QA assesses clinical reasoning through scenarios designed to exploit the Einstellung effect-the fixation of thought arising from prior experience, targeting LLM inductive biases toward inflexible pattern matching from their training data rather than engaging in flexible reasoning. We find that LLMs, including current state-of-the-art o1, Gemini, Claude, and DeepSeek models, perform poorly compared to physicians on mARC-QA, often demonstrating lack of commonsense medical reasoning and a propensity to hallucinate. In addition, uncertainty estimation analyses indicate that LLMs exhibit overconfidence in their answers, despite their limited accuracy. The failure modes revealed by mARC-QA in LLM medical reasoning underscore the need to exercise caution when deploying these models in clinical settings.

    View details for DOI 10.1038/s41598-025-22940-0

    View details for PubMedID 41219270

    View details for PubMedCentralID PMC12606185

  • Short-horizon neonatal seizure prediction using EEG-based deep learning. PLOS digital health Kim, J., Amorim, E., Rao, V. R., Glass, H. C., Bernardo, D. 2025; 4 (7): e0000890

    Abstract

    Strategies to predict neonatal seizure risk have typically focused on long-term static predictions with prediction horizons spanning days during the acute postnatal period. Higher temporal resolution or short-horizon neonatal seizure prediction, on the time-frame of minutes, remains unexplored. Here, we investigated quantitative electroencephalography (QEEG) based deep learning (DL) for short-horizon seizure prediction. We used two publicly available EEG seizure datasets with a total of 132 neonates containing a total of 281 hours of EEG data. We benchmarked current state-of-the-art time-series DL methods for seizure prediction, identifying convolutional LSTM (ConvLSTM) as having the strongest performance at preictal state classification. We assessed ConvLSTM performance in a seizure alarm system over varying short-range (1-7 minutes) seizure prediction horizons (SPH) and seizure occurrence periods (SOP) and identified optimal performance at SPH 3 min and SOP 7 min, with AUROC 0.8. At 80% sensitivity, false detection rate was 0.68 events/hour with time-in-warning of 0.36. Model calibration was moderate, with an expected calibration error of 0.106. These findings establish the feasibility of short-horizon neonatal seizure prediction and warrant the need for further validation.

    View details for DOI 10.1371/journal.pdig.0000890

    View details for PubMedID 40644380

    View details for PubMedCentralID PMC12250315

  • Machine learning for forecasting initial seizure onset in neonatal hypoxic-ischemic encephalopathy. Epilepsia Bernardo, D., Kim, J., Cornet, M. C., Numis, A. L., Scheffler, A., Rao, V. R., Amorim, E., Glass, H. C. 2024

    Abstract

    This study was undertaken to develop a machine learning (ML) model to forecast initial seizure onset in neonatal hypoxic-ischemic encephalopathy (HIE) utilizing clinical and quantitative electroencephalogram (QEEG) features.We developed a gradient boosting ML model (Neo-GB) that utilizes clinical features and QEEG to forecast time-dependent seizure risk. Clinical variables included cord blood gas values, Apgar scores, gestational age at birth, postmenstrual age (PMA), postnatal age, and birth weight. QEEG features included statistical moments, spectral power, and recurrence quantification analysis (RQA) features. We trained and evaluated Neo-GB on a University of California, San Francisco (UCSF) neonatal HIE dataset, augmenting training with publicly available neonatal electroencephalogram (EEG) datasets from Cork University and Helsinki University Hospitals. We assessed the performance of Neo-GB at providing dynamic and static forecasts with diagnostic performance metrics and incident/dynamic area under the receiver operating characteristic curve (iAUC) analyses. Model explanations were performed to assess contributions of QEEG features and channels to model predictions.The UCSF dataset included 60 neonates with HIE (30 with seizures). In subject-level static forecasting at 30 min after EEG initiation, baseline Neo-GB without time-dependent features had an area under the receiver operating characteristic curve (AUROC) of .76 and Neo-GB with time-dependent features had an AUROC of .89. In time-dependent evaluation of the initial seizure onset within a 24-h seizure occurrence period, dynamic forecast with Neo-GB demonstrated median iAUC = .79 (interquartile range [IQR] .75-.82) and concordance index (C-index) = .82, whereas baseline static forecast at 30 min demonstrated median iAUC = .75 (IQR .72-.76) and C-index = .69. Model explanation analysis revealed that spectral power, PMA, RQA, and cord blood gas values made the strongest contributions in driving Neo-GB predictions. Within the most influential EEG channels, as the preictal period advanced toward eventual seizure, there was an upward trend in broadband spectral power.This study demonstrates an ML model that combines QEEG with clinical features to forecast time-dependent risk of initial seizure onset in neonatal HIE. Spectral power evolution is an early EEG marker of seizure risk in neonatal HIE.

    View details for DOI 10.1111/epi.18163

    View details for PubMedID 39495029

  • Statistical Considerations and Tools to Improve Histopathologic Protocols with Spectroscopic Imaging. Applied spectroscopy Mittal, S., Kim, J., Bhargava, R. 2022; 76 (4): 428-438

    Abstract

    Advances in infrared (IR) spectroscopic imaging instrumentation and data science now present unique opportunities for large validation studies of the concept of histopathology using spectral data. In this study, we examine the discrimination potential of IR metrics for different histologic classes to estimate the sample size needed for designing validation studies to achieve a given statistical power and statistical significance. Next, we present an automated annotation transfer tool that can allow large-scale training/validation, overcoming the limitations of sparse ground truth data with current manual approaches by providing a tool to transfer pathologist annotations from stained images to IR images across diagnostic categories. Finally, the results of a combination of supervised and unsupervised analysis provide a scheme to identify diagnostic groups/patterns and isolating pure chemical pixels for each class to better train complex histopathological models. Together, these methods provide essential tools to take advantage of the emerging capabilities to record and utilize large spectroscopic imaging datasets.

    View details for DOI 10.1177/00037028211066327

    View details for PubMedID 35296146

    View details for PubMedCentralID PMC9202564

  • The Slip-Pad: A Haptic Display Using Interleaved Belts to Simulate Lateral and Rotational Slip Ho, C., Kim, J., Patil, S., Goldberg, K. edited by Colgate, J. E., Tan, H. Z., Choi, S. M., Gerling, G. J. IEEE. 2015: 189-195
  • Autonomous Multilateral Debridement with the Raven Surgical Robot Kehoe, B., Kahn, G., Mahler, J., Kim, J., Lee, A., Lee, A., Nakagawa, K., Patil, S., Boyd, W., Abbeel, P., Goldberg, K., IEEE IEEE. 2014: 1432-1439
  • Microparticle trapping in an ultrasonic Bessel beam. Applied physics letters Choe, Y., Kim, J. W., Shung, K. K., Kim, E. S. 2011; 99 (23): 233704-2337043

    Abstract

    This paper describes an acoustic trap consisting of a multi-foci Fresnel lens on 127 μm thick lead zirconate titanate sheet. The multi-foci Fresnel lens was designed to have similar working mechanism to an Axicon lens and generates an acoustic Bessel beam, and has negative axial radiation force capable of trapping one or more microparticle(s). The fabricated acoustic tweezers trapped lipid particles ranging in diameter from 50 to 200 μm and microspheres ranging in diameter from 70 to 90 μm at a distance of 2 to 5 mm from the tweezers without any contact between the transducer and microparticles.

    View details for DOI 10.1063/1.3665615

    View details for PubMedID 22247566

    View details for PubMedCentralID PMC3253744