All Publications


  • Genome modelling and design across all domains of life with Evo 2. Nature Brixi, G., Durrant, M. G., Ku, J., Naghipourfar, M., Poli, M., Sun, G., Brockman, G., Chang, D., Fanton, A., Gonzalez, G. A., King, S. H., Li, D. B., Merchant, A. T., Nguyen, E., Ricci-Tam, C., Romero, D. W., Schmok, J. C., Taghibakhshi, A., Vorontsov, A., Yang, B., Deng, M., Gorton, L., Nguyen, N., Wang, N. K., Pearce, M. T., Simon, E., Adams, E., Amador, Z. J., Ashley, E. A., Baccus, S. A., Dai, H., Dillmann, S., Ermon, S., Guo, D., Herschl, M. H., Ilango, R., Janik, K., Lu, A. X., Mehta, R., Mofrad, M. R., Ng, M. Y., Pannu, J., Ré, C., St John, J., Sullivan, J., Tey, J., Viggiano, B., Zhu, K., Zynda, G., Balsam, D., Collison, P., Costa, A. B., Hernandez-Boussard, T., Ho, E., Liu, M. Y., McGrath, T., Powell, K., Pinglay, S., Burke, D. P., Goodarzi, H., Hsu, P. D., Hie, B. L. 2026

    Abstract

    All of life encodes information with DNA. Although tools for genome sequencing, synthesis and editing have transformed biological research, we still lack sufficient understanding of the immense complexity encoded by genomes to predict the effects of many classes of genomic changes or to intelligently compose new biological systems. Artificial intelligence models that learn information from genomic sequences across diverse organisms have increasingly advanced prediction and design capabilities1,2. Here we introduce Evo 2, a biological foundation model trained on 9 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life to have a 1 million token context window with single-nucleotide resolution. Evo 2 learns to accurately predict the functional impacts of genetic variation-from noncoding pathogenic mutations to clinically significant BRCA1 variants-without task-specific fine-tuning. Mechanistic interpretability analyses reveal that Evo 2 learns representations associated with biological features, including exon-intron boundaries, transcription factor binding sites, protein structural elements and prophage genomic regions. The generative abilities of Evo 2 produce mitochondrial, prokaryotic and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods. Evo 2 also generates experimentally validated chromatin accessibility patterns when guided by predictive models3,4 and inference-time search. We have made Evo 2 fully open, including model parameters, training code5, inference code and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity.

    View details for DOI 10.1038/s41586-026-10176-5

    View details for PubMedID 41781614

    View details for PubMedCentralID 12057570

  • Steering Protein Generative Models at Test-Time for Guided AAV2 Capsid Design. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing Viggiano, B., Lu, W. S., Zhang, X., Mille-Fragoso, L. S., Gao, X. J., Ashley, E., Wong, W. H. 2026; 31: 438-451

    Abstract

    Recent advances in protein generative models have created new opportunities for protein engineering. However, a significant challenge remains in effectively steering these models to generate sequences with specific, desired functionalities, especially when these properties are defined by "black-box" or non-differentiable fitness functions. To address this, we present ProVADA+, a model-agnostic framework that guides pretrained generative models at testtime without costly retraining. Our approach introduces a reinforcement learning-based adaptive masking technique (MADA-DUCB) that significantly accelerates convergence. We demonstrate this framework on the challenging task of designing novel Adeno-Associated Virus 2 (AAV2) capsids. By coupling a ProteinMPNN generative prior with a fine-tuned AAV viability oracle, our method successfully navigates the rugged fitness landscape where unguided random mutagenesis is ineffective-with prior experiments showing as few as 0.3% of variants with six or more mutations are viable. In its final iterations, ProVADA generated a pool of novel candidates with a mean viral selection score of 2.72, consistently scoring highly viable variants while maintaining a diverse range of sequence similarity to the wildtype sequence. Our results show that ProVADA provides a powerful and efficient framework for accelerating the design of proteins with complex, user-defined properties.

    View details for DOI 10.1142/9789819824755_0031

    View details for PubMedID 41758159

  • Steering Protein Generative Models at Test-Time for Guided AAV2 Capsid Design Viggiano, B., Lu, W., Zhang, X., Mille-Fragoso, L. S., Gao, X. J., Ashley, E., Wong, W. edited by Altman, R. B., Hunter, L., Ritchie, M. D., Murray, T., Klein, T. E. WORLD SCIENTIFIC PUBL CO PTE LTD. 2026: 438-451

    Abstract

    Recent advances in protein generative models have created new opportunities for protein engineering. However, a significant challenge remains in effectively steering these models to generate sequences with specific, desired functionalities, especially when these properties are defined by "black-box" or non-differentiable fitness functions. To address this, we present ProVADA+, a model-agnostic framework that guides pretrained generative models at testtime without costly retraining. Our approach introduces a reinforcement learning-based adaptive masking technique (MADA-DUCB) that significantly accelerates convergence. We demonstrate this framework on the challenging task of designing novel Adeno-Associated Virus 2 (AAV2) capsids. By coupling a ProteinMPNN generative prior with a fine-tuned AAV viability oracle, our method successfully navigates the rugged fitness landscape where unguided random mutagenesis is ineffective-with prior experiments showing as few as 0.3% of variants with six or more mutations are viable. In its final iterations, ProVADA generated a pool of novel candidates with a mean viral selection score of 2.72, consistently scoring highly viable variants while maintaining a diverse range of sequence similarity to the wildtype sequence. Our results show that ProVADA provides a powerful and efficient framework for accelerating the design of proteins with complex, user-defined properties.

    View details for Web of Science ID 001796654100031

    View details for PubMedID 41758159

  • Approach to the Postmarket Evaluation of Consumer Wearable Technologies. JAMA cardiology Pundi, K., Bhavnani, S., Seninger, C., Zuckerman, B., Paulsen, J., Aguel, F., Din, N., Viggiano, B., Yoo, R. M., Dalal, N., Go, A. S., Granger, C., Krumholz, H., Lacar, K., Li, R., Lin, S., Mahaffey, K. W., Mahoney, M., McCall, D., Hills, M. T., Harrington, R. A., Hernandez-Boussard, T., Saha, A., Shah, N., Turakhia, M. P. 2025

    Abstract

    Consumer wearable technologies have wide applications, including some that have US Food and Drug Administration clearance for health-related notifications. While wearable technologies may have premarket testing, validation, and safety evaluation as part of a regulatory authorization process, information on their postmarket use remains limited. The Stanford Center for Digital Health organized 2 pan-stakeholder think tank meetings to develop an organizing concept for empirical research on the postmarket evaluation of consumer-facing wearables.The postmarket evaluation of consumer wearables involves broad consideration of an individual consumer's journey from acquisition, intended and unintended use of the wearable, and access to health care resources on receipt of a notification. For individuals who do access the health care system, a wearable's downstream effects can be studied through appropriate clinical evaluation, delivery of guideline-directed treatments, shared decision-making in areas of clinical equipoise, and analysis of clinical end points and patient harms. Effective postmarket research draws from denominators appropriate to the clinical question, with clearly defined parameters for success and failure. Generalizability related to data completeness and reliability should also be considered. As patients increasingly integrate wearables into their health monitoring, cross-platform data sharing with a focus on privacy and data quality can drive patient-centered innovation and identify opportunities to bridge gaps in medical care.The think tank identified priorities in postmarket research, comprising the journey from consumer to patient and accounting for patient, clinician, health care delivery system, and societal impacts of consumer wearables. Overall, this approach serves not only to organize the study of consumer wearables but also to act as a guidepost for using real-world data in postmarket research.

    View details for DOI 10.1001/jamacardio.2025.3006

    View details for PubMedID 40928810

  • Scalable Approach to Consumer Wearable Postmarket Surveillance: Development and Validation Study. JMIR medical informatics Yoo, R. M., Viggiano, B. T., Pundi, K. N., Fries, J. A., Zahedivash, A., Podchiyska, T., Din, N., Shah, N. H. 2024; 12: e51171

    Abstract

    Background: With the capability to render prediagnoses, consumer wearables have the potential to affect subsequent diagnoses and the level of care in the health care delivery setting. Despite this, postmarket surveillance of consumer wearables has been hindered by the lack of codified terms in electronic health records (EHRs) to capture wearable use.Objective: We sought to develop a weak supervision-based approach to demonstrate the feasibility and efficacy of EHR-based postmarket surveillance on consumer wearables that render atrial fibrillation (AF) prediagnoses.Methods: We applied data programming, where labeling heuristics are expressed as code-based labeling functions, to detect incidents of AF prediagnoses. A labeler model was then derived from the predictions of the labeling functions using the Snorkel framework. The labeler model was applied to clinical notes to probabilistically label them, and the labeled notes were then used as a training set to fine-tune a classifier called Clinical-Longformer. The resulting classifier identified patients with an AF prediagnosis. A retrospective cohort study was conducted, where the baseline characteristics and subsequent care patterns of patients identified by the classifier were compared against those who did not receive a prediagnosis.Results: The labeler model derived from the labeling functions showed high accuracy (0.92; F1-score=0.77) on the training set. The classifier trained on the probabilistically labeled notes accurately identified patients with an AF prediagnosis (0.95; F1-score=0.83). The cohort study conducted using the constructed system carried enough statistical power to verify the key findings of the Apple Heart Study, which enrolled a much larger number of participants, where patients who received a prediagnosis tended to be older, male, and White with higher CHA2DS2-VASc (congestive heart failure, hypertension, age ≥75 years, diabetes, stroke, vascular disease, age 65-74 years, sex category) scores (P<.001). We also made a novel discovery that patients with a prediagnosis were more likely to use anticoagulants (525/1037, 50.63% vs 5936/16,560, 35.85%) and have an eventual AF diagnosis (305/1037, 29.41% vs 262/16,560, 1.58%). At the index diagnosis, the existence of a prediagnosis did not distinguish patients based on clinical characteristics, but did correlate with anticoagulant prescription (P=.004 for apixaban and P=.01 for rivaroxaban).Conclusions: Our work establishes the feasibility and efficacy of an EHR-based surveillance system for consumer wearables that render AF prediagnoses. Further work is necessary to generalize these findings for patient populations at other sites.

    View details for DOI 10.2196/51171

    View details for PubMedID 38596848

  • WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks Wornow, M., Narayan, A., Viggiano, B., Khare, I. S., Verma, T., Thompson, T., Hernandez, M., Sundar, S., Trujillo, C., Chawla, K., Lu, R., Shen, J., Nagaraj, D., Martinez, J., Agrawal, V., Hudson, A., Shah, N. H., Re, C. edited by Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C. NEURAL INFORMATION PROCESSING SYSTEMS (NIPS). 2024
  • A Metric for Quantification of Iodine Contrast Enhancement (Q-ICE) in Computed Tomography JOURNAL OF COMPUTER ASSISTED TOMOGRAPHY Szczykutowicz, T. P., Viggiano, B., Rose, S., Pickhardt, P. J., Lubner, M. G. 2021; 45 (6): 870-876

    Abstract

    Poor contrast enhancement is related to issues with examination execution, contrast prescription, computed tomography (CT) protocols, and patient conditions. Currently, our community has no metric to monitor true enhancement on routine single-phase examinations because this requires knowledge of both pre- and postcontrast CT number.We propose an automatable solution to quantifying contrast enhancement without requiring a dedicated noncontrast series.The difference in CT number between a target region in an enhanced and unenhanced image defines the metric "quantification of iodine contrast enhancement" (Q-ICE). Quantification of iodine contrast enhancement uses the noncontrast bolus tracking baseline image from routine abdominal examinations, which mitigates the need for a dedicated noncontrast series. We applied this method retrospectively to 312 patient livers from 2 sites between 2017 and 2020. Each site used a weight-based contrast injection protocol for weights 60 to 113 kg and a constant volume less than 60 kg and greater than 113 kg. Hypothesis testing was performed to compare Q-ICE between sites and detect Q-ICE dependence on weight and kilovoltage (kV).Mean Q-ICE differed between sites (P = 0.004) by 4.96 Hounsfield unit with 95% confidence interval (1.63-8.28), albeit this difference was roughly 2 times smaller than the SD in Q-ICE across patients at a single site. For patients between 60 and 113 kg, we did not observe evidence of Q-ICE varying with patient weight (P = 0.920 and 0.064 for 120 and 140 kV, respectively). The Q-ICE did vary with patient weight for patients less than 60 kg (P = 0.003) and greater than 113 kg (P = 0.04). We observed a roughly 10 Hounsfield unit reduction in Q-ICE liver for patients scanned with 140 versus 120 kV. We observed several underenhancing examinations with an arterial phase appearance motivating our CT protocol optimization team to consider increasing the delay for slowly enhancing patients.A quality metric for quantifying CT contrast enhancement was developed and suggested tangible opportunities for quality improvement and potential financial savings.

    View details for DOI 10.1097/RCT.0000000000001215

    View details for Web of Science ID 000719012000011

    View details for PubMedID 34469906

  • Applying a New CT Quality Metric in Radiology: How CT Pulmonary Angiography Repeat Rates Compare Across Institutions JOURNAL OF THE AMERICAN COLLEGE OF RADIOLOGY Rose, S., Viggiano, B., Bour, R., Bartels, C., Kanne, J. P., Szczykutowicz, T. P. 2021; 18 (7): 962-968

    Abstract

    To quantify overall CT repeat and reject rates at five institutions and investigate repeat and reject rates for CT pulmonary angiography (CTPA).In this retrospective study, we apply an automated repeat rate analysis algorithm to 103,752 patient examinations performed at five institutions from July 2017 to August 2019. The algorithm identifies repeated scans for specific scanner and protocol combinations. For each institution, we compared repeat rates for CTPA to all other CT protocols. We used logistic regression and analysis of deviance to compare CTPA repeat rates across institutions and size-based protocols.Of 103,752 examinations, 1,447 contained repeated helical scans (1.4%). Overall repeat rates differed across institutions (P < .001) ranging from 0.8% to 1.8%. Large-patient CTPA repeat rates ranged from 3.0% to 11.2% with the odds (95% confidence intervals) of a repeat being 4.8 (3.5-6.6) times higher for large- relative to medium-patient CTPA protocols. CTPA repeat rates were elevated relative to all other CT protocols at four of five institutions, with strong evidence of an effect at two institutions (P < .001 for each; odds ratios: 2.0 [1.6-2.6] and 6.2 [4.4-8.9]) and somewhat weaker evidence at the others (P = .005 and P = 0.011; odds ratios: 2.2 [1.3-3.8] and 3.7 [1.5-9.1], respectively). Accounting for size-based protocols, CTPA repeat rates differed across institutions (P < .001).The results indicate low overall repeat rates (<2%) with CTPA rates elevated relative to other protocols. Large-patient CTPA rates were highest (eg, 11.2% at one institution). Differences in repeat rates across institutions suggest the potential for quality improvement.

    View details for DOI 10.1016/j.jacr.2021.02.014

    View details for Web of Science ID 000668925200013

    View details for PubMedID 33741373

  • Effect of contrast agent administration on water equivalent diameter in CT MEDICAL PHYSICS Viggiano, B., Rose, S., Szczykutowicz, T. P. 2021; 48 (3): 1117-1124

    Abstract

    Water equivalent diameter (WED) is the preferred surrogate for patient size in computed tomography (CT). It is better than geometric size surrogates and patient weight/height/BMI/age because it correlates the best with x-ray attenuation. The administration of oral/IV contrast agents increases a patient's attenuation and should therefore increase WED. Here we study the clinically relevant effect of oral and IV contrast agent on WED.We pulled 1703 routine adult abdominal/pelvis cases acquired at 100, 120, and 140 kV from our PACS under retrospective IRB approval. One hundred and forty cases cases had no oral or IV contrast (NONCON), 285 had just IV contrast (IV), 107 had just oral contrast (ORAL), and 1171 had both oral and IV contrast (BOTH). For each case, we measured the water equivalent and effective diameter (ED) from axial CT images. We plotted the WED versus the ED for each class of contrast. We used a linear regression model and omnibus F-test to determine if significant differences between WED distributions existed between the contrast groups for each kV. We then performed a post hoc analysis to determine if any significant differences existed in pairwise comparisons of the different contrast groups. Bonferroni correction was used to account for multiple comparisons.We found statistically significant changes at 100 and 120 kV with a maximum change of 2.1 mm. We measured a ~25 mm spread (i.e., prediction interval) of WEDs within all four contrast groups.While our sample size was large enough to detect statistically significant differences between some of the contrast groups, the differences were clinically irrelevant when one considers that the change in size-specific dose estimate (SSDE) caused by our observations is roughly 1%.

    View details for DOI 10.1002/mp.14721

    View details for Web of Science ID 000617057900001

    View details for PubMedID 33440020

  • A Multiinstitutional Study on Wasted CT Scans for Over 60,000 Patients AMERICAN JOURNAL OF ROENTGENOLOGY Rose, S., Viggiano, B., Bour, R., Bartels, C., Szczykutowicz, T. 2020; 215 (5): 1123-1128

    Abstract

    OBJECTIVE. Repeated imaging is an unnecessary source of patient radiation exposure, a detriment to patient satisfaction, and a waste of time and money. Although analysis of rates of repeated and rejected images is mandated in mammography and recommended in radiography, the available data on these rates for CT are limited. MATERIALS AND METHODS. In this retrospective study, an automated repeat-reject rate analysis algorithm was used to quantify repeat rates from 61,102 patient examinations obtained between 2015 and 2018. The algorithm used DICOM metadata to identify repeat acquisitions. We quantified rates for one academic site and one rural site. The method allows scanner-, technologist-, protocol-, and indication-specific rates to be determined. Positive predictive values and sensitivity were estimated for correctly identifying and classifying repeat acquisitions. Repeat rates were compared between sites to identify areas for targeted technologist training. RESULTS. Of 61,102 examinations, 4676 instances of repeat scanning contributed excess radiation dose to patients. Estimated helical overlap repeat rates were 1.4% (95% CI, 1.2-1.6%) for the rural site and 1.1% (95% CI, 1.0-1.2%) for the academic site. Significant differences in rates of repeat imaging required because of bolus tracking (11.6% vs 4.3%; p < 0.001) and helical extension (3.3% vs 1.8%; p < 0.001) were observed between sites. Positive predictive values ranged from 91% to 99% depending on the reason for repeat imaging and site location. Sensitivity of the algorithm was 92% (95% CI, 87-96%). Rates tended to be highest for emergent imaging procedures and exceeded 9% for certain protocols. CONCLUSION. Our multiinstitutional automated quantification of repeat rates for CT provided a useful metric for unnecessary radiation exposure and identification of technologists in need of training.

    View details for DOI 10.2214/AJR.19.22604

    View details for Web of Science ID 000582043500021

    View details for PubMedID 32960668