Bio


I'm a Stanford Data Science Fellow and postdoc with Rhiju Das at the Department of Biochemistry. I build lab-in-the-loop AI for RNA biology, pairing deep learning with wet-lab experiments at scale.

I did my PhD in Computer Science at the University of Cambridge with Pietro Liò, on geometric deep learning for molecular design. I built gRNAde, the first 3D generative model for RNA, and validated it in the wet lab as a visiting researcher in Phil Holliger's group at the MRC LMB. I've also interned at Prescient Design (Genentech) and FAIR Chemistry (Meta AI), and my work has been recognized by the Qualcomm Innovation Fellowship and the A*STAR National Science Scholarship.

Honors & Awards


  • Stanford Data Science Fellowship, Stanford Data Science (2026)
  • Qualcomm Innovation Fellowship, Qualcomm Inc. (2024)
  • National Science Scholarship, A*STAR, Singapore (2021)

Professional Education


  • Ph.D., University of Cambridge, UK, Computer Science (2026)
  • B.Eng., Nanyang Technological University, Singapore, Computer Science, Valedictorian (2019)

Stanford Advisors


All Publications


  • De novo design of RNA pseudoknots with deep learning. Science (New York, N.Y.) Townley, J., Kladwang, W., Baker, D., Blair, H. M., Choe, C. A., El Nesr, G., Favor, A., Fisker, E., Haack, D. B., He, S., Hingey, J., Huang, P. S., Huang, R., Joshi, C. K., Karagianes, T., Kubaney, A., Liò, P., Mancino, A., Romano, J., Rudolfs, B., Spellmon, N., Toor, N., Verma, J., Wu, V., Yu, Z., Participants, E., Das, R. 2026; 393 (6814): 931-937

    Abstract

    RNA design has been hindered by the limited accuracy of three-dimensional (3D) structure prediction. In this study, we show that intricate RNA structures can be generated with current deep learning tools through accurate de novo design of pseudoknot secondary structures. In an Eterna competition involving 57 pseudoknots, generative artificial intelligence (AI) methods matched experienced human designers in solving most blind challenges, evaluated by single nucleotide-resolution chemical mapping, compensatory mutagenesis, and cryo-electron microscopy. AI-generated molecules with accurate secondary structures formed well-ordered 3D folds stabilized by noncanonical tertiary interactions not modeled during design. Success was guided by an RNet foundation model trained on prior chemical mapping data, suggesting that some difficult RNA design tasks may be tractable without first solving RNA 3D structure prediction.

    View details for DOI 10.1126/science.aeg6829

    View details for PubMedID 42658941

  • Understanding biology with machine learning: compression, intelligibility, and dependency ARTIFICIAL INTELLIGENCE IN THE LIFE SCIENCES Lawrence, E., El-Shazly, A., Seal, S., Joshi, C. K., Lio, P., Bender, A., Singh, S., Sormanni, P., Spjuth, O., Greenig, M. 2026; 9
  • De novo design of RNA pseudoknots with deep learning. bioRxiv : the preprint server for biology Townley, J., Kladwang, W., Baker, D., Blair, H. M., Choe, C. A., Nesr, G. E., Favor, A., Fisker, E., Haack, D. B., He, S., Hingey, J., Huang, P. S., Huang, R., Joshi, C. K., Karagianes, T., Kubaney, A., Liò, P., Mancino, A., Romano, J., Rudolfs, B., Spellmon, N., Toor, N., Wu, V., Yu, Z., Participants, E., Das, R. 2026

    Abstract

    RNA design has been hindered by the limited accuracy of 3D structure prediction. Here, we show that intricate RNA structures can be generated with current deep learning tools through accurate de novo design of pseudoknot secondary structures. In an Eterna competition involving 57 pseudoknots, generative AI methods matched experienced human designers in solving most blind challenges, evaluated by single-nucleotide-resolution chemical mapping, compensatory mutagenesis, and cryogenic electron microscopy. Unexpectedly, AI-generated molecules with accurate secondary structures formed well-ordered 3D folds stabilized by noncanonical tertiary interactions not modeled during design. Success was guided by a RNet foundation model trained on prior chemical mapping data, suggesting that some difficult RNA design tasks may be tractable without first solving RNA 3D structure prediction.

    View details for DOI 10.64898/2026.05.21.726960

    View details for PubMedID 42239184

    View details for PubMedCentralID PMC13228335

  • Template-based RNA structure prediction advanced through a blind code competition. bioRxiv : the preprint server for biology Lee, Y., He, S., Oda, T., Rao, G. J., Kim, Y., Kim, R., Kim, H., Heng, C. K., Kowerko, D., Li, H., Nguyen, H., Sampathkumar, A., Gómez, R. E., Chen, M., Yoshizawa, A., Kuraishi, S., Ogawa, K., Zou, S., Paullier, A., Zhao, B., Chen, H. L., Hsu, T. A., Hirano, T., Chiu, W., Gezelle, J. G., Haack, D., Hong, Y., Jadhav, S., Koirala, D., Kretsch, R. C., Lewicka, A., Li, S., Marcia, M., Piccirilli, J., Rudolfs, B., Srivastava, Y., Steckelberg, A. L., Su, Z., Toor, N., Wang, L., Yang, Z., Zhang, K., Zou, J., Baker, D., Chen, S. J., Demkin, M., Favor, A., Hummer, A. M., Joshi, C. K., Kryshtafovych, A., Küçükbenli, E., Miao, Z., Moult, J., Munley, C., Reade, W., Viel, T., Westhof, E., Zhang, S., Das, R. 2025

    Abstract

    Automatically predicting RNA 3D structure from sequence remains an unsolved challenge in biology and biotechnology. Here, we describe a Kaggle code competition engaging over 1700 teams and 43 previously unreleased structures to tackle this challenge. The top three submitted algorithms achieved scores within statistical error of the winners of the recent CASP16 competition. Unexpectedly, the top Kaggle strategy involved a pipeline for discovering 3D templates, without the use of deep learning. We integrated this template-modeling pipeline and other Kaggle strategies to develop a single model RNAPro that retrospectively outperformed individual Kaggle models on the same test set. These results suggest a growing importance of template-based modeling in RNA structure prediction.

    View details for DOI 10.64898/2025.12.30.696949

    View details for PubMedID 41509375

    View details for PubMedCentralID PMC12776560

  • RNA-FrameFlow: Flow Matching for de novo 3D RNA Backbone Design. ArXiv Anand, R., Joshi, C. K., Morehead, A., Jamasb, A. R., Harris, C., Mathis, S. V., Didi, K., Ying, R., Hooi, B., Lio, P. 2025

    Abstract

    We introduce RNA-FrameFlow, the first generative model for 3D RNA backbone design. We build upon S E ( 3 ) flow matching for protein backbone generation and establish protocols for data preparation and evaluation to address unique challenges posed by RNA modeling. We formulate RNA structures as a set of rigid-body frames and associated loss functions which account for larger, more conformationally flexible RNA backbones (13 atoms per nucleotide) vs. proteins (4 atoms per residue). Toward tackling the lack of diversity in 3D RNA datasets, we explore training with structural clustering and cropping augmentations. Additionally, we define a suite of evaluation metrics to measure whether the generated RNA structures are globally self-consistent (via inverse folding followed by forward folding) and locally recover RNA-specific structural descriptors. The most performant version of RNA-FrameFlow generates locally realistic RNA backbones of 40-150 nucleotides, over 40% of which pass our validity criteria as measured by a self-consistency TM-score ≥ 0.45, at which two RNAs have the same global fold. Open-source code: github.com/rish-16/rna-backbone-design.

    View details for PubMedID 38947930

  • Machine Learning for Toxicity Prediction Using Chemical Structures: Pillars for Success in the Real World CHEMICAL RESEARCH IN TOXICOLOGY Seal, S., Mahale, M., Garcia-Ortegon, M., Joshi, C. K., Hosseini-Gerami, L., Beatson, A., Greenig, M., Shekhar, M., Patra, A., Weis, C., Mehrjou, A., Badre, A., Paisley, B., Lowe, R., Singh, S., Shah, F., Johannesson, B., Williams, D., Rouquie, D., Clevert, D., Schwab, P., Richmond, N., Nicolaou, C. A., Gonzalez, R. J., Naven, R., Schramm, C., Vidler, L. R., Mansouri, K., Walters, W., Wilk, D., Spjuth, O., Carpenter, A. E., Bender, A. 2025; 38 (5): 759-807

    Abstract

    Machine learning (ML) is increasingly valuable for predicting molecular properties and toxicity in drug discovery. However, toxicity-related end points have always been challenging to evaluate experimentally with respect to in vivo translation due to the required resources for human and animal studies; this has impacted data availability in the field. ML can augment or even potentially replace traditional experimental processes depending on the project phase and specific goals of the prediction. For instance, models can be used to select promising compounds for on-target effects or to deselect those with undesirable characteristics (e.g., off-target or ineffective due to unfavorable pharmacokinetics). However, reliance on ML is not without risks, due to biases stemming from nonrepresentative training data, incompatible choice of algorithm to represent the underlying data, or poor model building and validation approaches. This might lead to inaccurate predictions, misinterpretation of the confidence in ML predictions, and ultimately suboptimal decision-making. Hence, understanding the predictive validity of ML models is of utmost importance to enable faster drug development timelines while improving the quality of decisions. This perspective emphasizes the need to enhance the understanding and application of machine learning models in drug discovery, focusing on well-defined data sets for toxicity prediction based on small molecule structures. We focus on five crucial pillars for success with ML-driven molecular property and toxicity prediction: (1) data set selection, (2) structural representations, (3) model algorithm, (4) model validation, and (5) translation of predictions to decision-making. Understanding these key pillars will foster collaboration and coordination between ML researchers and toxicologists, which will help to advance drug discovery and development.

    View details for DOI 10.1021/acs.chemrestox.5c00033

    View details for Web of Science ID 001481083500001

    View details for PubMedID 40314361

    View details for PubMedCentralID PMC12093382

  • gRNAde: Geometric Deep Learning for 3D RNA inverse design. bioRxiv : the preprint server for biology Joshi, C. K., Jamasb, A. R., Vinas, R., Harris, C., Mathis, S. V., Morehead, A., Anand, R., Lio, P. 2025

    Abstract

    Computational RNA design tasks are often posed as inverse problems, where sequences are designed based on adopting a single desired secondary structure without considering 3D conformational diversity. We introduce gRNAde, a geometric RNA design pipeline operating on 3D RNA backbones to design sequences that explicitly account for structure and dynamics. gRNAde uses a multi-state Graph Neural Network and autoregressive decoding to generates candidate RNA sequences conditioned on one or more 3D backbone structures where the identities of the bases are unknown. On a single-state fixed backbone re-design benchmark of 14 RNA structures from the PDB identified by Das et al. (2010), gRNAde obtains higher native sequence recovery rates (56% on average) compared to Rosetta (45% on average), taking under a second to produce designs compared to the reported hours for Rosetta. We further demonstrate the utility of gRNAde on a new benchmark of multi-state design for structurally flexible RNAs, as well as zero-shot ranking of mutational fitness landscapes in a retrospective analysis of a recent ribozyme. Experimental wet lab validation on 10 different structured RNA backbones finds that gRNAde has a success rate of 50% at designing pseudoknotted RNA structures, a significant advance over 35% for Rosetta. Open source code and tutorials are available at: github.com/chaitjo/geometric-rna-design.

    View details for DOI 10.1101/2024.03.31.587283

    View details for PubMedID 38826198

  • gRNAde: A Geometric Deep Learning Pipeline for 3D RNA Inverse Design. Methods in molecular biology (Clifton, N.J.) Joshi, C. K., Lio, P. 2025; 2847: 121-135

    Abstract

    Fundamental to the diverse biological functions of RNA are its 3D structure and conformational flexibility, which enable single sequences to adopt a variety of distinct 3D states. Currently, computational RNA design tasks are often posed as inverse problems, where sequences are designed based on adopting a single desired secondary structure without considering 3D geometry and conformational diversity. In this tutorial, we present gRNAde, a geometric RNA design pipeline operating on sets of 3D RNA backbone structures to design sequences that explicitly account for RNA 3D structure and dynamics. gRNAde is a graph neural network that uses an SE (3) equivariant encoder-decoder framework for generating RNA sequences conditioned on backbone structures where the identities of the bases are unknown. We demonstrate the utility of gRNAde for fixed-backbone re-design of existing RNA structures of interest from the PDB, including riboswitches, aptamers, and ribozymes. gRNAde is more accurate in terms of native sequence recovery while being significantly faster compared to existing physics-based tools for 3D RNA inverse design, such as Rosetta.

    View details for DOI 10.1007/978-1-0716-4079-1_8

    View details for PubMedID 39312140

  • All-atom Diffusion Transformers: Unified generative modelling of molecules and materials Joshi, C. K., Fu, X., Liao, Y., Gharakhanyan, V., Miller, B., Sriram, A., Ulissi, Z. W. edited by Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J. JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2025: 28393-28417
  • Evaluating Representation Learning on the Protein Structure Universe. ArXiv Jamasb, A. R., Morehead, A., Joshi, C. K., Zhang, Z., Didi, K., Mathis, S., Harris, C., Tang, J., Cheng, J., Lio, P., Blundell, T. L. 2024

    Abstract

    We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models. We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop.

    View details for PubMedID 38947934

  • Hypergraph factorization for multi-tissue gene expression imputation NATURE MACHINE INTELLIGENCE Vinas, R., Joshi, C. K., Georgiev, D., Lin, P., Dumitrascu, B., Gamazon, E. R., Lio, P. 2023: 739-753

    Abstract

    Integrating gene expression across tissues and cell types is crucial for understanding the coordinated biological mechanisms that drive disease and characterise homeostasis. However, traditional multitissue integration methods cannot handle uncollected tissues or rely on genotype information, which is often unavailable and subject to privacy concerns. Here we present HYFA (Hypergraph Factorisation), a parameter-efficient graph representation learning approach for joint imputation of multi-tissue and cell-type gene expression. HYFA is genotype-agnostic, supports a variable number of collected tissues per individual, and imposes strong inductive biases to leverage the shared regulatory architecture of tissues and genes. In performance comparison on Genotype-Tissue Expression project data, HYFA achieves superior performance over existing methods, especially when multiple reference tissues are available. The HYFA-imputed dataset can be used to identify replicable regulatory genetic variations (eQTLs), with substantial gains over the original incomplete dataset. HYFA can accelerate the effective and scalable integration of tissue and cell-type transcriptome biorepositories.

    View details for DOI 10.1038/s42256-023-00684-8

    View details for Web of Science ID 001033508400002

    View details for PubMedID 37771758

    View details for PubMedCentralID PMC10538467

  • On the Expressive Power of Geometric Graph Neural Networks Joshi, C. K., Bodnar, C., Mathis, S., Cohen, T., Lio, P. edited by Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2023
  • On Representation Knowledge Distillation for Graph Neural Networks IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS Joshi, C. K., Liu, F., Xun, X., Lin, J., Foo, C. 2024; 35 (4): 4656-4667

    Abstract

    Knowledge distillation (KD) is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the local structure preserving (LSP) loss, which matches local structural relationships defined over edges across the student and teacher's node embeddings. This article studies whether preserving the global topology of how the teacher embeds graph data can be a more effective distillation objective for GNNs, as real-world graphs often contain latent interactions and noisy edges. We propose graph contrastive representation distillation (G-CRD), which uses contrastive learning to implicitly preserve global topology by aligning the student node embeddings to those of the teacher in a shared representation space. Additionally, we introduce an expanded set of benchmarks on large-scale real-world datasets where the performance gap between teacher and student GNNs is non-negligible. Experiments across four datasets and 14 heterogeneous GNN architectures show that G-CRD consistently boosts the performance and robustness of lightweight GNNs, outperforming LSP (and a global structure preserving (GSP) variant of LSP) as well as baselines from 2-D computer vision. An analysis of the representational similarity among teacher and student embedding spaces reveals that G-CRD balances preserving local and global relationships, while structure preserving approaches are best at preserving one or the other.

    View details for DOI 10.1109/TNNLS.2022.3223018

    View details for Web of Science ID 000912827500001

    View details for PubMedID 36459610

  • Learning the travelling salesperson problem requires rethinking generalization CONSTRAINTS Joshi, C. K., Cappart, Q., Rousseau, L., Laurent, T. 2022; 27 (1-2): 70-98
  • Point Discriminative Learning for Data-efficient 3D Point Cloud Analysis Liu, F., Lin, G., Foo, C., Joshi, C. K., Lin, J., IEEE IEEE. 2022: 42-51
  • Benchmarking Graph Neural Networks JOURNAL OF MACHINE LEARNING RESEARCH Dwivedi, V., Joshi, C. K., Luu, A., Laurent, T., Bengio, Y., Bresson, X. 2022; 23