Alejandro Buendia
Ph.D. Student in Biomedical Data Science, admitted Autumn 2024
All Publications
-
BenchRep-T: A Systematic Evaluation of T-Cell Repertoire-Based Disease Diagnostics.
bioRxiv : the preprint server for biology
2026
Abstract
Adaptive immune receptor repertoire sequencing data has emerged as a promising potential modality for disease diagnosis, relying on computational methods to analyze T-cell receptor (TCR) sequences from an individual's blood sample. Published methods rely on different cohorts, data preprocessing pipelines, and evaluation metrics, making direct comparison across methods challenging. We present BenchRep-T, a unified benchmark that standardizes multiple publicly available TCR repertoire datasets and evaluates nine computational approaches, spanning statistical enrichment of shared sequences, feature-engineered ensembles, deep learning, and sequence clustering. BenchRep-T evaluates methods on four tasks: disease classification across conditions, performance scaling under restricted sequence-sampling depth, recovery of known antigen-specific driver sequences, and evaluation of sensitivity to demographic confounding. Under controlled evaluation, simple baselines prove competitive, with tree-based models trained on V- and J-gene usage and short sequence motifs approaching the classification performance of more complex methods. Our findings underscore the complexity of modeling TCR repertoire data, and show that no single method dominates across all tasks. BenchRep-T provides a framework for rigorous and reproducible evaluation of TCR repertoire classification methods to accelerate the development of immune repertoire-based diagnostics.
View details for DOI 10.64898/2026.06.09.727013
View details for PubMedID 42327158
View details for PubMedCentralID PMC13277846
-
Bio-BLIP: A Multimodal Architecture for Transferable Reasoning in Genomic Variant Interpretation.
bioRxiv : the preprint server for biology
2026
Abstract
Developing scientific hypotheses in biology requires integrating heterogeneous evidence across DNA sequence, gene context, protein function, and prior literature. Existing multimodal AI systems expose biological evidence to reasoning models through textification or by projecting biological embeddings into fine-tuned language models. However, these models are typically highly optimized the specific set of tasks for which they are fine-tuned. Here we present Bio-BLIP, a multimodal Q-former based architecture which leverages biological embeddings and a LLM to generalize to complex reasoning tasks without task-specific fine-tuning. The key to Bio-BLIP is a new neural network architecture that integrates four data modalities - DNA, genes, proteins, and text - through a master Qformer model, which integrates the modality-specific information into a fixed-length prefix for the LLM backbone. Bio-BLIP is pretrained on the task of human genetic variant annotation and achieves a 29.8% increase in generating accurate variant features over frontier LLMs. We evaluate Bio-BLIP zero-shot on downstream genomic tasks of variant prioritization and target gene prediction. Bio-BLIP outperforms two alignment-free genomic language models on regulatory variant prioritization for Mendelian disease. Across the target gene prediction task, Bio-BLIP improves accuracy over LLMs by leveraging learned genomic variant knowledge in difficult cases. Our model produces rich, transparent reasoning traces. In biological domains characterized by multiple scales of data and varied downstream tasks, Bio-BLIP offers a step toward natively multimodal, generalizable reasoning.
View details for DOI 10.64898/2026.05.12.724740
View details for PubMedID 42182443
View details for PubMedCentralID PMC13192633
-
SpatialProp: tissue perturbation modeling with spatially resolved single-cell transcriptomics.
bioRxiv : the preprint server for biology
2025
Abstract
Perturbational studies are the gold standard for identifying causal relationships between components of biological systems. Recent technological advances, including Perturb-seq and related assays, have enabled high-throughput screening of genetic perturbation effects on single cells. Several machine learning tools have also been developed to infer the effect of single-cell perturbations. However, both approaches are generally limited to dissociated cells, and the effect of genetic perturbations on neighboring cells within intact tissue has not yet been explored. Here we introduce a computational framework using graph neural networks for predicting the effect of multi-gene, multi-cell type perturbations on cells in whole tissue sections. We leverage the natural heterogeneity in tissue microenvironments across spatially resolved single-cell transcriptomics datasets to train SpatialProp (Spatial Propagation of Single-cell Perturbations). We show that SpatialProp can predict gene expression from the tissue microenvironment and map fine-grained steering of tissue microenvironments to new target states. To assess for causal enrichment in spatial perturbation predictions, we propose CausalInteractionBench, a bidirectional benchmarking approach using curated cell-cell interactions. Under this benchmark, we evaluate the causal utility of SpatialProp in predicting the spatial effects of different perturbations. SpatialProp provides a framework towards rapid hypothesis generation and in silico perturbation experiments, particularly in the study of spatially patterned tissue biology.
View details for DOI 10.64898/2025.11.30.691355
View details for PubMedID 41573962
View details for PubMedCentralID PMC12822716
-
A foundation model of transcription across human cell types.
Nature
2025; 637 (8047): 965-973
Abstract
Transcriptional regulation, which involves a complex interplay between regulatory sequences and proteins, directs all biological processes. Computational models of transcription lack generalizability to accurately extrapolate to unseen cell types and conditions. Here we introduce GET (general expression transformer), an interpretable foundation model designed to uncover regulatory grammars across 213 human fetal and adult cell types1,2. Relying exclusively on chromatin accessibility data and sequence information, GET achieves experimental-level accuracy in predicting gene expression even in previously unseen cell types3. GET also shows remarkable adaptability across new sequencing platforms and assays, enabling regulatory inference across a broad range of cell types and conditions, and uncovers universal and cell-type-specific transcription factor interaction networks. We evaluated its performance in prediction of regulatory activity, inference of regulatory elements and regulators, and identification of physical interactions between transcription factors and found that it outperforms current models4 in predicting lentivirus-based massively parallel reporter assay readout5,6. In fetal erythroblasts7, we identified distal (greater than 1 Mbp) regulatory regions that were missed by previous models, and, in B cells, we identified a lymphocyte-specific transcription factor-transcription factor interaction that explains the functional significance of a leukaemia risk predisposing germline mutation8-10. In sum, we provide a generalizable and accurate model for transcription together with catalogues of gene regulation and transcription factor interactions, all with cell type specificity.
View details for DOI 10.1038/s41586-024-08391-z
View details for PubMedID 39779852
View details for PubMedCentralID PMC11754112
-
Deeper evaluation of a single-cell foundation model
NATURE MACHINE INTELLIGENCE
2024; 6 (12)
View details for DOI 10.1038/s42256-024-00949-w
View details for Web of Science ID 001377165100001
-
TabLLM: Few-shot Classification of Tabular Data with Large Language Models
edited by Ruiz, F., Dy, J., VanDeMeent, J. W.
JMLR-JOURNAL MACHINE LEARNING RESEARCH. 2023
View details for Web of Science ID 001222727705031