All Publications


  • Genomic scale analysis of assembly-line polyketide synthase diversity and evolution. Natural product reports Lee, S., Khosla, C. 2026

    Abstract

    Covering: up to 2026Assembly-line polyketide synthases (PKSs) are among the most sophisticated catalysts in nature, responsible for the biosynthesis of many medicinally important natural products including antibiotics, anticancer agents, immunosuppressants and veterinary agents. While conventional methods for polyketide natural product discovery have fallen out of favor, the rapid expansion of microbial genome sequencing has revealed a vast untapped diversity of assembly-line PKSs, most of which can be regarded as "orphans" in that their product identity is unknown. A clear picture of how this sequence diversity reflects the evolutionary history of assembly-line PKSs could enable judicious prioritization of experimental efforts aimed at decoding these orphans. To this end we have used a scalable curation workflow to hand-curate an updated catalogue, PKSClusterDB, that includes 16 633 non-redundant assembly-line PKSs. We have also formulated an automatable "anchor-window" framework based on conserved multimodular segments of assembly-line PKSs to identify and interrogate PKS families of interest. Application of this framework to three different PKS families extracted from PKSClusterDB revealed lineage-dependent patterns of diversification, ranging from broadly diversified families to more compact or sharply bounded lineages. PKSClusterDB has also proven useful in exploring other aspects of PKS diversity, including scaffold architecture, extender-unit specificity, reductive-state programming, stereochemical control and modular organization. In closing, we consider how advances in machine learning could be harnessed to accelerate our understanding of assembly-line PKS evolution, diversity and biosynthetic mechanisms.

    View details for DOI 10.1039/d6np00052e

    View details for PubMedID 42359791

  • Learning Correlations between Internal Coordinates to Improve 3D Cartesian Coordinates for Proteins JOURNAL OF CHEMICAL THEORY AND COMPUTATION Li, J., Zhang, O., Lee, S., Namini, A., Liu, Z., Teixeira, J. M. C., Forman-Kay, J. D., Head-Gordon, T. 2023; 19 (14): 4689-4700

    Abstract

    We consider a generic representation problem of internal coordinates (bond lengths, valence angles, and dihedral angles) and their transformation to 3-dimensional Cartesian coordinates of a biomolecule. We show that the internal-to-Cartesian process relies on correctly predicting chemically subtle correlations among the internal coordinates themselves, and learning these correlations increases the fidelity of the Cartesian representation. We developed a machine learning algorithm, Int2Cart, to predict bond lengths and bond angles from backbone torsion angles and residue types of a protein, which allows reconstruction of protein structures better than using fixed bond lengths and bond angles or a static library method that relies on backbone torsion angles and residue types in a local environment. The method is able to be used for structure validation, as we show that the agreement between Int2Cart-predicted bond geometries and those from an AlphaFold 2 model can be used to estimate model quality. Additionally, by using Int2Cart to reconstruct an IDP ensemble, we are able to decrease the clash rate during modeling. The Int2Cart algorithm has been implemented as a publicly accessible python package at https://github.com/THGLab/int2cart.

    View details for DOI 10.1021/acs.jctc.2c01270

    View details for Web of Science ID 000963364700001

    View details for PubMedID 36749957

    View details for PubMedCentralID PMC10404647