Seokyoung Lee
Ph.D. Student in Chemistry, admitted Autumn 2023
All Publications
-
Genomic scale analysis of assembly-line polyketide synthase diversity and evolution.
Natural product reports
2026
Abstract
Covering: up to 2026Assembly-line polyketide synthases (PKSs) are among the most sophisticated catalysts in nature, responsible for the biosynthesis of many medicinally important natural products including antibiotics, anticancer agents, immunosuppressants and veterinary agents. While conventional methods for polyketide natural product discovery have fallen out of favor, the rapid expansion of microbial genome sequencing has revealed a vast untapped diversity of assembly-line PKSs, most of which can be regarded as "orphans" in that their product identity is unknown. A clear picture of how this sequence diversity reflects the evolutionary history of assembly-line PKSs could enable judicious prioritization of experimental efforts aimed at decoding these orphans. To this end we have used a scalable curation workflow to hand-curate an updated catalogue, PKSClusterDB, that includes 16 633 non-redundant assembly-line PKSs. We have also formulated an automatable "anchor-window" framework based on conserved multimodular segments of assembly-line PKSs to identify and interrogate PKS families of interest. Application of this framework to three different PKS families extracted from PKSClusterDB revealed lineage-dependent patterns of diversification, ranging from broadly diversified families to more compact or sharply bounded lineages. PKSClusterDB has also proven useful in exploring other aspects of PKS diversity, including scaffold architecture, extender-unit specificity, reductive-state programming, stereochemical control and modular organization. In closing, we consider how advances in machine learning could be harnessed to accelerate our understanding of assembly-line PKS evolution, diversity and biosynthetic mechanisms.
View details for DOI 10.1039/d6np00052e
View details for PubMedID 42359791
-
Learning Correlations between Internal Coordinates to Improve 3D Cartesian Coordinates for Proteins
JOURNAL OF CHEMICAL THEORY AND COMPUTATION
2023; 19 (14): 4689-4700
Abstract
We consider a generic representation problem of internal coordinates (bond lengths, valence angles, and dihedral angles) and their transformation to 3-dimensional Cartesian coordinates of a biomolecule. We show that the internal-to-Cartesian process relies on correctly predicting chemically subtle correlations among the internal coordinates themselves, and learning these correlations increases the fidelity of the Cartesian representation. We developed a machine learning algorithm, Int2Cart, to predict bond lengths and bond angles from backbone torsion angles and residue types of a protein, which allows reconstruction of protein structures better than using fixed bond lengths and bond angles or a static library method that relies on backbone torsion angles and residue types in a local environment. The method is able to be used for structure validation, as we show that the agreement between Int2Cart-predicted bond geometries and those from an AlphaFold 2 model can be used to estimate model quality. Additionally, by using Int2Cart to reconstruct an IDP ensemble, we are able to decrease the clash rate during modeling. The Int2Cart algorithm has been implemented as a publicly accessible python package at https://github.com/THGLab/int2cart.
View details for DOI 10.1021/acs.jctc.2c01270
View details for Web of Science ID 000963364700001
View details for PubMedID 36749957
View details for PubMedCentralID PMC10404647
https://orcid.org/0000-0002-8525-8646