Jejoong Yoo

dblp:370/5504 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-7120-8464ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein structure prediction
1.522025
DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models · Bioinform. 2025
DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy function · Bioinform. 2023
Bioinformatics and computational biology › protein structure analysis › protein flexibility
conformational diversity
1.012026
StrucPTM: a database of structurally validated protein modifications and their conformational variation · Bioinform. 2026
Bioinformatics and computational biology › proteomics
post-translational modification
1.012026
StrucPTM: a database of structurally validated protein modifications and their conformational variation · Bioinform. 2026
Bioinformatics and computational biology › sequence alignment
protein sequence alignment
1.012026
OTalign: optimal transport alignment for remote protein homologs using protein language model embeddings · Bioinform. 2026
Bioinformatics and computational biology
protein structure analysis
1.012026
StrucPTM: a database of structurally validated protein modifications and their conformational variation · Bioinform. 2026
Bioinformatics and computational biology › structural bioinformatics
protein structure database
1.012026
StrucPTM: a database of structurally validated protein modifications and their conformational variation · Bioinform. 2026
Bioinformatics and computational biology › multiple sequence alignment
multiple sequence alignment construction
0.912025
DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models · Bioinform. 2025
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.912025
DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models · Bioinform. 2025
Bioinformatics and computational biology › protein structure prediction
protein complex structure prediction
0.312025
DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models · Bioinform. 2025
Bioinformatics and computational biology › protein structure prediction
protein structure refinement
0.212023
DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy function · Bioinform. 2023
Bioinformatics and computational biology › protein structure prediction
side-chain prediction
0.212023
DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy function · Bioinform. 2023

Methods — techniques the papers use, named apart from their topics

protein language model · 1.9structural descriptor analysis · 1.0optimal transport · 1.0homolog comparison · 1.0entropy-regularized unbalanced optimal transport · 1.0vector embedding database · 0.9contrastive learning · 0.9deep neural network · 0.7conformational space annealing · 0.7conditional random field · 0.7
YearPublicationVenuePosition
2026 StrucPTM: a database of structurally validated protein modifications and their conformational variation
abstract
MOTIVATION: Post-translational modifications (PTMs) alter functional states and interaction specificity largely through the conformational changes they impose on protein structure. However, most existing resources remain sequence-centric and cannot reveal how chemical modifications reshape 3D structures. To address this gap, we propose a structural database that systematically extracts and contextualizes modification sites within experimentally determined protein structures, providing a foundation for future studies of protein structure, function, and regulatory mechanisms. RESULTS: We present StrucPTM, a database that extracts modified residues directly from the Protein Data Bank (PDB) structures using atom-level composition rules, substantially expanding coverage beyond annotation-dependent methods. Each validated PTM modification is mapped onto a UniProt entry. The database further characterizes residues using key structural descriptors-including secondary structure, relative solvent accessibility (RSA), and whether the PTM site lies at an interchain interface. All chains associated with the same UniProt ID are compared and grouped into homolog sets based on sequence identity. This emphasizes structural conservation among homologs, allowing PTM-induced conformational deviations to be distinguished from unrelated sequence divergence. AVAILABILITY AND IMPLEMENTATION: StrucPTM offers searchable access, interactive 3D visualization, and homolog-based structural comparison through its web interface: https://prix.hanyang.ac.kr/strucptm. The source code and datasets are permanently archived on Zenodo (DOI: 10.5281/zenodo.18939125) and are accessible via GitHub (https://github.com/HanyangBISLab/StrucPTM.git).
Seonggwang Jeon, Jejoong Yoo, Keehyoung Joo, Eunok Paek
Bioinform.2
2026 OTalign: optimal transport alignment for remote protein homologs using protein language model embeddings
abstract
MOTIVATION: Protein sequence alignment is a crucial task in bioinformatics, yet aligning remote homologs with low sequence identity remains a longstanding challenge, particularly due to the difficulty of handling gaps. We introduce a new method that applies Optimal Transport (OT) theory to sequence alignment, providing a mathematically principled framework for modeling residue matches and gaps. RESULTS: OTalign formulates sequence alignment as an entropy-regularized unbalanced optimal transport (UOT) problem over embeddings derived from protein language models (PLMs). Unlike traditional methods, it introduces position-specific gap penalties that adapt to each sequence pair. On challenging remote-homolog benchmarks (SABmark, MALIDUP, MALISAM), OTalign consistently outperforms baselines (Needleman-Wunsch, HHalign) and recent PLM-based methods (PLMAlign, DeepBLAST), achieving F1 scores of 0.594 on SABmark Superfamily and 0.358 on SABmark Twilight. Furthermore, OTalign provides a quantitative and interpretable metric of how effectively PLM embeddings represent sequence similarity relationships. Finally, its differentiable nature enables end-to-end fine-tuning of PLMs, establishing a framework for learning embeddings explicitly optimized for alignment tasks. AVAILABILITY AND IMPLEMENTATION: This code is available at https://github.com/DeepFoldProtein/OTalign.
Hanjin Bae, Gyeongpil Jo, Kunwoo Kim, Jejoong Yoo, Keehyoung Joo
Bioinform.5
2025 DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models
abstract
MOTIVATION: Protein structure prediction has been revolutionized and generalized with the advent of cutting-edge AI methods such as AlphaFold, but reliance on computationally intensive multiple sequence alignments (MSA) remains a major limitation. RESULTS: We introduce DeepFold-PLM, a novel framework that integrates advanced protein language models with vector embedding databases to enhance ultra-fast MSA construction, remote homology detection, and protein structure prediction. DeepFold-PLM utilizes high-dimensional embeddings and contrastive learning, significantly accelerate MSA generation, achieving 47 times faster than standard methods, while maintaining prediction accuracy comparable to AlphaFold. In addition, it enhances structure prediction by extending modeling capabilities to multimeric protein complexes, provides a scalable PyTorch-based implementation for efficient large-scale prediction. Our method also effectively increases sequence diversity (Neff = 8.65 versus 4.83 with JackHMMER) enriching coevolutionary information critical for accurate structure prediction. DeepFold-PLM thus represents a versatile and practical resource that enables high-throughput applications in computational structural biology. AVAILABILITY AND IMPLEMENTATION: Source codes and user-friendly Python API of all modules of DeepFold-PLM publicly available at https://github.com/DeepFoldProtein/DeepFold-PLM.
Hanjin Bae, Gyeongpil Jo, Kunwoo Kim, Sung Jong Lee, Jejoong Yoo, Keehyoung Joo
Bioinform.6
2023 DeepFold: enhancing protein structure prediction through optimized loss functions, improved template features, and re-optimized energy function
abstract
MOTIVATION: Predicting protein structures with high accuracy is a critical challenge for the broad community of life sciences and industry. Despite progress made by deep neural networks like AlphaFold2, there is a need for further improvements in the quality of detailed structures, such as side-chains, along with protein backbone structures. RESULTS: Building upon the successes of AlphaFold2, the modifications we made include changing the losses of side-chain torsion angles and frame aligned point error, adding loss functions for side chain confidence and secondary structure prediction, and replacing template feature generation with a new alignment method based on conditional random fields. We also performed re-optimization by conformational space annealing using a molecular mechanics energy function which integrates the potential energies obtained from distogram and side-chain prediction. In the CASP15 blind test for single protein and domain modeling (109 domains), DeepFold ranked fourth among 132 groups with improvements in the details of the structure in terms of backbone, side-chain, and Molprobity. In terms of protein backbone accuracy, DeepFold achieved a median GDT-TS score of 88.64 compared with 85.88 of AlphaFold2. For TBM-easy/hard targets, DeepFold ranked at the top based on Z-scores for GDT-TS. This shows its practical value to the structural biology community, which demands highly accurate structures. In addition, a thorough analysis of 55 domains from 39 targets with publicly available structures indicates that DeepFold shows superior side-chain accuracy and Molprobity scores among the top-performing groups. AVAILABILITY AND IMPLEMENTATION: DeepFold tools are open-source software available at https://github.com/newtonjoo/deepfold.
Jae-Won Lee, Jong-Hyun Won, Seonggwang Jeon, Yujin Choo, Yubin Yeon, Jin-Seon Oh, Seonhwa Kim, InSuk Joung, Cheongjae Jang, Sung Jong Lee, Kyong Hwan Jin, Giltae Song, Eun-Sol Kim, Jejoong Yoo, Eunok Paek, Yung-Kyun Noh, Keehyoung Joo
Bioinform.16