Seungyeon Choi

dblp:212/9731 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TheSelective: Dual Affinity-Guided Diffusion for Selective Molecular Generation
Hyoungjoon Park, Hwanhee Kim, Seungyeon Choi, Yoonju Kim, Sanghyun Park 0003
PAKDD (2)3
2025 TheProperty: Euclidean 3D Molecular Representations for Explainable ADMET Prediction
abstract
Molecular representation learning has become a fundamental component of modern AI-driven drug discovery. Although 2D graph-based models have advanced this field, they often fail to capture the 3D geometric structures essential for many molecular properties, including ADMET. To address this limitation, we propose TheProperty, a self-supervised framework that integrates 2D topological information with 3D geometry. We introduce a distance denoising pre-training objective, which trains the model to recover the original euclidean distance matrix from a noised 3D conformer. For the fine-tuning, we designed a StructureInformed Attention mechanism that effectively fuses pre-trained 2D atom embeddings with the learned 3D distance embeddings, where the distance embeddings are injected as a structural pair bias into the attention mechanism. TheProperty demonstrated significant performance improvements across 22 tasks in the TDC benchmark. Furthermore, various analytical experiments verified that our proposed model effectively learns 3D structural consistency and pharmacological properties. TheProperty provides a robust, 3D-aware, and interpretable molecular representation that effectively integrates 2D topology with 3D geometry for enhanced ADMET prediction.
Yoonju Kim, Seungyeon Choi, Hwanhee Kim, Hyoungjoon Park
BIBM2
2025 Fixing Truncation-Induced Mode Collapse in GFlowNets via Pruning Loss
abstract
Generative Flow Networks (GFlowNets) are designed to generate diverse, high-quality samples by sampling proportionally to rewards using flow conservation constraints. However, they suffer from mode collapse, the very problem they were designed to address. We identify that the root cause is forced terminal states arising from artificial trajectory truncation in vast state spaces. Unlike natural terminal states, forced terminals violate flow conservation boundary constraints, causing flow leakage that biases generation toward maximum-length trajectories and triggers mode collapse. We propose Pruning Loss, a novel training objective that enforces flow conservation at forced terminals by requiring both sink flow and total outflow to equal the reward. This dual constraint implicitly drives unnecessary action flows to zero while maintaining non-vanishing gradients for stable convergence. Our theoretical analysis demonstrates that Pruning Loss recovers proper flow conservation in truncated spaces while guaranteeing gradient persistence. Empirical evaluation on molecular generation tasks validates our theoretical predictions. On sparse-reward kinase protein targets, our method achieves substantial improvements over standard objectives. On dense-reward drug-likeness tasks, all methods perform comparably well, validating that flow leakage specifically limits performance in sparse reward landscapes where diverse exploration is critical. Our results establish that correcting boundary constraints at forced terminals is more fundamental than refining balance equations. This principle provides a new foundation for addressing mode collapse in GFlowNets.
Hwanhee Kim, Seungyeon Choi, Hyoungjoon Park, Yoonju Kim
BIBM2
2025 TheProtein: Evolutionarily Informed Graph-Surface Protein Representation
abstract
Protein representation learning is a research field that converts proteins into numerical representations that computational models can process, playing a crucial role in areas such as structure-based drug development and protein function prediction. However, existing methods struggle to integrate global information from protein language models with local structure and surface data, particularly on protein surfaces which are key to interactions. We propose a novel surface feature initialization method that maps ESM embeddings onto the protein surface and TheProtein, a novel multimodal architecture that integrates global evolutionary context with local structure and surface information. This is achieved through a new surface feature initialization method using ESM embeddings and dedicated TheProtein blocks. TheProtein achieved new state-of-the-art performance on the Atom3D benchmark, with ablation studies and visualizations confirming its effective integration of global ESM information with local data. Our work highlights the importance of evolutionary context in multimodal protein representation and offers a significant methodological advance.
Seungyeon Choi, Hwanhee Kim, Hyoungjoon Park, Yoonju Kim, Hyunseo Yang
BIBM2
2025 LatentTune: Efficient Tuning of High Dimensional Database Parameters via Latent Representation Learning
abstract
As data volumes continue to grow, optimizing database performance has become increasingly critical, making the implementation of effective tuning methods essential. Among various approaches, database parameter tuning has proven to be a highly effective means of enhancing performance. Recent studies have shown that machine learning techniques can successfully optimize database parameters, leading to significant performance improvements. However, existing methods still face several limitations. First, they require substantial time to generate large training datasets. Second, to cope with the challenges of highdimensional optimization, they typically optimize only a subset of parameters rather than the full configuration space. Third, they often rely on information from similar workloads instead of directly leveraging information from the target workload. To address these limitations, we propose LatentTune, a novel approach that differs fundamentally from traditional methods. To reduce the time required for data generation, LatentTune incorporates a data augmentation strategy. Furthermore, it constructs a latent space that compresses information from all database parameters, enabling the optimization of the full configuration space. In addition, LatentTune integrates external metric information into the latent space, allowing for precise tuning tailored to the actual target workload. Experimental results demonstrate that LatentTune outperforms baseline models across four workloads on MySQL and RocksDB, achieving up to 1332 % improvement for RocksDB and 11.82 % throughput gain with 46.01 % latency reduction for MySQL.
Sein Kwon, Youngwan Jo, Seungyeon Choi, Jieun Lee 0006, Huijun Jin, Sanghyun Park 0003
HiPC3
2025 Controllable 3D Molecular Generation for Structure-Based Drug Design Through Bayesian Flow Networks and Gradient Integration
abstract
Recent advances in Structure-based Drug Design (SBDD) have leveraged generative models for 3D molecular generation, predominantly evaluating model performance by binding affinity to target proteins. However, practical drug discovery necessitates high binding affinity along with synthetic feasibility and selectivity, critical properties that were largely neglected in previous evaluations. To address this gap, we identify fundamental limitations of conventional diffusion-based generative models in effectively guiding molecule generation toward these diverse pharmacological properties. We propose $\texttt{CByG}$, a novel framework extending Bayesian Flow Network into a gradient-based conditional generative model that robustly integrates property-specific guidance. Additionally, we introduce a comprehensive evaluation scheme incorporating practical benchmarks for binding affinity, synthetic feasibility, and selectivity, overcoming the limitations of conventional evaluation methods. Extensive experiments demonstrate that our proposed $\texttt{CByG}$, framework significantly outperforms baseline models across multiple essential evaluation criteria, highlighting its effectiveness and practicality for real-world drug discovery applications.
Seungyeon Choi, Hwanhee Kim, Chihyun Park, Dahyeon Lee, Yoonju Kim, Hyoungjoon Park, Sein Kwon, Youngwan Jo
NeurIPS1
2025 Exploring the potential of compound-protein complex structure-free models in virtual screening using BlendNet
abstract
Identifying new compounds that interact with a target is a crucial time-limiting step in the initial phases of drug discovery. Compound-protein complex structure-based affinity prediction models can expedite this process; however, their dependence on high-quality three-dimensional (3D) complex structures limits their practical application. Prediction models that do not require 3D complex structures for binding-affinity estimation offer a theoretically attractive alternative; however, accurately predicting affinity without interaction information presents significant challenges. We introduce BlendNet, a framework that employs a knowledge transfer strategy to improve affinity prediction accuracy by learning the interdependent relationships between compounds and proteins without relying on 3D complex structures. Compared with state-of-the-art models for affinity prediction, BlendNet demonstrated superior performance across various cold-start cases. The ability of BlendNet to interpret compound-protein interactions without utilizing complex structure data highlights its potential to accelerate and streamline drug development.
Hwanhee Kim, Jieun Lee 0006, Seungyeon Choi
Briefings Bioinform.4
2024 PretrainedBA: Enhancing Compound-Protein Binding Affinity Prediction Accuracy via Pre-training Large-Scale Interaction Information
abstract
Finding potential drug candidates with high binding affinity for the specific target protein presents an important goal in early drug discovery. Although compound-protein complex structure-based affinity prediction methods have shown promising prediction accuracy, their dependency on high-resolution three-dimensional (3D) complex structure data considerably limits their practical application. Alternatively, many complex-free binding affinity prediction methods have been proposed; however, there is still room for improvement to compensate effectively for the lack of binding information. In particular, the interpretability of compound-protein interactions is a significant challenge that needs to be addressed. To alleviate the limitations of current complex-free models, we propose PretrainedBA, a predictive model that uses pre-training strategies on large-scale datasets, including interaction data. PretrainedBA pre-trains the interdependent relationships between compounds and proteins, rather than the independent pre-training of compounds and proteins utilized in existing studies. PretrainedBA consists of six modules and is designed to effectively model compound-protein interactions within identified binding pockets. Comparisons with state-of-the-art complex-free models on seven external benchmark datasets demonstrate that this pre-training strategy improves binding affinity prediction accuracy. In particular, the outstanding interpretive power of compound-protein interaction mechanisms compared with the previous method further emphasizes the value of PretrainedBA. Real-world application evaluation using the Database of Useful Decoys-Enhanced (DUD-E) dataset confirmed PretrainedBA’s practical applicability, demonstrating its utility in drug discovery.
Sangmin Seo 0001, Seungyeon Choi, Hwanhee Kim, Sanghyun Park 0003
BIBM2
2024 SPIN: SE(3)-Invariant Physics Informed Network for Binding Affinity Prediction
abstract
Accurate prediction of protein-ligand binding affinity is crucial for rapid and efficient drug development. Recently, the importance of predicting binding affinity has led to increased attention on research that models the three-dimensional structure of protein-ligand complexes using graph neural networks to predict binding affinity. However, traditional methods often fail to accurately model the complex’s spatial information or rely solely on geometric features, neglecting the principles of protein-ligand binding. This can lead to overfitting, resulting in models that perform poorly on independent datasets and ultimately reducing their usefulness in real drug development. To address this issue, we propose SPIN, a model designed to achieve superior generalization by incorporating various inductive biases applicable to this task, beyond merely training on empirical data from datasets. For prediction, we defined two types of inductive biases: a geometric perspective that maintains consistent binding affinity predictions regardless of the complex’s rotations and translations, and a physicochemical perspective that necessitates minimal binding free energy along their reaction coordinate for effective protein-ligand binding. These prior knowledge inputs enable the SPIN to outperform comparative models in benchmark sets such as CASF-2016 and CSAR HiQ. Furthermore, we demonstrated the practicality of our model through virtual screening experiments and validated the reliability and potential of our proposed model based on experiments assessing its interpretability.
Seungyeon Choi, Sangmin Seo 0001, Sanghyun Park 0003
ECAI1
2024 DrDiff: Drug Response Prediction Through Controllable Diffusion-GE and Graph Attention Network
abstract
The accurate prediction of drug responses based on the genomic profile of a patient is essential to progress in the field of precision medicine. The advent of various deep-learning algorithms based on publicly available large-scale omics datasets is the driving force behind research in this field. The characteristics of biological datasets, characterized by high dimensions and low sample sizes, pose challenges of overfitting and limited generalization in prediction models. Additionally, constructing prediction models using biological data such as gene expression is further complicated by the need to account for the complex relationships among genes, which exacerbates the aforementioned challenges. To address these challenges, we propose a drug response prediction framework (DrDiff) that integrates a denoising diffusion probabilistic model (DDPM) based data augmentation module with a graph attention network based drug response prediction module. The proposed model showed a 10% higher AUC than the state-of-the-art models for drug response prediction for the six drugs considered in the study, suggesting the superior generalization performance of DrDiff over other baseline models. Furthermore, we demonstrated the feasibility of generative models, which form one of the modules of the proposed framework, in overcoming the fundamental limitations of omics datasets. Further experiments bear out the feasibility of generative models, which form one of the modules of the proposed framework, in augmenting gene expression data.
Seungyeon Choi, Sangmin Seo 0001, Jonghwan Choi, Chihyun Park, Sanghyun Park 0003
ECAI1
2024 Pseq2Sites: Enhancing protein sequence-based ligand binding-site prediction accuracy via the deep convolutional network and attention mechanism
abstract
Protein-ligand interactions play an essential role in many biological processes, and prior knowledge of ligand binding sites is necessary for successful drug design. Many 3D structure- and sequence-based methods have been proposed for identifying ligand binding sites. The 3D structure-based methods typically achieve better binding site prediction than the sequence-based methods. However, as deep-learning techniques that can extract structural information from large-scale sequence data have been developed, the performance gap between 3D structure- and sequence-based methods is narrowing. Nonetheless, there remains room for improvement in sequence-based prediction. We propose Pseq2Sites, a sequence-based deep-learning model for predicting ligand binding sites. Pseq2Sites comprises a 1D convolutional neural network that extracts local features from the protein sequence, and a position-based attention mechanism that captures long-distance dependencies between binding residues. To verify the effectiveness of the proposed method, we compared it with other state-of-the-art methods using three public datasets: COACH420, HOLO4K, and CSAR-NRC HiQ. Utilizing solely protein sequence information, Pseq2Sites outperformed 3D structure-based state-of-the-art methods on external test datasets; within the COACH420 dataset, Pseq2Sites remarkably identified 97% of the binding pockets (at a significance level δ = 0.5), which was 27% higher than the second highest-performing model. Pseq2Sites also achieved outstanding binding site prediction, even for proteins with low similarity to the training dataset. Our code is available at https://github.com/Blue1993/Pseq2Sites.
Sangmin Seo 0001, Jonghwan Choi, Seungyeon Choi, Jieun Lee 0006, Chihyun Park, Sanghyun Park 0003
Eng. Appl. Artif. Intell.3