VLDB 2026 Research / reviewers in the wild / expert
Dongmin Bang
dblp:358/8445
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-9217-8380ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 89% Computational science and engineering · 11% | |
| Artificial intelligence
3 papers |
Generative modeling · 90% Learning paradigms · 10% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
drug discovery |
1.7 | 2 | 2025 | ADME-drug-likeness: enriching molecular foundation models via pharmacokinetics-guided multi-task learning for drug-likeness prediction · Bioinform. 2025 BounDr.E: Predicting Drug-likeness via Biomedical Knowledge Alignment and EM-like One-Class Boundary Optimization · ICML 2025 |
Bioinformatics and computational biology › drug discovery
drug-likeness prediction |
1.7 | 2 | 2025 | ADME-drug-likeness: enriching molecular foundation models via pharmacokinetics-guided multi-task learning for drug-likeness prediction · Bioinform. 2025 BounDr.E: Predicting Drug-likeness via Biomedical Knowledge Alignment and EM-like One-Class Boundary Optimization · ICML 2025 |
Bioinformatics and computational biology › drug discovery › drug-target interaction prediction
drug-target binding affinity prediction |
0.9 | 1 | 2025 | MixingDTA: improved drug-target affinity prediction by extending mixup with guilt-by-association · Bioinform. 2025 |
Bioinformatics and computational biology › drug discovery
virtual screening |
0.9 | 1 | 2025 | BounDr.E: Predicting Drug-likeness via Biomedical Knowledge Alignment and EM-like One-Class Boundary Optimization · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization · AAAI 2024 |
Machine learning › Generative modeling › diffusion model
molecular conformation generation |
0.8 | 1 | 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization · AAAI 2024 |
Machine learning › Generative modeling › diffusion model
schrödinger bridge |
0.8 | 1 | 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization · AAAI 2024 |
Computational science and engineering
computational chemistry |
0.8 | 1 | 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization · AAAI 2024 |
Bioinformatics and computational biology › drug discovery
drug response prediction |
0.8 | 1 | 2024 | Transfer learning of condition-specific perturbation in gene interactions improves drug response prediction · Bioinform. 2024 |
Machine learning › Learning paradigms
multi-task learning |
0.3 | 1 | 2025 | ADME-drug-likeness: enriching molecular foundation models via pharmacokinetics-guided multi-task learning for drug-likeness prediction · Bioinform. 2025 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.3 | 1 | 2025 | MixingDTA: improved drug-target affinity prediction by extending mixup with guilt-by-association · Bioinform. 2025 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 1.7molecular foundation models · 1.7schrödinger bridge · 1.5transformer · 0.9pre-trained language model · 0.9one-class boundary optimization · 0.9mixup · 0.9knowledge alignment · 0.9guilt-by-association · 0.9expectation-maximization · 0.9parameter-frozen fine-tuning · 0.8gene-gene attention · 0.8SE(3)-equivariant diffusion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BounDr.E: Predicting Drug-likeness via Biomedical Knowledge Alignment and EM-like One-Class Boundary OptimizationabstractThe advent of generative AI now enables large-scale $\textit{de novo}$ design of molecules, but identifying viable drug candidates among them remains an open problem. Existing drug-likeness prediction methods often rely on ambiguous negative sets or purely structural features, limiting their ability to accurately classify drugs from non-drugs. In this work, we introduce BounDr.E: a novel modeling of drug-likeness as a compact space surrounding approved drugs through a dynamic one-class boundary approach. Specifically, we enrich the chemical space through biomedical knowledge alignment, and then iteratively tighten the drug-like boundary by pushing non-drug-like compounds outside via an Expectation-Maximization (EM)-like process. Empirically, BounDr.E achieves 10% F1-score improvement over the previous state-of-the-art and demonstrates robust cross-dataset performance, including zero-shot toxic compound filtering. Additionally, we showcase its effectiveness through comprehensive case studies in large-scale $\textit{in silico}$ screening. Our codes and constructed benchmark data under various schemes are provided at: https://github.com/eugenebang/boundr_e. Dongmin Bang, Inyoung Sung, Yinhua Piao, Sangseon Lee, Sun Kim |
ICML | 1 |
| 2025 | ADME-drug-likeness: enriching molecular foundation models via pharmacokinetics-guided multi-task learning for drug-likeness predictionabstractSUMMARY: Recent breakthroughs in AI-driven generative models enable the rapid design of extensive molecular libraries, creating an urgent need for fast and accurate drug-likeness evaluation. Traditional approaches, however, rely heavily on structural descriptors and overlook pharmacokinetic (PK) factors such as absorption, distribution, metabolism, and excretion (ADME). Furthermore, existing deep-learning models neglect the complex interdependencies among ADME tasks, which play a pivotal role in determining clinical viability. We introduce ADME-DL (drug likeness), a novel two-step pipeline that first enhances diverse range of molecular foundation models (MFMs) via sequential ADME multi-task learning. By enforcing an A→D→M→E flow-grounded in a data-driven task dependency analysis that aligns with established PK principles-our method more accurately encodes PK information into the learned embedding space. In Step 2, the resulting ADME-informed embeddings are leveraged for drug-likeness classification, distinguishing approved drugs from negative sets drawn from chemical libraries. Through comprehensive experiments, our sequential ADME multi-task learning achieves up to +2.4% improvement over state-of-the-art baselines, and enhancing performance across tested MFMs by up to +18.2%. Case studies with clinically annotated drugs validate that respecting the PK hierarchy produces more relevant predictions, reflecting drug discovery phases. These findings underscore the potential of ADME-DL to significantly enhance the early-stage filtering of candidate molecules, bridging the gap between purely structural screening methods and PK-aware modeling. AVAILABILITY AND IMPLEMENTATION: The source code for ADME-DL is available at https://github.com/eugenebang/ADME-DL. Dongmin Bang, Haerin Song, Sun Kim |
Bioinform. | 1 |
| 2025 | MixingDTA: improved drug-target affinity prediction by extending mixup with guilt-by-associationabstractSUMMARY: Drug-target affinity (DTA) prediction is an important regression task for drug discovery, which can provide richer information than traditional drug-target interaction prediction as a binary prediction task. To achieve accurate DTA prediction, quite large amount of data are required for each drug, which is not available as of now. Thus, data scarcity and sparsity is a major challenge. Another important task is "cold-start" DTA prediction for unseen drug or protein. In this work, we introduce MixingDTA, a novel framework to tackle data scarcity by incorporating domain-specific pretrained language models for molecules and proteins with our MEETA (MolFormer and ESM-based Efficient aggregation Transformer for Affinity) model. We further address the label sparsity and cold-start challenges through a novel data augmentation strategy named GBA-Mixup, which interpolates embeddings of neighboring entities based on the guilt-by-association (GBA) principle, to improve prediction accuracy even in sparse regions of DTA space. Our experiments on benchmark datasets demonstrate that the MEETA backbone alone provides up to a 19% improvement of mean squared error over current state-of-the-art baseline, and the addition of GBA-Mixup contributes a further 8.4% improvement. Importantly, GBA-Mixup is model-agnostic, delivering performance gains across all tested backbone models of up to 16.9%. Case studies shows how MixingDTA interpolates between drugs and targets in the embedding space, demonstrating generalizability for unseen drug-target pairs while effectively focusing on functionally critical residues. These results highlight MixingDTA's potential to accelerate drug discovery by offering accurate, scalable, and biologically informed DTA predictions. AVAILABILITY AND IMPLEMENTATION: The code for MixingDTA is available at https://github.com/rokieplayer20/MixingDTA. Youngoh Kim, Dongmin Bang, Bonil Koo, Jungseob Yi, Changyun Cho, Jeonguk Choi, Sun Kim |
Bioinform. | 2 |
| 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer OptimizationabstractThe generation of energetically optimal 3D molecular conformers is crucial in cheminformatics and drug discovery. While deep generative models have been utilized for direct generation in Euclidean space, this approach encounters challenges, including the complexity of navigating a vast search space. Recent generative models that implement simplifications to circumvent these challenges have achieved state-of-the-art results, but this simplified approach unavoidably creates a gap between the generated conformers and the ground-truth conformational landscape. To bridge this gap, we introduce DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization, a novel diffusion framework that enables direct learning of nonlinear diffusion processes in prior-constrained Euclidean space for the optimization of 3D molecular conformers. Through the incorporation of an SE(3)-equivariant Schrödinger bridge, we establish the roto-translational equivariance of the generated conformers. Our framework is model-agnostic and offers an easily implementable solution for the post hoc optimization of conformers produced by any generation method. Through comprehensive evaluations and analyses, we establish the strengths of our framework, substantiating the application of the Schrödinger bridge for molecular conformer optimization. First, our approach consistently outperforms four baseline approaches, producing conformers with higher diversity and improved quality. Then, we show that the intermediate conformers generated during our diffusion process exhibit valid and chemically meaningful characteristics. We also demonstrate the robustness of our method when starting from conformers of diverse quality, including those unseen during training. Lastly, we show that the precise generation of low-energy conformers via our framework helps in enhancing the downstream prediction of molecular properties. The code is available at https://github.com/Danyeong-Lee/DiSCO. Danyeong Lee, Dohoon Lee, Dongmin Bang, Sun Kim |
AAAI | 3 |
| 2024 | Transfer learning of condition-specific perturbation in gene interactions improves drug response predictionabstractSUMMARY: Drug response is conventionally measured at the cell level, often quantified by metrics like IC50. However, to gain a deeper understanding of drug response, cellular outcomes need to be understood in terms of pathway perturbation. This perspective leads us to recognize a challenge posed by the gap between two widely used large-scale databases, LINCS L1000 and GDSC, measuring drug response at different levels-L1000 captures information at the gene expression level, while GDSC operates at the cell line level. Our study aims to bridge this gap by integrating the two databases through transfer learning, focusing on condition-specific perturbations in gene interactions from L1000 to interpret drug response integrating both gene and cell levels in GDSC. This transfer learning strategy involves pretraining on the transcriptomic-level L1000 dataset, with parameter-frozen fine-tuning to cell line-level drug response. Our novel condition-specific gene-gene attention (CSG2A) mechanism dynamically learns gene interactions specific to input conditions, guided by both data and biological network priors. The CSG2A network, equipped with transfer learning strategy, achieves state-of-the-art performance in cell line-level drug response prediction. In two case studies, well-known mechanisms of drugs are well represented in both the learned gene-gene attention and the predicted transcriptomic profiles. This alignment supports the modeling power in terms of interpretability and biological relevance. Furthermore, our model's unique capacity to capture drug response in terms of both pathway perturbation and cell viability extends predictions to the patient level using TCGA data, demonstrating its expressive power obtained from both gene and cell levels. AVAILABILITY AND IMPLEMENTATION: The source code for the CSG2A network is available at https://github.com/eugenebang/CSG2A. Dongmin Bang, Bonil Koo, Sun Kim |
Bioinform. | 1 |
| 2023 | A model-agnostic framework to enhance knowledge graph-based drug combination prediction with drug-drug interaction data and supervised contrastive learningabstractCombination therapies have brought significant advancements to the treatment of various diseases in the medical field. However, searching for effective drug combinations remains a major challenge due to the vast number of possible combinations. Biomedical knowledge graph (KG)-based methods have shown potential in predicting effective combinations for wide spectrum of diseases, but the lack of credible negative samples has limited the prediction performance of machine learning models. To address this issue, we propose a novel model-agnostic framework that leverages existing drug-drug interaction (DDI) data as a reliable negative dataset and employs supervised contrastive learning (SCL) to transform drug embedding vectors to be more suitable for drug combination prediction. We conducted extensive experiments using various network embedding algorithms, including random walk and graph neural networks, on a biomedical KG. Our framework significantly improved performance metrics compared to the baseline framework. We also provide embedding space visualizations and case studies that demonstrate the effectiveness of our approach. This work highlights the potential of using DDI data and SCL in finding tighter decision boundaries for predicting effective drug combinations. Jeonghyeon Gu, Dongmin Bang, Jungseob Yi, Sangseon Lee, Dong Kyu Kim, Sun Kim |
Briefings Bioinform. | 2 |