EDBT 2026 Demo / reviewers in the wild / expert
Keonwoo Kim 0002
dblp:58/2926-2 · also Mogan Gim
· DBLP profile ↗
14ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-6458-7723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cradle-VAE: Enhancing Single-Cell Gene Perturbation Modeling with Counterfactual Reasoning-based Artifact DisentanglementabstractPredicting cellular responses to various perturbations is a critical focus in drug discovery and personalized therapeutics, with deep learning models playing a significant role in this endeavor. Single-cell datasets contain technical artifacts that may hinder the predictability of such models, which poses quality control issues highly regarded in this area. To address this, we propose Cradle-VAE, a causal generative framework tailored for single-cell gene perturbation modeling, enhanced with counterfactual reasoning-based artifact disentanglement. Throughout training, Cradle-VAE models the underlying latent distribution of technical artifacts and perturbation effects present in single-cell datasets. It employs counterfactual reasoning to effectively disentangle such artifacts by modulating the latent basal spaces and learns robust features for generating cellular response data with improved quality. Experimental results demonstrate that this approach improves not only treatment effect estimation performance but also generative quality as well. Seungheun Baek, Soyon Park, Yan Ting Chok, Jun-Hyun Lee, Jueon Park, Keonwoo Kim 0002, Jaewoo Kang |
AAAI | 6 |
| 2025 | DeepAries: Adaptive Rebalancing Interval Selection for Enhanced Portfolio Selection
Jinkyu Kim 0004, Hyungjung Yi, Keonwoo Kim 0002, Donghee Choi, Jaewoo Kang |
CIKM | 3 |
| 2025 | GPO-VAE: modeling explainable gene perturbation responses utilizing GRN-aligned parameter optimizationabstractMOTIVATION: Predicting cellular responses to genetic perturbations is essential for understanding biological systems and developing targeted therapeutic strategies. While variational autoencoders (VAEs) have shown promise in modeling perturbation responses, their limited explainability poses a significant challenge, as the learned features often lack clear biological meaning. Nevertheless, model explainability is one of the most important aspects in the realm of biological AI. One of the most effective ways to achieve explainability is incorporating the concept of gene regulatory networks (GRNs) in designing deep learning models such as VAEs. GRNs elicit the underlying causal relationships between genes and are capable of explaining the transcriptional responses caused by genetic perturbation treatments. RESULTS: We propose GPO-VAE, an explainable VAE enhanced by GRN-aligned Parameter Optimization that explicitly models gene regulatory networks in the latent space. Our key approach is to optimize the learnable parameters related to latent perturbation effects toward GRN-aligned explainability. Experimental results on perturbation prediction show our model achieves state-of-the-art performance in predicting transcriptional responses across multiple benchmark datasets. Furthermore, additional results on evaluating the GRN inference task reveal our model's ability to generate meaningful GRNs compared to other methods. According to qualitative analysis, GPO-VAE possesses the ability to construct biologically explainable GRNs that align with experimentally validated regulatory pathways. AVAILABILITY AND IMPLEMENTATION: GPO-VAE is available at https://github.com/dmis-lab/GPO-VAE. Seungheun Baek, Soyon Park, Yan Ting Chok, Keonwoo Kim 0002, Jaewoo Kang |
Bioinform. | 4 |
| 2025 | ArcDFI: Attention regularization guided by CYP450 interactions for predicting drug-food interactionsabstractCYP450 isoenzymes are known to be deeply involved in the formation of drug-food interactions (DFI). Previously introduced computational approaches for predicting DFIs do not take drug-CYP450 interactions (DCI) into account and have limited generalizability in handling compounds unseen during model training. We introduce ArcDFI, a model that utilizes attention regularization guided by CYP450 interactions to predict drug-food interactions. Experiments on DFI prediction-evaluated under stringent cold-drug and cold-food settings-show that our model outperforms ten baseline approaches, demonstrating the effectiveness of incorporating CYP450 interactions. Analysis of its attention mechanism provides insight into its current understanding of DCI and how they are related to its DFI predictions. To the best of our knowledge, ArcDFI is the first DFI prediction model that incorporates the concept of DCI, resulting in improved predictive generalizability and model explainability. ArcDFI is available at https://github.com/KU-MedAI/ArcDFI. Keonwoo Kim 0002, Jaewoo Kang, Donghyeon Park, Minji Jeon |
PLoS Comput. Biol. | 1 |
| 2024 | DeepClair: Utilizing Market Forecasts for Effective Portfolio SelectionabstractUtilizing market forecasts is pivotal in optimizing portfolio selection strategies. We introduce DeepClair, a novel framework for portfolio selection. DeepClair leverages a transformer-based time-series forecasting model to predict market trends, facilitating more informed and adaptable portfolio decisions. To integrate the forecasting model into a deep reinforcement learning-driven portfolio selection framework, we introduced a two-step strategy: first, pre-training the time-series model on market data, followed by fine-tuning the portfolio selection architecture using this model. Additionally, we investigated the optimization technique, Low-Rank Adaptation (LoRA), to enhance the pre-trained forecasting model for fine-tuning in investment scenarios. This work bridges market forecasting and portfolio selection, facilitating the advancement of investment strategies. Donghee Choi, Jinkyu Kim 0004, Keonwoo Kim 0002, Jaewoo Kang |
CIKM | 3 |
| 2024 | LAPIS: Language Model-Augmented Police Investigation System
Heedou Kim, Jiwoo Lee, Chanwoong Yoon, Donghee Choi, Keonwoo Kim 0002, Jaewoo Kang |
CIKM | 6 |
| 2024 | CookingSense: A Culinary Knowledgebase with Multidisciplinary AssertionsabstractThis paper introduces CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes, from which knowledge covering a broad range of aspects is acquired. CookingSense is constructed through a series of dictionary-based filtering and language model-based semantic filtering techniques, which results in a rich knowledgebase of multidisciplinary food-related assertions. Additionally, we present FoodBench, a novel benchmark to evaluate culinary decision support systems. From evaluations with FoodBench, we empirically prove that CookingSense improves the performance of retrieval augmented language models. We also validate the quality and variety of assertions in CookingSense through qualitative analysis. Donghee Choi, Keonwoo Kim 0002, Donghyeon Park, Mujeen Sung, Hyunjae Kim, Jaewoo Kang |
LREC/COLING | 2 |
| 2024 | MulinforCPI: enhancing precision of compound-protein interaction prediction through novel perspectives on multi-level information integrationabstractForecasting the interaction between compounds and proteins is crucial for discovering new drugs. However, previous sequence-based studies have not utilized three-dimensional (3D) information on compounds and proteins, such as atom coordinates and distance matrices, to predict binding affinity. Furthermore, numerous widely adopted computational techniques have relied on sequences of amino acid characters for protein representations. This approach may constrain the model's ability to capture meaningful biochemical features, impeding a more comprehensive understanding of the underlying proteins. Here, we propose a two-step deep learning strategy named MulinforCPI that incorporates transfer learning techniques with multi-level resolution features to overcome these limitations. Our approach leverages 3D information from both proteins and compounds and acquires a profound understanding of the atomic-level features of proteins. Besides, our research highlights the divide between first-principle and data-driven methods, offering new research prospects for compound-protein interaction tasks. We applied the proposed method to six datasets: Davis, Metz, KIBA, CASF-2016, DUD-E and BindingDB, to evaluate the effectiveness of our approach. Ngoc-Quang Nguyen, Se-Jeong Park, Keonwoo Kim 0002, Jaewoo Kang |
Briefings Bioinform. | 3 |
| 2024 | MolPLA: a molecular pretraining framework for learning cores, R-groups and their linker jointsabstractMOTIVATION: Molecular core structures and R-groups are essential concepts in drug development. Integration of these concepts with conventional graph pre-training approaches can promote deeper understanding in molecules. We propose MolPLA, a novel pre-training framework that employs masked graph contrastive learning in understanding the underlying decomposable parts in molecules that implicate their core structure and peripheral R-groups. Furthermore, we formulate an additional framework that grants MolPLA the ability to help chemists find replaceable R-groups in lead optimization scenarios. RESULTS: Experimental results on molecular property prediction show that MolPLA exhibits predictability comparable to current state-of-the-art models. Qualitative analysis implicate that MolPLA is capable of distinguishing core and R-group sub-structures, identifying decomposable regions in molecules and contributing to lead optimization scenarios by rationally suggesting R-group replacements given various query core templates. AVAILABILITY AND IMPLEMENTATION: The code implementation for MolPLA and its pre-trained model checkpoint is available at https://github.com/dmis-lab/MolPLA. Keonwoo Kim 0002, Jueon Park, Soyon Park, Seungheun Baek, Jun-Hyun Lee, Ngoc-Quang Nguyen, Jaewoo Kang |
Bioinform. | 1 |
| 2023 | ArkDTA: attention regularization guided by non-covalent interactions for explainable drug-target binding affinity predictionabstractMOTIVATION: Protein-ligand binding affinity prediction is a central task in drug design and development. Cross-modal attention mechanism has recently become a core component of many deep learning models due to its potential to improve model explainability. Non-covalent interactions (NCIs), one of the most critical domain knowledge in binding affinity prediction task, should be incorporated into protein-ligand attention mechanism for more explainable deep drug-target interaction models. We propose ArkDTA, a novel deep neural architecture for explainable binding affinity prediction guided by NCIs. RESULTS: Experimental results show that ArkDTA achieves predictive performance comparable to current state-of-the-art models while significantly improving model explainability. Qualitative investigation into our novel attention mechanism reveals that ArkDTA can identify potential regions for NCIs between candidate drug compounds and target proteins, as well as guiding internal operations of the model in a more interpretable and domain-aware manner. AVAILABILITY: ArkDTA is available at https://github.com/dmis-lab/ArkDTA. CONTACT: [email protected]. Keonwoo Kim 0002, Junseok Choe, Seungheun Baek, Jueon Park, Chaeeun Lee, Minjae Ju, Jaewoo Kang |
Bioinform. | 1 |
| 2023 | KitchenScale: Learning to predict ingredient quantities from recipe contexts
Donghee Choi, Keonwoo Kim 0002, Samy Badreddine, Hajung Kim, Donghyeon Park, Jaewoo Kang |
Expert Syst. Appl. | 2 |
| 2022 | RecipeMind: Guiding Ingredient Choices from Food Pairing to Recipe Completion using Cascaded Set TransformerabstractWe propose a computational approach for recipe ideation, a downstream task that helps users select and gather ingredients for creating dishes. To perform this task, we developed RecipeMind, a food affinity score prediction model that quantifies the suitability of adding an ingredient to set of other ingredients. We constructed a large-scale dataset containing ingredient co-occurrence based scores to train and evaluate RecipeMind on food affinity score prediction. Deployed in recipe ideation, RecipeMind helps the user expand an initial set of ingredients by suggesting additional ingredients. Experiments and qualitative analysis show RecipeMind's potential in fulfilling its assistive role in cuisine domain. Keonwoo Kim 0002, Donghee Choi, Kana Maruyama, Hajung Kim, Donghyeon Park, Jaewoo Kang |
CIKM | 1 |
| 2020 | Improved survival analysis by learning shared genomic information from pan-cancer dataabstractMOTIVATION: Recent advances in deep learning have offered solutions to many biomedical tasks. However, there remains a challenge in applying deep learning to survival analysis using human cancer transcriptome data. As the number of genes, the input variables of survival model, is larger than the amount of available cancer patient samples, deep-learning models are prone to overfitting. To address the issue, we introduce a new deep-learning architecture called VAECox. VAECox uses transfer learning and fine tuning. RESULTS: We pre-trained a variational autoencoder on all RNA-seq data in 20 TCGA datasets and transferred the trained weights to our survival prediction model. Then we fine-tuned the transferred weights during training the survival model on each dataset. Results show that our model outperformed other previous models such as Cox Proportional Hazard with LASSO and ridge penalty and Cox-nnet on the 7 of 10 TCGA datasets in terms of C-index. The results signify that the transferred information obtained from entire cancer transcriptome data helped our survival prediction model reduce overfitting and show robust performance in unseen cancer patient samples. AVAILABILITY AND IMPLEMENTATION: Our implementation of VAECox is available at https://github.com/dmis-lab/VAECox. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sunkyu Kim, Keonwoo Kim 0002, Junseok Choe, Inggeol Lee, Jaewoo Kang |
Bioinform. | 2 |
| 2019 | KitcheNette: Predicting and Ranking Food Ingredient Pairings using Siamese Neural NetworkabstractAs a vast number of ingredients exist in the culinary world, there are countless food ingredient pairings, but only a small number of pairings have been adopted by chefs and studied by food researchers. In this work, we propose KitcheNette which is a model that predicts food ingredient pairing scores and recommends optimal ingredient pairings. KitcheNette employs Siamese neural networks and is trained on our annotated dataset containing 300K scores of pairings generated from numerous ingredients in food recipes. As the results demonstrate, our model not only outperforms other baseline models, but also can recommend complementary food pairings and discover novel ingredient pairings. Donghyeon Park, Keonwoo Kim 0002, Yonggyu Park, Jungwoon Shin, Jaewoo Kang |
IJCAI | 2 |