VLDB 2026 Research / reviewers in the wild / expert
Seunghun Lee 0001
dblp:77/7676-1
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-9377-2832ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Optimization for machine learning · 34% Language models and text generation · 19% Graph learning · 18% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.7 | 3 | 2025 | Inversion-based Latent Bayesian Optimization · NeurIPS 2024 Advancing Bayesian Optimization via Learning Correlated Latent Space · NeurIPS 2023 PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs · NeurIPS 2025 |
Machine learning › Graph learning
graph neural network |
1.2 | 2 | 2023 | NuTrea: Neural Tree Search for Context-guided Multi-hop KGQA · NeurIPS 2023 Metropolis-Hastings Data Augmentation for Graph Neural Networks · NeurIPS 2021 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization |
0.9 | 1 | 2025 | PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs · NeurIPS 2025 |
Mathematical optimization › bayesian optimization
acquisition function optimization |
0.9 | 1 | 2025 | Latent Bayesian Optimization via Autoregressive Normalizing Flows · ICLR 2025 |
Mathematical optimization
bayesian optimization |
0.9 | 1 | 2025 | Latent Bayesian Optimization via Autoregressive Normalizing Flows · ICLR 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
latent space bayesian optimization |
0.8 | 1 | 2024 | Inversion-based Latent Bayesian Optimization · NeurIPS 2024 |
Machine learning › Optimization for machine learning
black-box optimization |
0.7 | 1 | 2023 | Advancing Bayesian Optimization via Learning Correlated Latent Space · NeurIPS 2023 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.7 | 1 | 2023 | NuTrea: Neural Tree Search for Context-guided Multi-hop KGQA · NeurIPS 2023 |
Machine learning › Generative modeling
latent space optimization |
0.7 | 1 | 2023 | Advancing Bayesian Optimization via Learning Correlated Latent Space · NeurIPS 2023 |
Machine learning › Generative modeling
variational autoencoder |
0.7 | 1 | 2023 | Advancing Bayesian Optimization via Learning Correlated Latent Space · NeurIPS 2023 |
Machine learning › Graph learning › graph neural network
graph data augmentation |
0.5 | 1 | 2021 | Metropolis-Hastings Data Augmentation for Graph Neural Networks · NeurIPS 2021 |
Machine learning › Learning paradigms
semi-supervised learning |
0.5 | 1 | 2021 | Metropolis-Hastings Data Augmentation for Graph Neural Networks · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.2 | 1 | 2024 | Inversion-based Latent Bayesian Optimization · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
soft prompts · 0.9score consistency regularization · 0.9preimage structure · 0.9normalizing flow · 0.9autoregressive model · 0.9trust region anchor selection · 0.8inversion method · 0.8bi-level optimization · 0.8relation frequency-inverse entity frequency · 0.7neural tree search · 0.7message passing · 0.7lipschitz regularization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Bayesian Optimization via Autoregressive Normalizing FlowsabstractBayesian Optimization (BO) has been recognized for its effectiveness in optimizing expensive and complex objective functions.
Recent advancements in Latent Bayesian Optimization (LBO) have shown promise by integrating generative models such as variational autoencoders (VAEs) to manage the complexity of high-dimensional and structured data spaces.
However, existing LBO approaches often suffer from the value discrepancy problem, which arises from the reconstruction gap between input and latent spaces.
This value discrepancy problem propagates errors throughout the optimization process, leading to suboptimal outcomes.
To address this issue, we propose a Normalizing Flow-based Bayesian Optimization (NF-BO), which utilizes normalizing flow as a generative model to establish one-to-one encoding function from the input space to the latent space, along with its left-inverse decoding function, eliminating the reconstruction gap. Specifically, we introduce SeqFlow, an autoregressive normalizing flow for sequence data.
In addition, we develop a new candidate sampling strategy that dynamically adjusts the exploration probability for each token based on its importance.
Through extensive experiments, our NF-BO method demonstrates superior performance in molecule generation tasks, significantly outperforming both traditional and recent LBO approaches. Seunghun Lee 0001, Jinyoung Park 0005, Jaewon Chu, Minseo Yoon, Hyunwoo J. Kim |
ICLR | 1 |
| 2025 | PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMsabstractLarge language models (LLMs) have achieved remarkable success across diverse domains, due to their strong instruction-following capabilities. This raised interest in optimizing instructions for black-box LLMs, whose internal parameters are inaccessible but popular for their strong performance and ease of use. Recent approaches leverage white-box LLMs to assist instruction optimization for black-box LLMs by generating instructions from soft prompts. However, white-box LLMs often map different soft prompts to the same instruction, leading to redundant queries to the black-box model. While previous studies regarded this many-to-one mapping as a redundancy to be avoided, we reinterpret it as useful prior knowledge that can enhance the optimization performance. To this end, we introduce PREimage-informed inSTruction Optimization (PRESTO), a novel framework that leverages the preimage structure of soft prompts to improve query efficiency. PRESTO consists of three key components: (1) score sharing, which shares the evaluation score with all soft prompts in a preimage; (2) preimage-based initialization, which select initial data points that maximize search space coverage using preimage information; and (3) score consistency regularization, which enforces prediction consistency within each preimage. By leveraging preimages, PRESTO observes 14 times more scored data under the same query budget, resulting in more efficient optimization. Experimental results on 33 instruction optimization tasks demonstrate the superior performance of PRESTO. Jaewon Chu, Seunghun Lee 0001, Hyunwoo J. Kim |
NeurIPS | 2 |
| 2024 | Inversion-based Latent Bayesian OptimizationabstractLatent Bayesian optimization (LBO) approaches have successfully adopted Bayesian optimization over a continuous latent space by employing an encoder-decoder architecture to address the challenge of optimization in a high dimensional or discrete input space. LBO learns a surrogate model to approximate the black-box objective function in the latent space. However, we observed that most LBO methods suffer from the `misalignment problem', which is induced by the reconstruction error of the encoder-decoder architecture. It hinders learning an accurate surrogate model and generating high-quality solutions. In addition, several trust region-based LBO methods select the anchor, the center of the trust region, based solely on the objective function value without considering the trust region's potential to enhance the optimization process. To address these issues, we propose $\textbf{Inv}$ersion-based Latent $\textbf{B}$ayesian $\textbf{O}$ptimization (InvBO), a plug-and-play module for LBO. InvBO consists of two components: an inversion method and a potential-aware trust region anchor selection. The inversion method searches the latent code that completely reconstructs the given target data. The potential-aware trust region anchor selection considers the potential capability of the trust region for better local optimization. Experimental results demonstrate the effectiveness of InvBO on nine real-world benchmarks, such as molecule design and arithmetic expression fitting tasks. Code is available at https://github.com/mlvlab/InvBO. Jaewon Chu, Jinyoung Park 0005, Seunghun Lee 0001, Hyunwoo J. Kim |
NeurIPS | 3 |
| 2023 | NuTrea: Neural Tree Search for Context-guided Multi-hop KGQAabstractMulti-hop Knowledge Graph Question Answering (KGQA) is a task that involves retrieving nodes from a knowledge graph (KG) to answer natural language questions. Recent GNN-based approaches formulate this task as a KG path searching problem, where messages are sequentially propagated from the seed node towards the answer nodes. However, these messages are past-oriented, and they do not consider the full KG context. To make matters worse, KG nodes often represent pronoun entities and are sometimes encrypted, being uninformative in selecting between paths. To address these problems, we propose Neural Tree Search (NuTrea), a tree search-based GNN model that incorporates the broader KG context. Our model adopts a message-passing scheme that probes the unreached subtree regions to boost the past-oriented embeddings. In addition, we introduce the Relation Frequency-Inverse Entity Frequency (RF-IEF) node embedding that considers the global KG context to better characterize ambiguous KG nodes. The general effectiveness of our approach is demonstrated through experiments on three major multi-hop KGQA benchmark datasets, and our extensive analyses further validate its expressiveness and robustness. Overall, NuTrea provides a powerful means to query the KG with complex natural language questions. Code is available at https://github.com/mlvlab/NuTrea. Hyeong Kyu Choi, Seunghun Lee 0001, Jaewon Chu, Hyunwoo J. Kim |
NeurIPS | 2 |
| 2023 | Advancing Bayesian Optimization via Learning Correlated Latent SpaceabstractBayesian optimization is a powerful method for optimizing black-box functions with limited function evaluations. Recent works have shown that optimization in a latent space through deep generative models such as variational autoencoders leads to effective and efficient Bayesian optimization for structured or discrete data. However, as the optimization does not take place in the input space, it leads to an inherent gap that results in potentially suboptimal solutions. To alleviate the discrepancy, we propose Correlated latent space Bayesian Optimization (CoBO), which focuses on learning correlated latent spaces characterized by a strong correlation between the distances in the latent space and the distances within the objective function. Specifically, our method introduces Lipschitz regularization, loss weighting, and trust region recoordination to minimize the inherent gap around the promising areas. We demonstrate the effectiveness of our approach on several optimization tasks in discrete data, such as molecule design and arithmetic expression fitting, and achieve high performance within a small budget. Seunghun Lee 0001, Jaewon Chu, Sihyeon Kim, Juyeon Ko, Hyunwoo J. Kim |
NeurIPS | 1 |
| 2022 | Graph Transformer Networks: Learning meta-path graphs to improve GNNsabstractGraph Neural Networks (GNNs) have been widely applied to various fields due to their powerful representations of graph-structured data. Despite the success of GNNs, most existing GNNs are designed to learn node representations on the fixed and homogeneous graphs. The limitations especially become problematic when learning representations on a misspecified graph or a heterogeneous graph that consists of various types of nodes and edges. To address these limitations, we propose Graph Transformer Networks (GTNs) that are capable of generating new graph structures, which preclude noisy connections and include useful connections (e.g., meta-paths) for tasks, while learning effective node representations on the new graphs in an end-to-end fashion. We further propose enhanced version of GTNs, Fast Graph Transformer Networks (FastGTNs), that improve scalability of graph transformations. Compared to GTNs, FastGTNs are up to 230× and 150× faster in inference and training, and use up to 100× and 148× less memory while allowing the identical graph transformations as GTNs. In addition, we extend graph transformations to the semantic proximity of nodes allowing non-local operations beyond meta-paths. Extensive experiments on both homogeneous graphs and heterogeneous graphs show that GTNs and FastGTNs with non-local operations achieve the state-of-the-art performance for node classification tasks. The code is available: https://github.com/seongjunyun/Graph_Transformer_Networks. Seongjun Yun, Minbyul Jeong, Sungdong Yoo, Seunghun Lee 0001, Sean S. Yi, Raehyun Kim, Jaewoo Kang, Hyunwoo J. Kim |
Neural Networks | 4 |
| 2021 | Metropolis-Hastings Data Augmentation for Graph Neural NetworksabstractGraph Neural Networks (GNNs) often suffer from weak-generalization due to sparsely labeled data despite their promising results on various graph-based tasks. Data augmentation is a prevalent remedy to improve the generalization ability of models in many domains. However, due to the non-Euclidean nature of data space and the dependencies between samples, designing effective augmentation on graphs is challenging. In this paper, we propose a novel framework Metropolis-Hastings Data Augmentation (MH-Aug) that draws augmented graphs from an explicit target distribution for semi-supervised learning. MH-Aug produces a sequence of augmented graphs from the target distribution enables flexible control of the strength and diversity of augmentation. Since the direct sampling from the complex target distribution is challenging, we adopt the Metropolis-Hastings algorithm to obtain the augmented samples. We also propose a simple and effective semi-supervised learning strategy with generated samples from MH-Aug. Our extensive experiments demonstrate that MH-Aug can generate a sequence of samples according to the target distribution to significantly improve the performance of GNNs. Hyeon-Jin Park, Seunghun Lee 0001, Sihyeon Kim, Jinyoung Park 0005, Jisu Jeong, Jung-Woo Ha 0001, Hyunwoo J. Kim |
NeurIPS | 2 |