VLDB 2026 Research / reviewers in the wild / expert
Sanghyeok Choi
dblp:338/9899
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0004-7589-9425ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 49% Reinforcement learning · 42% Motion planning and robot control · 7% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 93% Algorithms and data structures · 7% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
off-policy reinforcement learning |
2.5 | 3 | 2025 | Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025 Adaptive teachers for amortized samplers · ICLR 2025 Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024 |
Machine learning › Generative modeling
generative flow networks |
1.6 | 2 | 2025 | Adaptive teachers for amortized samplers · ICLR 2025 Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024 |
Mathematical optimization
combinatorial optimization |
1.6 | 2 | 2025 | RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization Benchmark · KDD (2) 2025 Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024 |
Machine learning › Generative modeling
amortized sampling |
0.9 | 1 | 2025 | Adaptive teachers for amortized samplers · ICLR 2025 |
Machine learning › Reinforcement learning
biological sequence design |
0.9 | 1 | 2025 | Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Adaptive teachers for amortized samplers · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | Adaptive teachers for amortized samplers · ICLR 2025 |
Bioinformatics and computational biology › synthetic biology
biological sequence design |
0.9 | 1 | 2025 | Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025 |
Mathematical optimization › evolutionary computation
genetic algorithm |
0.9 | 1 | 2025 | Neural Genetic Search in Discrete Spaces · ICML 2025 |
Machine learning › Generative modeling
molecular generation |
0.8 | 1 | 2024 | Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024 |
Robotics › Motion planning and robot control › motion planning
sequential planning |
0.8 | 1 | 2024 | Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024 |
Mathematical optimization › combinatorial optimization
routing problems |
0.8 | 1 | 2024 | Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024 |
Machine learning › Generative modeling › generative model evaluation
mode coverage |
0.3 | 1 | 2025 | Adaptive teachers for amortized samplers · ICLR 2025 |
Machine learning › Reinforcement learning
reinforcement learning for combinatorial optimization |
0.3 | 1 | 2025 | RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization Benchmark · KDD (2) 2025 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.3 | 1 | 2025 | Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 3.3uncertainty estimation · 1.7proxy model · 1.7parent-conditioned generation · 1.7genetic algorithm · 1.7crossover · 1.7conservative search · 1.7off-policy reinforcement learning · 0.9generative flow networks · 0.9adaptive teacher-student curriculum · 0.9transformer · 0.8heuristic search · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ant Colony Sampling with GFlowNets for Combinatorial OptimizationabstractWe present the Generative Flow Ant Colony Sampler (GFACS), a novel meta-heuristic method that hierarchically combines amortized inference and parallel stochastic search. Our method first leverages Generative Flow Networks (GFlowNets) to amortize a multi-modal prior distribution over combinatorial solution space that encompasses both high-reward and diversified solutions. This prior is iteratively updated via parallel stochastic search in the spirit of Ant Colony Optimization (ACO), leading to the posterior distribution that generates near-optimal solutions. Extensive experiments across seven combinatorial optimization problems demonstrate GFACS’s promising performances. Minsu Kim 0004, Sanghyeok Choi, Hyeonah Kim, Jiwoo Son, Jinkyoo Park, Yoshua Bengio |
AISTATS | 2 |
| 2025 | Adaptive teachers for amortized samplersabstractAmortized inference is the task of training a parametric model, such as a neural network, to approximate a distribution with a given unnormalized density where exact sampling is intractable. When sampling is modeled as a sequential decision-making process, reinforcement learning (RL) methods, such as generative flow networks, can be used to train the sampling policy. Off-policy RL training facilitates the discovery of diverse, high-reward candidates, but existing methods still face challenges in efficient exploration. We propose to use an adaptive training distribution (the Teacher) to guide the training of the primary amortized sampler (the Student). The Teacher, an auxiliary behavior model, is trained to sample high-loss regions of the Student and can generalize across unexplored modes, thereby enhancing mode coverage by providing an efficient training curriculum. We validate the effectiveness of this approach in a synthetic environment designed to present an exploration challenge, two diffusion-based sampling tasks, and four biochemical discovery tasks demonstrating its ability to improve sample efficiency and mode coverage. Source code is available at https://github.com/alstn12088/adaptive-teacher. Minsu Kim 0004, Sanghyeok Choi, Taeyoung Yun, Emmanuel Bengio, Leo Feng, Jarrid Rector-Brooks, Sungsoo Ahn, Jinkyoo Park, Nikolay Malkin, Yoshua Bengio |
ICLR | 2 |
| 2025 | Improved Off-policy Reinforcement Learning in Biological Sequence DesignabstractDesigning biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $\delta$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $\delta$, then denoise them using our policy. We further adapt $\delta$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design. Hyeonah Kim, Minsu Kim 0004, Taeyoung Yun, Sanghyeok Choi, Emmanuel Bengio, Alex Hernández-García, Jinkyoo Park |
ICML | 4 |
| 2025 | Neural Genetic Search in Discrete SpacesabstractEffective search methods are crucial for improving the performance of deep generative models at test time. In this paper, we introduce a novel test-time search method, Neural Genetic Search (NGS), which incorporates the evolutionary mechanism of genetic algorithms into the generation procedure of deep models. The core idea behind NGS is its crossover, which is defined as parent-conditioned generation using trained generative models. This approach offers a versatile and easy-to-implement search algorithm for deep generative models. We demonstrate the effectiveness and flexibility of NGS through experiments across three distinct domains: routing problems, adversarial prompt generation for language models, and molecular design. Hyeonah Kim, Sanghyeok Choi, Jiwoo Son, Jinkyoo Park, Changhyun Kwon 0001 |
ICML | 2 |
| 2025 | RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization BenchmarkabstractCombinatorial optimization (CO) is fundamental to several realworld applications, from logistics and scheduling to hardware design and resource allocation.Deep reinforcement learning (RL) has recently shown significant benefits in solving CO problems, reducing reliance on domain expertise and improving computational efficiency.However, the absence of a unified benchmarking framework leads to inconsistent evaluations, limits reproducibility, and increases engineering overhead, raising barriers to adoption for new researchers.To address these challenges, we introduce RL4CO, a unified and extensive benchmark with in-depth library coverage of 27 CO problem environments and 23 state-of-the-art baselines.Built on efficient software libraries and best practices in implementation, RL4CO features modularized implementation and flexible configurations of diverse environments, policy architectures, RL algorithms, and utilities with extensive documentation.RL4CO helps researchers build on existing successes while exploring and developing their own designs, facilitating the entire research process by decoupling science from heavy engineering.We finally provide extensive benchmark studies to inspire new insights and future work.RL4CO has already attracted numerous researchers in the community and is open-sourced at https://github.com/ai4co/rl4co 1 . Federico Berto, Chuanbo Hua, Junyoung Park 0002, Laurin Luttmann, Yining Ma 0001, Fanchen Bu, Jiarui Wang 0002, Haoran Ye, Minsu Kim 0004, Sanghyeok Choi, Nayeli Gast Zepeda, André Hottung, Jianan Zhou 0002, Jieyi Bi, Fei Liu 0044, Hyeonah Kim, Jiwoo Son, Haeyeon Kim, Davide Angioni, Wouter Kool 0001, Zhiguang Cao, Qingfu Zhang 0001, Joungho Kim, Jie Zhang 0002, Kijung Shin, Cathy Wu 0002, Sungsoo Ahn, Guojie Song, Changhyun Kwon 0001, Kevin Tierney, Jinkyoo Park |
KDD (2) | 10 |
| 2024 | Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity ContextabstractMin-max routing problems aim to minimize the maximum tour length among multiple agents as they collaboratively visit all cities, i.e., the completion time. These problems include impactful real-world applications but are known as NP-hard. Existing methods are facing challenges, particularly in large-scale problems that require the coordination of numerous agents to cover thousands of cities. This paper proposes Equity-Transformer to solve large-scale min-max routing problems. First, we model min-max routing problems into sequential planning, reducing the complexity and enabling the use of a powerful Transformer architecture. Second, we propose key inductive biases that ensure equitable workload distribution among agents. The effectiveness of Equity-Transformer is demonstrated through its superior performance in two representative min-max routing tasks: the min-max multi-agent traveling salesman problem (min-max mTSP) and the min-max multi-agent pick-up and delivery problem (min-max mPDP). Notably, our method achieves significant reductions of runtime, approximately 335 times, and cost values of about 53% compared to a competitive heuristic (LKH3) in the case of 100 vehicles with 1,000 cities of mTSP. We provide reproducible source code: https://github.com/kaist-silab/equity-transformer. Jiwoo Son, Minsu Kim 0004, Sanghyeok Choi, Hyeonah Kim, Jinkyoo Park |
AAAI | 3 |
| 2024 | Genetic-guided GFlowNets for Sample Efficient Molecular OptimizationabstractThe challenge of discovering new molecules with desired properties is crucial in domains like drug discovery and material design. Recent advances in deep learning-based generative methods have shown promise but face the issue of sample efficiency due to the computational expense of evaluating the reward function. This paper proposes a novel algorithm for sample-efficient molecular optimization by distilling a powerful genetic algorithm into deep generative policy using GFlowNets training, the off-policy method for amortized inference. This approach enables the deep generative policy to learn from domain knowledge, which has been explicitly integrated into the genetic algorithm. Our method achieves state-of-the-art performance in the official molecular optimization benchmark, significantly outperforming previous methods. It also demonstrates effectiveness in designing inhibitors against SARS-CoV-2 with substantially fewer reward calls. Hyeonah Kim, Minsu Kim 0004, Sanghyeok Choi, Jinkyoo Park |
NeurIPS | 3 |