Sanghyeok Choi

dblp:338/9899 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0004-7589-9425ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 49% Reinforcement learning · 42% Motion planning and robot control · 7%
Theoretical computer science
3 papers
Mathematical optimization · 93% Algorithms and data structures · 7%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
off-policy reinforcement learning
2.532025
Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025
Adaptive teachers for amortized samplers · ICLR 2025
Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024
Machine learning › Generative modeling
generative flow networks
1.622025
Adaptive teachers for amortized samplers · ICLR 2025
Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024
Mathematical optimization
combinatorial optimization
1.622025
RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization Benchmark · KDD (2) 2025
Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024
Machine learning › Generative modeling
amortized sampling
0.912025
Adaptive teachers for amortized samplers · ICLR 2025
Machine learning › Reinforcement learning
biological sequence design
0.912025
Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
Adaptive teachers for amortized samplers · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.912025
Adaptive teachers for amortized samplers · ICLR 2025
Bioinformatics and computational biology › synthetic biology
biological sequence design
0.912025
Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025
Mathematical optimization › evolutionary computation
genetic algorithm
0.912025
Neural Genetic Search in Discrete Spaces · ICML 2025
Machine learning › Generative modeling
molecular generation
0.812024
Genetic-guided GFlowNets for Sample Efficient Molecular Optimization · NeurIPS 2024
Robotics › Motion planning and robot control › motion planning
sequential planning
0.812024
Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024
Mathematical optimization › combinatorial optimization
routing problems
0.812024
Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context · AAAI 2024
Machine learning › Generative modeling › generative model evaluation
mode coverage
0.312025
Adaptive teachers for amortized samplers · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning for combinatorial optimization
0.312025
RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization Benchmark · KDD (2) 2025
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.312025
Improved Off-policy Reinforcement Learning in Biological Sequence Design · ICML 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 3.3uncertainty estimation · 1.7proxy model · 1.7parent-conditioned generation · 1.7genetic algorithm · 1.7crossover · 1.7conservative search · 1.7off-policy reinforcement learning · 0.9generative flow networks · 0.9adaptive teacher-student curriculum · 0.9transformer · 0.8heuristic search · 0.8
YearPublicationVenuePosition
2025 Ant Colony Sampling with GFlowNets for Combinatorial Optimization
abstract
We present the Generative Flow Ant Colony Sampler (GFACS), a novel meta-heuristic method that hierarchically combines amortized inference and parallel stochastic search. Our method first leverages Generative Flow Networks (GFlowNets) to amortize a multi-modal prior distribution over combinatorial solution space that encompasses both high-reward and diversified solutions. This prior is iteratively updated via parallel stochastic search in the spirit of Ant Colony Optimization (ACO), leading to the posterior distribution that generates near-optimal solutions. Extensive experiments across seven combinatorial optimization problems demonstrate GFACS’s promising performances.
Minsu Kim 0004, Sanghyeok Choi, Hyeonah Kim, Jiwoo Son, Jinkyoo Park, Yoshua Bengio
AISTATS2
2025 Adaptive teachers for amortized samplers
abstract
Amortized inference is the task of training a parametric model, such as a neural network, to approximate a distribution with a given unnormalized density where exact sampling is intractable. When sampling is modeled as a sequential decision-making process, reinforcement learning (RL) methods, such as generative flow networks, can be used to train the sampling policy. Off-policy RL training facilitates the discovery of diverse, high-reward candidates, but existing methods still face challenges in efficient exploration. We propose to use an adaptive training distribution (the Teacher) to guide the training of the primary amortized sampler (the Student). The Teacher, an auxiliary behavior model, is trained to sample high-loss regions of the Student and can generalize across unexplored modes, thereby enhancing mode coverage by providing an efficient training curriculum. We validate the effectiveness of this approach in a synthetic environment designed to present an exploration challenge, two diffusion-based sampling tasks, and four biochemical discovery tasks demonstrating its ability to improve sample efficiency and mode coverage. Source code is available at https://github.com/alstn12088/adaptive-teacher.
Minsu Kim 0004, Sanghyeok Choi, Taeyoung Yun, Emmanuel Bengio, Leo Feng, Jarrid Rector-Brooks, Sungsoo Ahn, Jinkyoo Park, Nikolay Malkin, Yoshua Bengio
ICLR2
2025 Improved Off-policy Reinforcement Learning in Biological Sequence Design
abstract
Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $\delta$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $\delta$, then denoise them using our policy. We further adapt $\delta$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design.
Hyeonah Kim, Minsu Kim 0004, Taeyoung Yun, Sanghyeok Choi, Emmanuel Bengio, Alex Hernández-García, Jinkyoo Park
ICML4
2025 Neural Genetic Search in Discrete Spaces
abstract
Effective search methods are crucial for improving the performance of deep generative models at test time. In this paper, we introduce a novel test-time search method, Neural Genetic Search (NGS), which incorporates the evolutionary mechanism of genetic algorithms into the generation procedure of deep models. The core idea behind NGS is its crossover, which is defined as parent-conditioned generation using trained generative models. This approach offers a versatile and easy-to-implement search algorithm for deep generative models. We demonstrate the effectiveness and flexibility of NGS through experiments across three distinct domains: routing problems, adversarial prompt generation for language models, and molecular design.
Hyeonah Kim, Sanghyeok Choi, Jiwoo Son, Jinkyoo Park, Changhyun Kwon 0001
ICML2
2025 RL4CO: An Extensive Reinforcement Learning for Combinatorial Optimization Benchmark
abstract
Combinatorial optimization (CO) is fundamental to several realworld applications, from logistics and scheduling to hardware design and resource allocation.Deep reinforcement learning (RL) has recently shown significant benefits in solving CO problems, reducing reliance on domain expertise and improving computational efficiency.However, the absence of a unified benchmarking framework leads to inconsistent evaluations, limits reproducibility, and increases engineering overhead, raising barriers to adoption for new researchers.To address these challenges, we introduce RL4CO, a unified and extensive benchmark with in-depth library coverage of 27 CO problem environments and 23 state-of-the-art baselines.Built on efficient software libraries and best practices in implementation, RL4CO features modularized implementation and flexible configurations of diverse environments, policy architectures, RL algorithms, and utilities with extensive documentation.RL4CO helps researchers build on existing successes while exploring and developing their own designs, facilitating the entire research process by decoupling science from heavy engineering.We finally provide extensive benchmark studies to inspire new insights and future work.RL4CO has already attracted numerous researchers in the community and is open-sourced at https://github.com/ai4co/rl4co 1 .
Federico Berto, Chuanbo Hua, Junyoung Park 0002, Laurin Luttmann, Yining Ma 0001, Fanchen Bu, Jiarui Wang 0002, Haoran Ye, Minsu Kim 0004, Sanghyeok Choi, Nayeli Gast Zepeda, André Hottung, Jianan Zhou 0002, Jieyi Bi, Fei Liu 0044, Hyeonah Kim, Jiwoo Son, Haeyeon Kim, Davide Angioni, Wouter Kool 0001, Zhiguang Cao, Qingfu Zhang 0001, Joungho Kim, Jie Zhang 0002, Kijung Shin, Cathy Wu 0002, Sungsoo Ahn, Guojie Song, Changhyun Kwon 0001, Kevin Tierney, Jinkyoo Park
KDD (2)10
2024 Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequential Generation with Equity Context
abstract
Min-max routing problems aim to minimize the maximum tour length among multiple agents as they collaboratively visit all cities, i.e., the completion time. These problems include impactful real-world applications but are known as NP-hard. Existing methods are facing challenges, particularly in large-scale problems that require the coordination of numerous agents to cover thousands of cities. This paper proposes Equity-Transformer to solve large-scale min-max routing problems. First, we model min-max routing problems into sequential planning, reducing the complexity and enabling the use of a powerful Transformer architecture. Second, we propose key inductive biases that ensure equitable workload distribution among agents. The effectiveness of Equity-Transformer is demonstrated through its superior performance in two representative min-max routing tasks: the min-max multi-agent traveling salesman problem (min-max mTSP) and the min-max multi-agent pick-up and delivery problem (min-max mPDP). Notably, our method achieves significant reductions of runtime, approximately 335 times, and cost values of about 53% compared to a competitive heuristic (LKH3) in the case of 100 vehicles with 1,000 cities of mTSP. We provide reproducible source code: https://github.com/kaist-silab/equity-transformer.
Jiwoo Son, Minsu Kim 0004, Sanghyeok Choi, Hyeonah Kim, Jinkyoo Park
AAAI3
2024 Genetic-guided GFlowNets for Sample Efficient Molecular Optimization
abstract
The challenge of discovering new molecules with desired properties is crucial in domains like drug discovery and material design. Recent advances in deep learning-based generative methods have shown promise but face the issue of sample efficiency due to the computational expense of evaluating the reward function. This paper proposes a novel algorithm for sample-efficient molecular optimization by distilling a powerful genetic algorithm into deep generative policy using GFlowNets training, the off-policy method for amortized inference. This approach enables the deep generative policy to learn from domain knowledge, which has been explicitly integrated into the genetic algorithm. Our method achieves state-of-the-art performance in the official molecular optimization benchmark, significantly outperforming previous methods. It also demonstrates effectiveness in designing inhibitors against SARS-CoV-2 with substantially fewer reward calls.
Hyeonah Kim, Minsu Kim 0004, Sanghyeok Choi, Jinkyoo Park
NeurIPS3