Jungtaek Kim 0001

dblp:31/3193-1 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0002-1905-1399ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
abstract
Process Reward Models (PRMs) have proven effective at enhancing mathematical reasoning for Large Language Models (LLMs) by leveraging increased inference-time computation. However, they are predominantly trained on mathematical data and their generalizability to non-mathematical domains has not been rigorously studied. In response, this work first shows that current PRMs have poor performance in other domains. To address this limitation, we introduce ***VersaPRM***, a multi-domain PRM trained on synthetic reasoning data generated using our novel data generation and annotation method. VersaPRM achieves consistent performance gains across diverse domains. For instance, in the MMLU-Pro category of Law, VersaPRM via weighted majority voting, achieves a 7.9% performance gain over the majority voting baseline–surpassing Qwen2.5-Math-PRM's gain of 1.3%. We further contribute to the community by open-sourcing all data, code and models for VersaPRM.
Thomas Zeng 0003, Shuibai Zhang, Shutong Wu, Christian Classen, Daewon Chae, Ethan Ewer, Heeju Kim, Wonjun Kang, Jackson Kunde, Jungtaek Kim 0001, Hyung Il Koo, Kannan Ramchandran, Dimitris S. Papailiopoulos, Kangwook Lee 0001
ICML12
2024 Exploiting Preferences in Loss Functions for Sequential Recommendation via Weak Transitivity
abstract
A choice of optimization objective is immensely pivotal in the design of a recommender system as it affects the general modeling process of a user's intent from previous interactions. Existing approaches mainly adhere to three categories of loss functions: pairwise, pointwise, and setwise loss functions. Despite their effectiveness, a critical and common drawback of such objectives is viewing the next observed item as a unique positive while considering all remaining items equally negative. Such a binary label assignment is generally limited to assuring a higher recommendation score of the positive item, neglecting potential structures induced by varying preferences between other unobserved items. To alleviate this issue, we propose a novel method that extends original objectives to explicitly leverage the different levels of preferences as relative orders between their scores. Finally, we demonstrate the superior performance of our method compared to baseline objectives.
Hyunsoo Chung, Jungtaek Kim 0001, Hyungeun Jo, Hyungwon Choi
CIKM2
2024 Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions
abstract
Sorting is a fundamental operation of all computer systems, having been a long-standing significant research topic. Beyond the problem formulation of traditional sorting algorithms, we consider sorting problems for more abstract yet expressive inputs, e.g., multi-digit images and image fragments, through a neural sorting network. To learn a mapping from a high-dimensional input to an ordinal variable, the differentiability of sorting networks needs to be guaranteed. In this paper we define a softening error by a differentiable swap function, and develop an error-free swap function that holds a non-decreasing condition and differentiability. Furthermore, a permutation-equivariant Transformer network with multi-head attention is adopted to capture dependency between given inputs and also leverage its model capacity with self-attention. Experiments on diverse sorting benchmarks show that our methods perform better than or comparable to baseline methods.
Jungtaek Kim 0001, Jeongbeen Yoon, Minsu Cho
ICLR1
2024 Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
abstract
Adapting to a priori unknown noise level is a very important but challenging problem in sequential decision-making as efficient exploration typically requires knowledge of the noise level, which is often loosely specified. We report significant progress in addressing this issue in linear bandits in two respects. First, we propose a novel confidence set that is ’semi-adaptive’ to the unknown sub-Gaussian parameter $\sigma_*^2$ in the sense that the (normalized) confidence width scales with $\sqrt{d\sigma_*^2 + \sigma_0^2}$ where $d$ is the dimension and $\sigma_0^2$ is the specified sub-Gaussian parameter (known) that can be much larger than $\sigma_*^2$. This is a significant improvement over $\sqrt{d\sigma_0^2}$ of the standard confidence set of Abbasi-Yadkori et al. (2011), especially when $d$ is large. We show that this leads to an improved regret bound in linear bandits. Second, for bounded rewards, we propose a novel variance-adaptive confidence set that has a much improved numerical performance upon prior art. We then apply this confidence set to develop, as we claim, the first practical variance-adaptive linear bandit algorithm via an optimistic approach, which is enabled by our novel regret analysis technique. Both of our confidence sets rely critically on ‘regret equality’ from online learning. Our empirical evaluation in Bayesian optimization tasks shows that our algorithms demonstrate better or comparable performance compared to existing methods.
Kwang-Sung Jun, Jungtaek Kim 0001
ICML2
2024 Model Fusion through Bayesian Optimization in Language Model Fine-Tuning
abstract
Fine-tuning pre-trained models for downstream tasks is a widely adopted technique known for its adaptability and reliability across various domains. Despite its conceptual simplicity, fine-tuning entails several troublesome engineering choices, such as selecting hyperparameters and determining checkpoints from an optimization trajectory. To tackle the difficulty of choosing the best model, one effective solution is model fusion, which combines multiple models in a parameter space. However, we observe a large discrepancy between loss and metric landscapes during the fine-tuning of pre-trained language models. Building on this observation, we introduce a novel model fusion technique that optimizes both the desired metric and loss through multi-objective Bayesian optimization. In addition, to effectively select hyperparameters, we establish a two-stage procedure by integrating Bayesian optimization processes into our framework. Experiments across various downstream tasks show considerable performance improvements using our Bayesian optimization-guided method.
Chaeyun Jang, Hyungi Lee, Jungtaek Kim 0001, Juho Lee 0001
NeurIPS3
2023 Datasets and Benchmarks for Nanophotonic Structure and Parametric Design Simulations
abstract
Nanophotonic structures have versatile applications including solar cells, anti-reflective coatings, electromagnetic interference shielding, optical filters, and light emitting diodes. To design and understand these nanophotonic structures, electrodynamic simulations are essential. These simulations enable us to model electromagnetic fields over time and calculate optical properties. In this work, we introduce frameworks and benchmarks to evaluate nanophotonic structures in the context of parametric structure design problems. The benchmarks are instrumental in assessing the performance of optimization algorithms and identifying an optimal structure based on target optical properties. Moreover, we explore the impact of varying grid sizes in electrodynamic simulations, shedding light on how evaluation fidelity can be strategically leveraged in enhancing structure designs.
Jungtaek Kim 0001, Oliver Hinder, Paul W. Leu
NeurIPS1
2023 Generative Neural Fields by Mixtures of Neural Implicit Functions
abstract
We propose a novel approach to learning the generative neural fields represented by linear combinations of implicit basis networks. Our algorithm learns basis networks in the form of implicit neural representations and their coefficients in a latent space by either conducting meta-learning or adopting auto-decoding paradigms. The proposed method easily enlarges the capacity of generative neural fields by increasing the number of basis networks while maintaining the size of a network for inference to be small through their weighted model averaging. Consequently, sampling instances using the model is efficient in terms of latency and memory footprint. Moreover, we customize denoising diffusion probabilistic model for a target task to sample latent mixture coefficients, which allows our final model to generate unseen data effectively. Experiments show that our approach achieves competitive generation performance on diverse benchmarks for images, voxel data, and NeRF scenes without sophisticated designs for specific modalities and domains.
Tackgeun You, Mijeong Kim 0002, Jungtaek Kim 0001, Bohyung Han
NeurIPS3
2022 On Uncertainty Estimation by Tree-based Surrogate Models in Sequential Model-based Optimization
abstract
Sequential model-based optimization sequentially selects a candidate point by constructing a surrogate model with the history of evaluations, to solve a black-box optimization problem. Gaussian process (GP) regression is a popular choice as a surrogate model, because of its capability of calculating prediction uncertainty analytically. On the other hand, an ensemble of randomized trees is another option and has practical merits over GPs due to its scalability and easiness of handling continuous/discrete mixed variables. In this paper we revisit various ensembles of randomized trees to investigate their behavior in the perspective of prediction uncertainty estimation. Then, we propose a new way of constructing an ensemble of randomized trees, referred to as BwO forest, where bagging with oversampling is employed to construct bootstrapped samples that are used to build randomized trees with random splitting. Experimental results demonstrate the validity and good performance of BwO forest over existing tree-based models in various circumstances.
Jungtaek Kim 0001, Seungjin Choi 0001
AISTATS1
2022 On Evaluation Metrics for Graph Generative Models
Rylee Thompson, Boris Knyazev 0001, Elaheh Ghalebi, Jungtaek Kim 0001, Graham W. Taylor
ICLR4
2022 Learning to Assemble Geometric Shapes
abstract
Assembling parts into an object is a combinatorial problem that arises in a variety of contexts in the real world and involves numerous applications in science and engineering. Previous related work tackles limited cases with identical unit parts or jigsaw-style parts of textured shapes, which greatly mitigate combinatorial challenges of the problem. In this work, we introduce the more challenging problem of shape assembly, which involves textureless fragments of arbitrary shapes with indistinctive junctions, and then propose a learning-based approach to solving it. We demonstrate the effectiveness on shape assembly tasks with various scenarios, including the ones with abnormal fragments (e.g., missing and distorted), the different number of fragments, and different rotation discretization.
Jinhwi Lee, Jungtaek Kim 0001, Hyunsoo Chung, Jaesik Park, Minsu Cho
IJCAI2
2022 Combinatorial Bayesian optimization with random mapping functions to convex polytopes
abstract
Bayesian optimization is a popular method for solving the problem of global optimization of an expensive-to-evaluate black-box function. It relies on a probabilistic surrogate model of the objective function, upon which an acquisition function is built to determine where next to evaluate the objective function. In general, Bayesian optimization with Gaussian process regression operates on a continuous space. When input variables are categorical or discrete, an extra care is needed. A common approach is to use one-hot encoded or Boolean representation for categorical variables which might yield a combinatorial explosion problem. In this paper we present a method for Bayesian optimization in a combinatorial space, which can operate well in a large combinatorial space. The main idea is to use a random mapping which embeds the combinatorial space into a convex polytope in a continuous space, on which all essential process is performed to determine a solution to the black-box optimization in the combinatorial space. We describe our combinatorial Bayesian optimization algorithm and present its regret analysis. Numerical experiments demonstrate that our method shows satisfactory performance compared to existing methods.
Jungtaek Kim 0001, Seungjin Choi 0001, Minsu Cho
UAI1
2021 Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning
abstract
Discovering a solution in a combinatorial space is prevalent in many real-world problems but it is also challenging due to diverse complex constraints and the vast number of possible combinations. To address such a problem, we introduce a novel formulation, combinatorial construction, which requires a building agent to assemble unit primitives (i.e., LEGO bricks) sequentially -- every connection between two bricks must follow a fixed rule, while no bricks mutually overlap. To construct a target object, we provide incomplete knowledge about the desired target (i.e., 2D images) instead of exact and explicit volumetric information to the agent. This problem requires a comprehensive understanding of partial information and long-term planning to append a brick sequentially, which leads us to employ reinforcement learning. The approach has to consider a variable-sized action space where a large number of invalid actions, which would cause overlap between bricks, exist. To resolve these issues, our model, dubbed Brick-by-Brick, adopts an action validity prediction network that efficiently filters invalid actions for an actor-critic network. We demonstrate that the proposed method successfully learns to construct an unseen object conditioned on a single image or multiple views of a target object.
Hyunsoo Chung, Jungtaek Kim 0001, Boris Knyazev 0001, Jinhwi Lee, Graham W. Taylor, Jaesik Park, Minsu Cho
NeurIPS2
2021 Bayesian optimization with approximate set kernels
Jungtaek Kim 0001, Michael McCourt, Tackgeun You, Saehoon Kim, Seungjin Choi 0001
Mach. Learn.1
2020 Bootstrapping neural processes
abstract
Unlike in the traditional statistical modeling for which a user typically hand-specify a prior, Neural Processes (NPs) implicitly define a broad class of stochastic processes with neural networks. Given a data stream, NP learns a stochastic process that best describes the data. While this ``data-driven'' way of learning stochastic processes has proven to handle various types of data, NPs still relies on an assumption that uncertainty in stochastic processes is modeled by a single latent variable, which potentially limits the flexibility. To this end, we propose the Bootstrapping Neural Process (BNP), a novel extension of the NP family using the bootstrap. The bootstrap is a classical data-driven technique for estimating uncertainty, which allows BNP to learn the stochasticity in NPs without assuming a particular form. We demonstrate the efficacy of BNP on various types of data and its robustness in the presence of model-data mismatch.
Juho Lee 0001, Yoonho Lee 0001, Jungtaek Kim 0001, Eunho Yang, Sung Ju Hwang, Yee Whye Teh
NeurIPS3
2020 On Local Optimizers of Acquisition Functions in Bayesian Optimization
Jungtaek Kim 0001, Seungjin Choi 0001
ECML/PKDD (2)1
2019 Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks
abstract
Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set, models used to address them should be permutation invariant. We present an attention-based neural network module, the Set Transformer, specifically designed to model interactions among elements in the input set. The model consists of an encoder and a decoder, both of which rely on attention mechanisms. In an effort to reduce computational complexity, we introduce an attention scheme inspired by inducing point methods from sparse Gaussian process literature. It reduces the computation time of self-attention from quadratic to linear in the number of elements in the set. We show that our model is theoretically attractive and we evaluate it on a range of tasks, demonstrating the state-of-the-art performance compared to recent methods for set-structured data.
Juho Lee 0001, Yoonho Lee 0001, Jungtaek Kim 0001, Adam R. Kosiorek, Seungjin Choi 0001, Yee Whye Teh
ICML3
2018 On the Optimal Bit Complexity of Circulant Binary Embedding
Saehoon Kim, Jungtaek Kim 0001, Seungjin Choi 0001
AAAI2
2018 Open Set Recognition by Regularising Classifier with Fake Data Generated by Generative Adversarial Networks
abstract
We present a new method to generate fake data in unknown classes in generative adversarial networks (GANs) framework. The generator in GANs is trained to generate somewhat similar to data in known classes but the different one by modelling noisy distribution on feature space of a classifier using proposed marginal denoising autoencoder. The generated data are treated as fake instances in unknown classes and given to the classifier to make it be robust to the real unknown classes. Our results show that synthetic data can act as fake unknown classes and keep down the certainty of the classifier on real unknown classes meanwhile the classification capability of known classes is not degenerated, even improved.
Inhyuk Jo, Jungtaek Kim 0001, Hyohyeong Kang, Yong-Deok Kim, Seungjin Choi 0001
ICASSP2
2018 Clustering-Guided Gp-Ucb for Bayesian Optimization
abstract
Bayesian optimization is a powerful technique for finding extrema of an objective function, a closed-form expression of which is not given but expensive evaluations at query points are available. Gaussian Process (GP) regression is often used to estimate the objective function and uncertainty estimates that guide GP-Upper Confidence Bound (GP-UCB) to determine where next to sample from the objective function, balancing exploration and exploitation. In general, it requires an auxiliary optimization to tune the hyperparameter in GP-UCB, which is sometimes not easy to carry out in practice. In this paper we present a simple practical method which improves GP-UCB, especially in cases where the objective function is not smooth with sharp peaks and valleys. We first present a geometric interpretation of GP-UCB on which we base our development of the clustering-guided method to select the next observation. Clustering is applied to two-dimensional vectors whose entries correspond to the posterior mean and standard deviation computed by GP regression, which is followed by utility maximization with GP-UCB, in order to determine where next to sample from the objective function. Experiments on various functions demonstrate our method alleviates the chance of being trapped in local extrema, making small efforts for auxiliary optimization.
Jungtaek Kim 0001, Seungjin Choi 0001
ICASSP1