VLDB 2026 Research / reviewers in the wild / expert
Ji Won Park
dblp:83/10554
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Optimization for machine learning · 54% Probabilistic and Bayesian machine learning · 34% Trustworthy machine learning · 12% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.4 | 2 | 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranks · ICML 2024 GAUCHE: A Library for Gaussian Processes in Chemistry · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
prediction-powered inference |
0.9 | 1 | 2025 | Reliable Algorithm Selection for Machine Learning-Guided Design · ICML 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function |
0.8 | 1 | 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranks · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.8 | 1 | 2024 | Chain of Log-Concave Markov Chains · ICLR 2024 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
multi-objective bayesian optimization |
0.8 | 1 | 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranks · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
unnormalized density sampling |
0.8 | 1 | 2024 | Chain of Log-Concave Markov Chains · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.7 | 1 | 2023 | GAUCHE: A Library for Gaussian Processes in Chemistry · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning
copula models |
0.2 | 1 | 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranks · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
density ratio estimation · 0.9wasserstein metric · 0.8multivariate rank · 0.8langevin dynamics · 0.8empirical bayes · 0.8copula · 0.8CDF indicator · 0.8gaussian process · 0.7bayesian optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semiparametric conformal predictionabstractMany risk-sensitive applications require well-calibrated prediction sets over multiple, potentially correlated target variables, for which the prediction algorithm may report correlated errors. In this work, we aim to construct the conformal prediction set accounting for the joint correlation structure of the vector-valued non-conformity scores. Drawing from the rich literature on multivariate quantiles and semiparametric statistics, we propose an algorithm to estimate the $1-\alpha$ quantile of the scores, where $\alpha$ is the user-specified miscoverage rate. In particular, we flexibly estimate the joint cumulative distribution function (CDF) of the scores using nonparametric vine copulas and improve the asymptotic efficiency of the quantile estimate using its influence function. The vine decomposition allows our method to scale well to a large number of targets. As well as guaranteeing asymptotically exact coverage, our method yields desired coverage and competitive efficiency on a range of real-world regression problems, including those with missing-at-random labels in the calibration set. Ji Won Park, Kyunghyun Cho |
AISTATS | 1 |
| 2025 | Reliable Algorithm Selection for Machine Learning-Guided DesignabstractAlgorithms for machine learning-guided design, or design algorithms, use machine learning-based predictions to propose novel objects with desired property values. Given a new design task—for example, to design novel proteins with high binding affinity to a therapeutic target—one must choose a design algorithm and specify any hyperparameters and predictive and/or generative models involved. How can these decisions be made such that the resulting designs are successful? This paper proposes a method for design algorithm selection, which aims to select design algorithms that will produce a distribution of design labels satisfying a user-specified success criterion—for example, that at least ten percent of designs’ labels exceed a threshold. It does so by combining designs’ predicted property values with held-out labeled data to reliably forecast characteristics of the label distributions produced by different design algorithms, building upon techniques from prediction-powered inference (Angelopoulos et al., 2023). The method is guaranteed with high probability to return design algorithms that yield successful label distributions (or the null set if none exist), if the density ratios between the design and labeled data distributions are known. We demonstrate the method’s effectiveness in simulated protein and RNA design tasks, in settings with either known or estimated density ratios. Clara Fannjiang, Ji Won Park |
ICML | 2 |
| 2025 | Efficient semantic uncertainty quantification in language models via diversity-steered samplingabstractAccurately estimating *semantic* aleatoric and epistemic uncertainties in large language models (LLMs) is particularly challenging in free-form question answering (QA), where obtaining stable estimates often requires many expensive generations. We introduce a **diversity-steered sampler** that discourages semantically redundant outputs during decoding, covers both autoregressive and masked diffusion paradigms, and yields substantial sample-efficiency gains. The key idea is to inject a continuous semantic-similarity penalty into the model’s proposal distribution using a natural language inference (NLI) model lightly finetuned on partial prefixes or intermediate diffusion states. We debias downstream uncertainty estimates with importance reweighting and shrink their variance with control variates. Across four QA benchmarks, our method matches or surpasses baselines while covering more semantic clusters with the same number of samples. Being modular and requiring no gradient access to the base LLM, the framework promises to serve as a drop-in enhancement for uncertainty estimation in risk-sensitive model deployments. Ji Won Park, Kyunghyun Cho |
NeurIPS | 1 |
| 2024 | Chain of Log-Concave Markov ChainsabstractWe introduce a theoretical framework for sampling from unnormalized densities based on a smoothing scheme that uses an isotropic Gaussian kernel with a single fixed noise scale. We prove one can decompose sampling from a density (minimal assumptions made on the density) into a sequence of sampling from log-concave conditional densities via accumulation of noisy measurements with equal noise levels. Our construction is unique in that it keeps track of a history of samples, making it non-Markovian as a whole, but it is lightweight algorithmically as the history only shows up in the form of a running empirical mean of samples. Our sampling algorithm generalizes walk-jump sampling (Saremi & Hyvärinen, 2019). The "walk" phase becomes a (non-Markovian) chain of (log-concave) Markov chains. The "jump" from the accumulated measurements is obtained by empirical Bayes. We study our sampling algorithm quantitatively using the 2-Wasserstein metric and compare it with various Langevin MCMC algorithms. We also report a remarkable capacity of our algorithm to "tunnel" between modes of a distribution. Saeed Saremi, Ji Won Park, Francis R. Bach |
ICLR | 2 |
| 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranksabstractMany scientific and industrial applications require the joint optimization of multiple, potentially competing objectives. Multi-objective Bayesian optimization (MOBO) is a sample-efficient framework for identifying Pareto-optimal solutions. At the heart of MOBO is the acquisition function, which determines the next candidate to evaluate by navigating the best compromises among the objectives. Acquisition functions that rely on integrating over the objective space scale poorly to a large number of objectives. In this paper, we show a natural connection between the non-dominated solutions and the highest multivariate rank, which coincides with the extreme level line of the joint cumulative distribution function (CDF). Motivated by this link, we propose the CDF indicator, a Pareto-compliant metric for evaluating the quality of approximate Pareto sets, that can complement the popular hypervolume indicator. We then introduce an acquisition function based on the CDF indicator, called BOtied. BOtied can be implemented efficiently with copulas, a statistical tool for modeling complex, high-dimensional distributions. Our experiments on a variety of synthetic and real-world experiments demonstrate that BOtied outperforms state-of-the-art MOBO algorithms while being computationally efficient for many objectives. Ji Won Park, Natasa Tagasovska, Michael Maser, Stephen Ra, Kyunghyun Cho |
ICML | 1 |
| 2023 | GAUCHE: A Library for Gaussian Processes in ChemistryabstractWe introduce GAUCHE, an open-source library for GAUssian processes in CHEmistry. Gaussian processes have long been a cornerstone of probabilistic machine learning, affording particular advantages for uncertainty quantification and Bayesian optimisation. Extending Gaussian processes to molecular representations, however, necessitates kernels defined over structured inputs such as graphs, strings and bit vectors. By providing such kernels in a modular, robust and easy-to-use framework, we seek to enable expert chemists and materials scientists to make use of state-of-the-art black-box optimization techniques. Motivated by scenarios frequently encountered in practice, we showcase applications for GAUCHE in molecular discovery, chemical reaction optimisation and protein design. The codebase is made available at https://github.com/leojklarner/gauche. Ryan-Rhys Griffiths, Leo Klarner, Henry B. Moss, Aditya Ravuri, Sang Truong, Yuanqi Du, Samuel Stanton, Gary Tom, Bojana Rankovic, Arian Rokkum Jamasb, Aryan Deshwal, Julius Schwartz, Austin Tripp, Gregory Kell, Simon Frieder, Anthony Bourached, Alex Chan, Jacob Moss, Chengzhi Guo, Johannes Peter Dürholt, Saudamini Chaurasia, Ji Won Park, Felix Strieth-Kalthoff, Alpha A. Lee, Bingqing Cheng, Alán Aspuru-Guzik, Philippe Schwaller, Jian Tang 0005 |
NeurIPS | 22 |