VLDB 2026 Research / reviewers in the wild / expert
Vincent Zhuang
dblp:200/8345
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0007-2931-3069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 29% Language models and text generation · 21% Probabilistic and Bayesian machine learning · 18% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.1 | 2 | 2024 | Scalable Bayesian Optimization via Focalized Sparse Gaussian Processes · NeurIPS 2024 Stagewise Safe Bayesian Optimization with Gaussian Processes · ICML 2018 |
Natural language and speech › Language models and text generation › large language model inference
inference-time computation |
0.9 | 1 | 2025 | Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models · ICLR 2025 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.9 | 1 | 2025 | Training Language Models to Self-Correct via Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.9 | 1 | 2025 | Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
self-correction |
0.9 | 1 | 2025 | Training Language Models to Self-Correct via Reinforcement Learning · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.8 | 1 | 2024 | Scalable Bayesian Optimization via Focalized Sparse Gaussian Processes · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse gaussian process |
0.8 | 1 | 2024 | Scalable Bayesian Optimization via Focalized Sparse Gaussian Processes · NeurIPS 2024 |
Query processing and optimization
cardinality estimation |
0.7 | 1 | 2023 | Kepler: Robust Learning for Parametric Query Optimization · Proc. ACM Manag. Data 2023 |
Query processing and optimization › query optimization
learned query optimization |
0.7 | 1 | 2023 | Kepler: Robust Learning for Parametric Query Optimization · Proc. ACM Manag. Data 2023 |
Query processing and optimization › adaptive query processing › adaptive query optimization
parametric query optimization |
0.7 | 1 | 2023 | Kepler: Robust Learning for Parametric Query Optimization · Proc. ACM Manag. Data 2023 |
Query processing and optimization › query planning
query plan selection |
0.7 | 1 | 2023 | Kepler: Robust Learning for Parametric Query Optimization · Proc. ACM Manag. Data 2023 |
Machine learning › Reinforcement learning
exploration |
0.3 | 1 | 2018 | Stagewise Safe Bayesian Optimization with Gaussian Processes · ICML 2018 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
safe bayesian optimization |
0.3 | 1 | 2018 | Stagewise Safe Bayesian Optimization with Gaussian Processes · ICML 2018 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration |
0.3 | 1 | 2018 | Stagewise Safe Bayesian Optimization with Gaussian Processes · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7supervised fine-tuning · 0.9sampling-based planning · 0.9morphology-aware proportional control · 0.9model predictive control · 0.9imitation learning · 0.9black-box optimization · 0.9best-of-n sampling · 0.9variational inference · 0.8acquisition function optimization · 0.8row count evolution · 0.7neural network uncertainty estimation · 0.7gaussian process · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language ModelsabstractRecent studies indicate that effectively utilizing inference-time compute is crucial for attaining good performance from large language models (LLMs). Specifically, the Best-of-N (BoN) inference strategy, where an LLM generates multiple responses and a verifier selects the best, has shown strong empirical performance. Motivated by this, we develop a novel inference-aware fine-tuning paradigm, which encompasses the BoN-aware inference framework as a special case. We devise the first imitation learning and reinforcement learning (RL) methods for fine-tuning LLMs using BoN, overcoming the challenging, non-differentiable argmax operator in BoN. We empirically demonstrate that our BoN-aware models implicitly learn a per-example "meta-strategy", which interleaves best responses with more diverse responses that might be better suited to a test-time input—a process reminiscent of the exploration-exploitation trade-off in RL. Our experiments demonstrate the effectiveness of BoN-aware fine-tuning in terms of improved performance and inference-time compute. In particular, we show that our methods improve the BoN performance of Gemma 2B on Hendrycks MATH from 26.8% to 30.8%, and Pass@K from 60% to 67%. Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai 0001, Aviral Kumar, Rishabh Agarwal, Sridhar Thiagarajan, Craig Boutilier, Aleksandra Faust |
ICLR | 4 |
| 2025 | Training Language Models to Self-Correct via Reinforcement LearningabstractSelf-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs. Current methods for training self-correction typically depend on either multiple models, a more advanced model, or additional forms of supervision. To address these shortcomings, we develop a multi-turn online reinforcement learning (RL) approach, SCoRe, that significantly improves an LLM's self-correction ability using entirely self-generated data. To build SCoRe, we first show that variants of supervised fine-tuning (SFT) on offline model-generated correction traces are insufficient for instilling self-correction behavior. In particular, we observe that training via SFT either suffers from a distribution mismatch between the training data and the model's own responses or implicitly prefers only a certain mode of correction behavior that is often not effective at test time. SCoRe addresses these challenges by training under the model's own distribution of self-generated correction traces and using appropriate regularization to steer the learning process into learning a self-correction strategy that is effective at test time as opposed to simply fitting high-reward responses for a given prompt. This regularization prescribes running a first phase of RL on a base model to generate a policy initialization that is less susceptible to collapse and then using a reward bonus to amplify self-correction during training. When applied to Gemini 1.0 Pro and 1.5 Flash models, we find that SCoRe achieves state-of-the-art self-correction performance, improving the base models' self-correction by 15.6% and 9.1% respectively on the MATH and HumanEval benchmarks. Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, Aleksandra Faust |
ICLR | 2 |
| 2025 | Motion Control of High-Dimensional Musculoskeletal Systems with Hierarchical Model-Based PlanningabstractControlling high-dimensional nonlinear systems, such as those found in biological and robotic applications, is challenging due to large state and action spaces. While deep reinforcement learning has achieved a number of successes in these domains, it is computationally intensive and time consuming, and therefore not suitable for solving large collections of tasks that require significant manual tuning. In this work, we introduce Model Predictive Control with Morphology-aware Proportional Control (MPC$^2$), a hierarchical model-based learning algorithm for zero-shot and near-real-time control of high-dimensional complex dynamical systems. MPC$^2$ uses a sampling-based model predictive controller for target posture planning, and enables robust control for high-dimensional tasks by incorporating a morphology-aware proportional controller for actuator coordination. The algorithm enables motion control of a high-dimensional human musculoskeletal model in a variety of motion tasks, such as standing, walking on different terrains, and imitating sports activities. The reward function of MPC$^2$ can be tuned via black-box optimization, drastically reducing the need for human-intensive reward engineering. Yunyue Wei, Shanning Zhuang, Vincent Zhuang, Yanan Sui |
ICLR | 3 |
| 2024 | The Design of the Barkour Benchmark for Robot AgilityabstractIn this paper, we describe the design of the Barkour benchmark for measuring robot agility in navigating complex environments. Despite the growing interest in developing agile robot locomotion skills, the field lacks systematic benchmarks to measure the performance of robotic control systems and hardware in agility-focused tasks. This motivated us to propose the Barkour benchmark, an obstacle course designed to quantify agility across various robotic platforms. Inspired by dog agility competitions, the course features diverse obstacles and a time-based scoring mechanism, encouraging researchers to develop controllers that enable robots to move quickly, precisely, and with adaptability. This benchmark is challenging as it demands diverse motion skills and the time-based scoring requires control precision at high speed. Along with the design details presented in the paper, we release our simulated environment setups in MuJoCo-XLA and the CAD model of a custom-designed quadruped robot to facilitate future research to reproduce the Barkour setup (available at sites.google.com/view/barkour). We hope these together will accelerate the pace of robot agility research. Wenhao Yu 0003, Ken Caluwaerts, Atil Iscen, J. Chase Kew, Tingnan Zhang, Daniel Freeman, Lisa Lee, Stefano Saliceti, Vincent Zhuang, Nathan Batchelor, Steven Bohez, Federico Casarini, José Enrique Chen, Erwin Coumans, Adil Dostmohamed, Gabriel Dulac-Arnold, Alejandro Escontrela, Erik Frey, Roland Hafner, Deepali Jain, Bauyrjan Jyenis, Yuheng Kuang, Tsang-Wei Edward Lee, Ofir Nachum, Kenneth Oslund, Francesco Romano, Fereshteh Sadeghi, Baruch Tabanpour, Daniel Zheng, Michael Neunert, Raia Hadsell, Nicolas Heess, Francesco Nori, Jeff Seto, Carolina Parada, Vikas Sindhwani, Vincent Vanhoucke, Jie Tan 0001, Kuang-Huei Lee |
IROS | 9 |
| 2024 | Scalable Bayesian Optimization via Focalized Sparse Gaussian ProcessesabstractBayesian optimization is an effective technique for black-box optimization, but its applicability is typically limited to low-dimensional and small-budget problems due to the cubic complexity of computing the Gaussian process (GP) surrogate. While various approximate GP models have been employed to scale Bayesian optimization to larger sample sizes, most suffer from overly-smooth estimation and focus primarily on problems that allow for large online samples. In this work, we argue that Bayesian optimization algorithms with sparse GPs can more efficiently allocate their representational power to relevant regions of the search space. To achieve this, we propose focalized GP, which leverages a novel variational loss function to achieve stronger local prediction, as well as FocalBO, which hierarchically optimizes the focalized GP acquisition function over progressively smaller search spaces. Experimental results demonstrate that FocalBO can efficiently leverage large amounts of offline and online data to achieve state-of-the-art performance on robot morphology design and to control a 585-dimensional musculoskeletal system. Yunyue Wei, Vincent Zhuang, Saraswati Soedarmadji, Yanan Sui |
NeurIPS | 2 |
| 2023 | Kepler: Robust Learning for Parametric Query OptimizationabstractMost existing parametric query optimization (PQO) techniques rely on traditional query optimizer cost models, which are often inaccurate and result in suboptimal query performance. We propose Kepler, an end-to-end learning-based approach to PQO that demonstrates significant speedups in query latency over a traditional query optimizer. Central to our method is Row Count Evolution (RCE), a novel plan generation algorithm based on perturbations in the sub-plan cardinality space. While previous approaches require accurate cost models, we bypass this requirement by evaluating candidate plans via actual execution data and training anML model to predict the fastest plan given parameter binding values. Our models leverage recent advances in neural network uncertainty in order to robustly predict faster plans while avoiding regressions in query performance. Experimentally, we show that Kepler achieves significant improvements in query runtime on multiple datasets on PostgreSQL. Lyric Doshi, Vincent Zhuang, Gaurav Jain, Ryan Marcus, Deniz Altinbüken, Eugene Brevdo, Campbell Fraser |
Proc. ACM Manag. Data | 2 |
| 2021 | No-Regret Reinforcement Learning with Heavy-Tailed RewardsabstractReinforcement learning algorithms typically assume rewards to be sampled from light-tailed distributions, such as Gaussian or bounded. However, a wide variety of real-world systems generate rewards that follow heavy-tailed distributions. We consider such scenarios in the setting of undiscounted reinforcement learning. By constructing a lower bound, we show that the difficulty of learning heavy-tailed rewards asymptotically dominates the difficulty of learning transition probabilities. Leveraging techniques from robust mean estimation, we propose Heavy-UCRL2 and Heavy-Q-Learning, and show that they achieve near-optimal regret bounds in this setting. Our algorithms also naturally generalize to deep reinforcement learning applications; we instantiate Heavy-DQN as an example of this. We demonstrate that all of our algorithms outperform baselines on both synthetic MDPs and standard RL benchmarks. Vincent Zhuang, Yanan Sui |
AISTATS | 1 |
| 2018 | Stagewise Safe Bayesian Optimization with Gaussian ProcessesabstractEnforcing safety is a key aspect of many problems pertaining to sequential decision making under uncertainty, which require the decisions made at every step to be both informative of the optimal decision and also safe. For example, we value both efficacy and comfort in medical therapy, and efficiency and safety in robotic control. We consider this problem of optimizing an unknown utility function with absolute feedback or preference feedback subject to unknown safety constraints. We develop an efficient safe Bayesian optimization algorithm, StageOpt, that separates safe region expansion and utility function maximization into two distinct stages. Compared to existing approaches which interleave between expansion and optimization, we show that StageOpt is more efficient and naturally applicable to a broader class of problems. We provide theoretical guarantees for both the satisfaction of safety constraints as well as convergence to the optimal utility value. We evaluate StageOpt on both a variety of synthetic experiments, as well as in clinical practice. We demonstrate that StageOpt is more effective than existing safe optimization approaches, and is able to safely and effectively optimize spinal cord stimulation therapy in our clinical experiments. Yanan Sui, Vincent Zhuang, Joel W. Burdick, Yisong Yue |
ICML | 2 |
| 2017 | Multi-dueling Bandits with Dependent Arms
Yanan Sui, Vincent Zhuang, Joel W. Burdick, Yisong Yue |
UAI | 2 |