VLDB 2026 Research / reviewers in the wild / expert
Wenpin Tang
dblp:240/4543
· DBLP profile ↗
10ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-7228-1954ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 34% Reinforcement learning · 32% Probabilistic and Bayesian machine learning · 21% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
preference optimization |
1.7 | 2 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.7 | 2 | 2025 | Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025 MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025 |
Machine learning › Reinforcement learning
continuous-time reinforcement learning |
1.5 | 2 | 2025 | Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025 Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
1.0 | 2 | 2023 | Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023 Mallows ranking models: maximum likelihood estimate and regeneration · ICML 2019 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning |
0.9 | 1 | 2025 | Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
covariance function |
0.7 | 1 | 2023 | Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.7 | 1 | 2023 | Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023 |
Machine learning › Reinforcement learning
policy optimization |
0.7 | 1 | 2023 | Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023 |
Graph algorithms and graph theory
random graph models |
0.4 | 1 | 2020 | The Buckley-Osthus model and the block preferential attachment model: statistical analysis and application · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
ranking model |
0.4 | 1 | 2019 | Mallows ranking models: maximum likelihood estimate and regeneration · ICML 2019 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.2 | 1 | 2023 | Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023 |
Computational social science and digital humanities › network science
network evolution |
0.1 | 1 | 2020 | The Buckley-Osthus model and the block preferential attachment model: statistical analysis and application · ICML 2020 |
Methods — techniques the papers use, named apart from their topics
direct preference optimization · 1.7unified objective · 0.9stochastic control · 0.9score matching · 0.9policy optimization · 0.9mallows ranking model · 0.9simulation study · 0.9maximum likelihood estimation · 0.9stochastic differential equation · 0.7policy gradient · 0.7occupation time · 0.7best linear unbiased predictor · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contractive Diffusion Probabilistic ModelsabstractAbstract. Diffusion probabilistic models (DPMs) have emerged as a promising technique in generative modeling. The success of DPMs relies on two ingredients: time reversal of diffusion processes and score matching. In view of possibly unguaranteed score matching, we propose a new criterion—the contraction property of backward sampling in the design of DPMs, leading to a novel class of contractive DPMs (CDPMs). Our key insight is that the contraction property can provably narrow score-matching errors and discretization errors; thus our proposed CDPMs are robust to both sources of error. For practical use, we showcase that CDPM can leverage weights of pretrained DPMs by a simple transformation, without the necessity of further training. We corroborated our approach by experiments on Swiss Roll, MNIST, CIFAR-10 32[Formula: see text]32, and AFHQ 64[Formula: see text]64 dataset. Notably, CDPM steadily improves the performance of baseline score-based diffusion models. Wenpin Tang, Hanyang Zhao |
SIAM J. Imaging Sci. | 1 |
| 2025 | MallowsPO: Fine-Tune Your LLM with Preference DispersionsabstractDirect Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning from human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakness of DPO, however, lies in its lack of capability to characterize the diversity of human preferences. Inspired by Mallows' theory of preference ranking, we develop in this paper a new approach, the *MallowsPO*. A distinct feature of this approach is a *dispersion index*, which reflects the dispersion of human preference to prompts. We show that existing DPO models can be reduced to special cases of this dispersion index, thus unified with MallowsPO. More importantly, we demonstrate empirically how to use this dispersion index to enhance the performance of DPO in a broad array of benchmark tasks, from synthetic bandit selection to controllable generation and dialogues, while maintaining great generalization capabilities. MallowsPO is also compatible with other SOTA offline preference optimization methods, boosting nearly 2\% extra LC win rate when used as a plugin for fine-tuning Llama3-Instruct. Haoxian Chen 0002, Hanyang Zhao, Henry Lam, David D. Yao, Wenpin Tang |
ICLR | 5 |
| 2025 | RainbowPO: A Unified Framework for Combining Improvements in Preference OptimizationabstractRecently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations. Hanyang Zhao, Genta Indra Winata, David D. Yao, Wenpin Tang, Sambit Sahu |
ICLR | 6 |
| 2025 | Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement LearningabstractReinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models. Most works in this area uses a discrete-time formulation, which is prone to induced errors, and often not applicable to models with higher-order/black-box solvers. The objective of this study is to develop a disciplined approach to fine-tuning diffusion models using continuous-time RL, formulated as a stochastic control problem with a reward function that aligns the end result (terminal state) with input prompt. The key idea is to treat score matching as controls or actions, and thereby connecting to policy optimization and regularization in continuous-time RL. To carry out this idea, we lay out a new policy optimization framework for continuous-time RL, and illustrate its potential in enhancing the value networks design space via leveraging the structural property of diffusion models. We validate the advantages of our method by experiments in downstream tasks of fine-tuning large-scale Text2Image models, Stable Diffusion v1.5. Hanyang Zhao, Haoxian Chen 0002, David D. Yao, Wenpin Tang |
ICML | 5 |
| 2025 | Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A SurveyabstractPreference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is organized into three main sections: 1) introduction and preliminaries: an introduction to reinforcement learning frameworks, preference tuning tasks, models, and datasets across various modalities: language, speech, and vision, as well as different policy approaches, 2) in-depth exploration of each preference tuning approach: a detailed analysis of the methods used in preference tuning, and 3) applications, discussion, and future directions: an exploration of the applications of preference tuning in downstream tasks, including evaluation methods for different modalities, and an outlook on future research directions. Our objective is to present the latest methodologies in preference tuning and model alignment, enhancing the understanding of this field for researchers and practitioners. We hope to encourage further engagement and innovation in this area. Additionally, we provide a GitHub link https://github.com/hanyang1999/Preference-Tuning-with-Human-Feedback. Genta Indra Winata, Hanyang Zhao, Wenpin Tang, David D. Yao, Sambit Sahu |
J. Artif. Intell. Res. | 4 |
| 2023 | Policy Optimization for Continuous Reinforcement LearningabstractWe study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance difference and local approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach. Hanyang Zhao, Wenpin Tang, David D. Yao |
NeurIPS | 2 |
| 2023 | Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian ManifoldsabstractGaussian processes are widely employed as versatile modelling and predictive tools in spatial statistics, functional data analysis, computer modelling and diverse applications of machine learning. They have been widely studied over Euclidean spaces, where they are specified using covariance functions or covariograms for modelling complex dependencies. There is a growing literature on Gaussian processes over Riemannian manifolds in order to develop richer and more flexible inferential frameworks for non-Euclidean data. While numerical approximations through graph representations have been well studied for the Matern covariogram and heat kernel, the behaviour of asymptotic inference on the parameters of the covariogram has received relatively scant attention. We focus on asymptotic behaviour for Gaussian processes constructed over compact Riemannian manifolds. Building upon a recently introduced Matern covariogram on a compact Riemannian manifold, we employ formal notions and conditions for the equivalence of two Matern Gaussian random measures on compact manifolds to derive the parameter that is identifiable, also known as the microergodic parameter, and formally establish the consistency of the maximum likelihood estimate and the asymptotic optimality of the best linear unbiased predictor. The circle is studied as a specific example of compact Riemannian manifolds with numerical experiments to illustrate and corroborate the theory. Didong Li, Wenpin Tang, Sudipto Banerjee |
J. Mach. Learn. Res. | 2 |
| 2022 | Asset selection via correlation blockmodel clustering
Wenpin Tang, Xun Yu Zhou |
Expert Syst. Appl. | 1 |
| 2020 | The Buckley-Osthus model and the block preferential attachment model: statistical analysis and applicationabstractThis paper is concerned with statistical estimation of two preferential attachment models: the Buckley-Osthus model and the block preferential attachment model. We prove that the maximum likelihood estimates for both models are consistent. We perform simulation studies to corroborate our theoretical findings. We also apply both models to study the evolution of a real-world network. A list of open problems are presented. Wenpin Tang, Fengmin Tang |
ICML | 1 |
| 2019 | Mallows ranking models: maximum likelihood estimate and regenerationabstractThis paper is concerned with various Mallows ranking models. We study the statistical properties of the MLE of Mallows’ $\phi$ model. We also make connections of various Mallows ranking models, encompassing recent progress in mathematics. Motivated by the infinite top-$t$ ranking model, we propose an algorithm to select the model size $t$ automatically. The key idea relies on the renewal property of such an infinite random permutation. Our algorithm shows good performance on several data sets. Wenpin Tang |
ICML | 1 |