Wenpin Tang

dblp:240/4543 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-7228-1954ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 34% Reinforcement learning · 32% Probabilistic and Bayesian machine learning · 21%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
preference optimization
1.722025
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025
MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.722025
Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025
MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025
Machine learning › Reinforcement learning
continuous-time reinforcement learning
1.522025
Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025
Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
1.022023
Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023
Mallows ranking models: maximum likelihood estimate and regeneration · ICML 2019
Natural language and speech › Language models and text generation
alignment
0.912025
MallowsPO: Fine-Tune Your LLM with Preference Dispersions · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning
0.912025
Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning · ICML 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
covariance function
0.712023
Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.712023
Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning
policy optimization
0.712023
Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023
Graph algorithms and graph theory
random graph models
0.412020
The Buckley-Osthus model and the block preferential attachment model: statistical analysis and application · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
ranking model
0.412019
Mallows ranking models: maximum likelihood estimate and regeneration · ICML 2019
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.212023
Policy Optimization for Continuous Reinforcement Learning · NeurIPS 2023
Computational social science and digital humanities › network science
network evolution
0.112020
The Buckley-Osthus model and the block preferential attachment model: statistical analysis and application · ICML 2020

Methods — techniques the papers use, named apart from their topics

direct preference optimization · 1.7unified objective · 0.9stochastic control · 0.9score matching · 0.9policy optimization · 0.9mallows ranking model · 0.9simulation study · 0.9maximum likelihood estimation · 0.9stochastic differential equation · 0.7policy gradient · 0.7occupation time · 0.7best linear unbiased predictor · 0.7
YearPublicationVenuePosition
2026 Contractive Diffusion Probabilistic Models
abstract
Abstract. Diffusion probabilistic models (DPMs) have emerged as a promising technique in generative modeling. The success of DPMs relies on two ingredients: time reversal of diffusion processes and score matching. In view of possibly unguaranteed score matching, we propose a new criterion—the contraction property of backward sampling in the design of DPMs, leading to a novel class of contractive DPMs (CDPMs). Our key insight is that the contraction property can provably narrow score-matching errors and discretization errors; thus our proposed CDPMs are robust to both sources of error. For practical use, we showcase that CDPM can leverage weights of pretrained DPMs by a simple transformation, without the necessity of further training. We corroborated our approach by experiments on Swiss Roll, MNIST, CIFAR-10 32[Formula: see text]32, and AFHQ 64[Formula: see text]64 dataset. Notably, CDPM steadily improves the performance of baseline score-based diffusion models.
Wenpin Tang, Hanyang Zhao
SIAM J. Imaging Sci.1
2025 MallowsPO: Fine-Tune Your LLM with Preference Dispersions
abstract
Direct Preference Optimization (DPO) has recently emerged as a popular approach to improve reinforcement learning from human feedback (RLHF), leading to better techniques to fine-tune large language models (LLM). A weakness of DPO, however, lies in its lack of capability to characterize the diversity of human preferences. Inspired by Mallows' theory of preference ranking, we develop in this paper a new approach, the *MallowsPO*. A distinct feature of this approach is a *dispersion index*, which reflects the dispersion of human preference to prompts. We show that existing DPO models can be reduced to special cases of this dispersion index, thus unified with MallowsPO. More importantly, we demonstrate empirically how to use this dispersion index to enhance the performance of DPO in a broad array of benchmark tasks, from synthetic bandit selection to controllable generation and dialogues, while maintaining great generalization capabilities. MallowsPO is also compatible with other SOTA offline preference optimization methods, boosting nearly 2\% extra LC win rate when used as a plugin for fine-tuning Llama3-Instruct.
Haoxian Chen 0002, Hanyang Zhao, Henry Lam, David D. Yao, Wenpin Tang
ICLR5
2025 RainbowPO: A Unified Framework for Combining Improvements in Preference Optimization
abstract
Recently, numerous preference optimization algorithms have been introduced as extensions to the Direct Preference Optimization (DPO) family. While these methods have successfully aligned models with human preferences, there is a lack of understanding regarding the contributions of their additional components. Moreover, fair and consistent comparisons are scarce, making it difficult to discern which components genuinely enhance downstream performance. In this work, we propose RainbowPO, a unified framework that demystifies the effectiveness of existing DPO methods by categorizing their key components into seven broad directions. We integrate these components into a single cohesive objective, enhancing the performance of each individual element. Through extensive experiments, we demonstrate that RainbowPO outperforms existing DPO variants. Additionally, we provide insights to guide researchers in developing new DPO methods and assist practitioners in their implementations.
Hanyang Zhao, Genta Indra Winata, David D. Yao, Wenpin Tang, Sambit Sahu
ICLR6
2025 Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
abstract
Reinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models. Most works in this area uses a discrete-time formulation, which is prone to induced errors, and often not applicable to models with higher-order/black-box solvers. The objective of this study is to develop a disciplined approach to fine-tuning diffusion models using continuous-time RL, formulated as a stochastic control problem with a reward function that aligns the end result (terminal state) with input prompt. The key idea is to treat score matching as controls or actions, and thereby connecting to policy optimization and regularization in continuous-time RL. To carry out this idea, we lay out a new policy optimization framework for continuous-time RL, and illustrate its potential in enhancing the value networks design space via leveraging the structural property of diffusion models. We validate the advantages of our method by experiments in downstream tasks of fine-tuning large-scale Text2Image models, Stable Diffusion v1.5.
Hanyang Zhao, Haoxian Chen 0002, David D. Yao, Wenpin Tang
ICML5
2025 Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
abstract
Preference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is organized into three main sections: 1) introduction and preliminaries: an introduction to reinforcement learning frameworks, preference tuning tasks, models, and datasets across various modalities: language, speech, and vision, as well as different policy approaches, 2) in-depth exploration of each preference tuning approach: a detailed analysis of the methods used in preference tuning, and 3) applications, discussion, and future directions: an exploration of the applications of preference tuning in downstream tasks, including evaluation methods for different modalities, and an outlook on future research directions. Our objective is to present the latest methodologies in preference tuning and model alignment, enhancing the understanding of this field for researchers and practitioners. We hope to encourage further engagement and innovation in this area. Additionally, we provide a GitHub link https://github.com/hanyang1999/Preference-Tuning-with-Human-Feedback.
Genta Indra Winata, Hanyang Zhao, Wenpin Tang, David D. Yao, Sambit Sahu
J. Artif. Intell. Res.4
2023 Policy Optimization for Continuous Reinforcement Learning
abstract
We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance difference and local approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach.
Hanyang Zhao, Wenpin Tang, David D. Yao
NeurIPS2
2023 Inference for Gaussian Processes with Matern Covariogram on Compact Riemannian Manifolds
abstract
Gaussian processes are widely employed as versatile modelling and predictive tools in spatial statistics, functional data analysis, computer modelling and diverse applications of machine learning. They have been widely studied over Euclidean spaces, where they are specified using covariance functions or covariograms for modelling complex dependencies. There is a growing literature on Gaussian processes over Riemannian manifolds in order to develop richer and more flexible inferential frameworks for non-Euclidean data. While numerical approximations through graph representations have been well studied for the Matern covariogram and heat kernel, the behaviour of asymptotic inference on the parameters of the covariogram has received relatively scant attention. We focus on asymptotic behaviour for Gaussian processes constructed over compact Riemannian manifolds. Building upon a recently introduced Matern covariogram on a compact Riemannian manifold, we employ formal notions and conditions for the equivalence of two Matern Gaussian random measures on compact manifolds to derive the parameter that is identifiable, also known as the microergodic parameter, and formally establish the consistency of the maximum likelihood estimate and the asymptotic optimality of the best linear unbiased predictor. The circle is studied as a specific example of compact Riemannian manifolds with numerical experiments to illustrate and corroborate the theory.
Didong Li, Wenpin Tang, Sudipto Banerjee
J. Mach. Learn. Res.2
2022 Asset selection via correlation blockmodel clustering
Wenpin Tang, Xun Yu Zhou
Expert Syst. Appl.1
2020 The Buckley-Osthus model and the block preferential attachment model: statistical analysis and application
abstract
This paper is concerned with statistical estimation of two preferential attachment models: the Buckley-Osthus model and the block preferential attachment model. We prove that the maximum likelihood estimates for both models are consistent. We perform simulation studies to corroborate our theoretical findings. We also apply both models to study the evolution of a real-world network. A list of open problems are presented.
Wenpin Tang, Fengmin Tang
ICML1
2019 Mallows ranking models: maximum likelihood estimate and regeneration
abstract
This paper is concerned with various Mallows ranking models. We study the statistical properties of the MLE of Mallows’ $\phi$ model. We also make connections of various Mallows ranking models, encompassing recent progress in mathematics. Motivated by the infinite top-$t$ ranking model, we propose an algorithm to select the model size $t$ automatically. The key idea relies on the renewal property of such an infinite random permutation. Our algorithm shows good performance on several data sets.
Wenpin Tang
ICML1