Wenpeng Zhang 0003

dblp:203/4474 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-3796-161XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Hardness-aware Privileged Features Distillation with Latent Alignment for CVR Prediction
abstract
In computational advertising, predicting the post-click conversion rate (CVR) using deep neural networks (DNNs) benefits a lot from privileged features, which can be collected for offline training but are unavailable during online serving.To utilize these privileged signals, privileged features distillation (PFD) methods incorporate a teacher model with privileged features to guide the CVR model.However, existing PFD approaches fail to put more emphasis on poorly predicted instances where the teacher's guidance is most crucial, and thus suffer from overconfidence on "easy" instances.In this work, we propose Hardness-aware Privileged Features Distillation (HA-PFD) for enhancing CVR prediction in real-world advertising recommender systems.We specifically design focal-style distillation losses that adaptively adjust the weight of each instance based on its "hardness".This method prioritizes poorly predicted instances during the distillation process, resulting in improved ranking performance and better model calibration.Additionally, we incorporate latent-level distillation into the PFD framework for the first time, which facilitates the student's representation learning through a straightforward layer alignment approach.We also propose a method for selecting privileged features based on their relevance to the conversion label.We conduct extensive offline experiments on large-scale, real-world datasets and online experiments on Douyin, a short video platform with billions of live users.In the offline evaluation, HA-PFD exhibits competitive performance and superior model calibration compared to existing state-of-the-art methods.In the online experiments, HA-PFD significantly improves advertiser value and conversions.Now we have deployed HA-PFD as the main online serving model on our short video platform.
Huining Yuan 0002, Wenpeng Zhang 0003, Zijie Hao, Zengde Deng
KDD (2)2
2023 Multi-Objective Online Learning
Jiyan Jiang, Wenpeng Zhang 0003, Shiji Zhou, Lihong Gu, Xiaodong Zeng, Wenwu Zhu 0001
ICLR2
2023 Model-free Reinforcement Learning with Stochastic Reward Stabilization for Recommender Systems
abstract
Model-free RL-based recommender systems have recently received increasing research attention due to their capability to handle partial feedback and long-term rewards. However, most existing research has ignored a critical feature in recommender systems: one user's feedback on the same item at different times is random. The stochastic rewards property essentially differs from that in classic RL scenarios with deterministic rewards, which makes RL-based recommender systems much more challenging. In this paper, we first demonstrate in a simulator environment where using direct stochastic feedback results in a significant drop in performance. Then to handle the stochastic feedback more efficiently, we design two stochastic reward stabilization frameworks that replace the direct stochastic feedback with that learned by a supervised model. Both frameworks are model-agnostic, i.e., they can effectively utilize various supervised models. We demonstrate the superiority of the proposed frameworks over different RL-based recommendation baselines with extensive experiments on a recommendation simulator as well as an industrial-level recommender system.
Tianchi Cai, Shenliao Bao, Jiyan Jiang, Shiji Zhou, Wenpeng Zhang 0003, Lihong Gu, Jinjie Gu
SIGIR5
2023 Marketing Budget Allocation with Offline Constrained Deep Reinforcement Learning
abstract
We study the budget allocation problem in online marketing campaigns that utilize previously collected offline data. We first discuss the long-term effect of optimizing marketing budget allocation decisions in the offline setting. To overcome the challenge, we propose a novel game-theoretic offline value-based reinforcement learning method using mixed policies. The proposed method reduces the need to store infinitely many policies in previous methods to only constantly many policies, which achieves nearly optimal policy efficiency, making it practical and favorable for industrial usage. We further show that this method is guaranteed to converge to the optimal policy, which cannot be achieved by previous value-based reinforcement learning methods for marketing budget allocation. Our experiments on a large-scale marketing campaign with tens-of-millions users and more than one billion budget verify the theoretical results and show that the proposed method outperforms various baseline methods. The proposed method has been successfully deployed to serve all the traffic of this marketing campaign.
Tianchi Cai, Jiyan Jiang, Wenpeng Zhang 0003, Shiji Zhou, Xierui Song, Lihong Gu, Xiaodong Zeng, Jinjie Gu
WSDM3
2022 Imbalance-Aware Uplift Modeling for Observational Data
abstract
Uplift modeling aims to model the incremental impact of a treatment on an individual outcome, which has attracted great interests of researchers and practitioners from different communities. Existing uplift modeling methods rely on either the data collected from randomized controlled trials (RCTs) or the observational data which is more realistic. However, we notice that on the observational data, it is often the case that only a small number of subjects receive treatment, but finally infer the uplift on a much large group of subjects. Such highly imbalanced data is common in various fields such as marketing and medical treatment but it is rarely handled by existing works. In this paper, we theoretically and quantitatively prove that the existing representative methods, transformed outcome (TOM) and doubly robust (DR), suffer from large bias and deviation on highly imbalanced datasets with skewed propensity scores, mainly because they are proportional to the reciprocal of the propensity score. To reduce the bias and deviation of uplift modeling with an imbalanced dataset, we propose an imbalance-aware uplift modeling (IAUM) method via constructing a robust proxy outcome, which adaptively combines the doubly robust estimator and the imputed treatment effects based on the propensity score. We theoretically prove that IAUM can obtain a better bias-variance trade-off than existing methods on a highly imbalanced dataset. We conduct extensive experiments on a synthetic dataset and two real-world datasets, and the experimental results well demonstrate the superiority of our method over state-of-the-art.
Xuanying Chen, Zhining Liu 0001, Liuyi Yao, Wenpeng Zhang 0003, Lihong Gu, Xiaodong Zeng, Yize Tan, Jinjie Gu
AAAI5
2022 Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training
Wenpeng Zhang 0003
ICLR3
2022 On the Convergence of Stochastic Multi-Objective Gradient Manipulation and Beyond
abstract
The conflicting gradients problem is one of the major bottlenecks for the effective training of machine learning models that deal with multiple objectives. To resolve this problem, various gradient manipulation techniques, such as PCGrad, MGDA, and CAGrad, have been developed, which directly alter the conflicting gradients to refined ones with alleviated or even no conflicts. However, the existing design and analysis of these techniques are mainly conducted under the full-batch gradient setting, ignoring the fact that they are primarily applied with stochastic mini-batch gradients. In this paper, we illustrate that the stochastic gradient manipulation algorithms may fail to converge to Pareto optimal solutions. Firstly, we show that these different algorithms can be summarized into a unified algorithmic framework, where the descent direction is given by the composition of the gradients of the multiple objectives. Then we provide an explicit two-objective convex optimization instance to explicate the non-convergence issue under the unified framework, which suggests that the non-convergence results from the determination of the composite weights solely by the instantaneous stochastic gradients. To fix the non-convergence issue, we propose a novel composite weights determination scheme that exponentially averages the past calculated weights. Finally, we show the resulting new variant of stochastic gradient manipulation converges to Pareto optimal or critical solutions and yield comparable or improved empirical performance.
Shiji Zhou, Wenpeng Zhang 0003, Jiyan Jiang, Leon Wenliang Zhong, Jinjie Gu, Wenwu Zhu 0001
NeurIPS2
2021 Asynchronous Decentralized Online Learning
abstract
Most existing algorithms in decentralized online learning are conducted in the synchronous setting. However, synchronization makes these algorithms suffer from the straggler problem, i.e., fast learners have to wait for slow learners, which significantly reduces such algorithms' overall efficiency. To overcome this problem, we study decentralized online learning in the asynchronous setting, which allows different learners to work at their own pace. We first formulate the framework of Asynchronous Decentralized Online Convex Optimization, which specifies the whole process of asynchronous decentralized online learning using a sophisticated event indexing system. Then we propose the Asynchronous Decentralized Online Gradient-Push (AD-OGP) algorithm, which performs asymmetric gossiping communication and instantaneous model averaging. We further derive a regret bound of AD-OGP, which is a function of the network topology, the levels of processing delays, and the levels of communication delays. Extensive experiments show that AD-OGP runs significantly faster than its synchronous counterpart and also verify the theoretical results.
Jiyan Jiang, Wenpeng Zhang 0003, Jinjie Gu, Wenwu Zhu 0001
NeurIPS2
2019 AutoML and Meta-learning for Multimedia
abstract
AutoML and meta-learning are exciting and fast-growing research directions to the research community in both academia and industry. This tutorial is to disseminate and promote the recent research achievements on AutoML and meta-learning as well as their potential applications for multimedia. Specifically, we will first advocate novel, high-quality research findings and innovative solutions to the challenging problems in AutoML and meta-learning. Then we will discuss scenarios of multimedia where AutoML and meta-learning serve as candidates for solutions. Finally, we will point out future research directions on AutoML and meta-learning as well as their potential new applications for multimedia.
Wenwu Zhu 0001, Xin Wang 0019, Wenpeng Zhang 0003
ACM Multimedia3
2018 Sketched Follow-The-Regularized-Leader for Online Factorization Machine
abstract
Factorization Machine (FM) is a supervised machine learning model for feature engineering, which is widely used in many real-world applications. In this paper, we consider the case that the data samples arrive sequentially. The existing convex formulation for online FM has the strong theoretical guarantee and stable performance in practice, but the computational cost is typically expensive when the data is high-dimensional. To address this weakness, we devise a novel online learning algorithm called Sketched Follow-The-Regularizer-Leader (SFTRL). SFTRL presents the parameters of FM implicitly by maintaining low-rank matrices and updates the parameters via sketching. More specifically, we propose Generalized Frequent Directions to approximate indefinite symmetric matrices in a streaming way, making that the sum of historical gradients for FM could be estimated with tighter error bound efficiently. With mild assumptions, we prove that the regret bound of SFTRL is close to that of the standard FTRL. Experimental results show that SFTRL has better prediction quality than the state-of-the-art online FM algorithms in much lower time and space complexities.
Luo Luo, Wenpeng Zhang 0003, Zhihua Zhang 0004, Wenwu Zhu 0001, Tong Zhang 0001, Jian Pei 0001
KDD2
2018 Online Compact Convexified Factorization Machine
abstract
Factorization Machine (FM) is a supervised learning approach with a powerful capability of feature engineering. It yields state-of-the-art performances in various batch learning tasks where all the training data is made available prior to the training. However, in real-world applications where the data arrives sequentially in a streaming manner, the high cost of re-training with batch learning algorithms has posed formidable challenges in the online learning scenario. The initial challenge is that no prior formulations of FM could directly fulfill the requirements in Online Convex Optimization (OCO) -- the paramount framework for online learning algorithm design. To address this aforementioned challenge, we invent a new convexification scheme leading to a Compact Convexified FM (CCFM) that seamlessly meets the requirements in OCO. However for learning Compact Convexified FM (CCFM) in the online learning settings, most existing algorithms suffer from expensive projection operations. To address this subsequent challenge, we follow the general projection-free algorithmic framework of Online Conditional Gradient and propose an Online Compact Convex Factorization Machine (OCCFM) algorithm that eschews the projection operation with efficient linear optimization steps. In support of the proposed OCCFM in terms of its theoretical foundation, we prove that the developed algorithm achieves a sub-linear regret bound. To evaluate the empirical performance of OCCFM, we conduct extensive experiments on 6 real-world datasets for online regression and online classification tasks. The experimental results show that OCCFM outperforms the state-of-art online learning methods for FM.
Xiao Lin 0002, Wenpeng Zhang 0003, Min Zhang 0006, Wenwu Zhu 0001, Jian Pei 0001, Peilin Zhao, Junzhou Huang
WWW2
2017 Projection-free Distributed Online Learning in Networks
abstract
The conditional gradient algorithm has regained a surge of research interest in recent years due to its high efficiency in handling large-scale machine learning problems. However, none of existing studies has explored it in the distributed online learning setting, where locally light computation is assumed. In this paper, we fill this gap by proposing the distributed online conditional gradient algorithm, which eschews the expensive projection operation needed in its counterpart algorithms by exploiting much simpler linear optimization steps. We give a regret bound for the proposed algorithm as a function of the network size and topology, which will be smaller on smaller graphs or “well-connected” graphs. Experiments on two large-scale real-world datasets for a multiclass classification task confirm the computational benefit of the proposed algorithm and also verify the theoretical regret bound.
Wenpeng Zhang 0003, Peilin Zhao, Wenwu Zhu 0001, Steven C. H. Hoi, Tong Zhang 0001
ICML1