Fan Yang 0107

dblp:29/3081-107 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0005-9390-8791ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Non-autoregressive Generative Auction with Global Externalities for Online Advertising
abstract
Online advertising auctions play a critical role in internet commerce, requiring mechanisms that maximize revenue while ensuring incentive compatibility, user experiences, and real-time efficiency. Existing learning-based auction frameworks advance contextual modeling by considering intra-list dependencies among ads, but still face challenges of insufficient global externality modeling and inefficiencies due to sequential processing. In this paper, we propose the Non-autoregressive Generative Auction with global externalities (NGA), a novel end-to-end auction framework for industrial online advertising. NGA explicitly models global externalities by jointly encoding dependencies among ads and the influence of neighboring organic content. To achieve real-time efficiency, NGA employs a non-autoregressive, constraint-based decoding mechanism and a parallel multi-tower evaluator that unifies list-wise reward and payment computation. Extensive offline experiments and large-scale online A/B tests on commercial advertising platforms demonstrate that NGA achieves superior performance in both effectiveness and efficiency compared to the state-of-the-art baselines.
Zuowu Zheng, Ze Wang 0005, Fan Yang 0107, Wenqing Ye, Weihua Huang, Wenqiang He
CIKM3
2023 PIER: Permutation-Level Interest-Based End-to-End Re-ranking Framework in E-commerce
abstract
Re-ranking draws increased attention on both academics and industries, which rearranges the ranking list by modeling the mutual influence among items to better meet users' demands. Many existing re-ranking methods directly take the initial ranking list as input, and generate the optimal permutation through a well-designed context-wise model, which brings the evaluation-before-reranking problem. Meanwhile, evaluating all candidate permutations brings unacceptable computational costs in practice. Thus, to better balance efficiency and effectiveness, online systems usually use a two-stage architecture which uses some heuristic methods such as beam-search to generate a suitable amount of candidate permutations firstly, which are then fed into the evaluation model to get the optimal permutation. However, existing methods in both stages can be improved through the following aspects. As for generation stage, heuristic methods only use point-wise prediction scores and lack an effective judgment. As for evaluation stage, most existing context-wise evaluation models only consider the item context and lack more fine-grained feature context modeling.
Xiaowen Shi, Fan Yang 0107, Ze Wang 0005, Xiaoxu Wu, Muzhi Guan, Guogang Liao, Yongkang Wang 0011, Dong Wang 0022
KDD2
2023 MDDL: A Framework for Reinforcement Learning-based Position Allocation in Multi-Channel Feed
abstract
Nowadays, the mainstream approach in position allocation system is to utilize a reinforcement learning model to allocate appropriate locations for items in various channels and then mix them into the feed. There are two types of data employed to train reinforcement learning (RL) model for position allocation, named strategy data and random data. Strategy data is collected from the current online model, it suffers from an imbalanced distribution of stateaction pairs, resulting in severe overestimation problems during training. On the other hand, random data offers a more uniform distribution of state-action pairs, but is challenging to obtain in industrial scenarios as it could negatively impact platform revenue and user experience due to random exploration. As the two types of data have different distributions, designing an effective strategy to leverage both types of data to enhance the efficacy of the RL model training has become a highly challenging problem. In this study, we propose a framework namedMulti-Distribution Data Learning (MDDL) to address the challenge of effectively utilizing both strategy and random data for training RL models on mixed multi-distribution data. Specifically, MDDL incorporates a novel imitation learning signal to mitigate overestimation problems in strategy data and maximizes the RL signal for random data to facilitate effective learning. In our experiments, we evaluated the proposed MDDL framework in a real-world position allocation system and demonstrated its superior performance compared to the previous baseline. MDDL has been fully deployed on the Meituan food delivery platform and currently serves over 300 million users.
Xiaowen Shi, Ze Wang 0005, Yuanying Cai, Xiaoxu Wu, Fan Yang 0107, Guogang Liao, Yongkang Wang 0011, Dong Wang 0022
SIGIR5