Yanyan Zou 0003

dblp:48/6274-3 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-8153-9299ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Mix Preference Optimization for Generative Recommendation
abstract
Recommender systems aim to leverage user interaction signals to recommend items that users are likely to be interested in. Motivated by the success of Large Language Model (LLMs), generative recommendation (GR) has recently gained increasing attention, typically following a two-stage paradigm: supervised fine-tuning followed by preference alignment. However, aligning generative recommenders with users' personalized preferences remains challenging, as user feedback is inherently heterogeneous and uncertain. Different types of user interaction signals reflect varying levels of intent and should therefore be modeled differently. In this work, we propose Adaptive Mix Preference Optimization (AMPO), an adaptive alignment framework that mixes likelihood and preference objectives with self-calibrated, sample-wise confidence adjustment. AMPO introduces an adaptive target margin that leverages the model's own probability ratio to modulate optimization strength: confident pairs receive full margins that reinforce correct rankings, while uncertain pairs receive reduced margins that prevent overfitting to ambiguous signals. Additionally, AMPO incorporates negative log-likelihood regularization on preferred items to counteract likelihood displacement, a phenomenon where contrastive objectives cause preferred and non-preferred probabilities to collapse simultaneously. Such a design eliminates the need for a reference model, yielding up to 2.5x speedup and 35% memory reduction. Extensive experiments on public benchmarks and a large-scale industrial dataset demonstrate consistent improvements in ranking metrics. Online A/B tests on a major e-commerce platform further confirm statistically significant gains in click-through and conversion rates. The code is available at https://github.com/jumbo-q/ampo.
Junbo Qi, Yanyan Zou 0003, Xuanhua Yang, Sulong Xu, Ying Sun 0026, Shengjie Li 0001
SIGIR2
2026 Breaking the Relevance-Diversity Seesaw: Hierarchical LLM Reasoning with RL for Industrial Novelty Recommendation
abstract
Novelty recommendation sustains long-term user engagement by exposing users to content that is both relevant and meaningfully different from their recent consumption. In large-scale e-commerce, this requires composing coherent yet non-redundant recommendation lists, a task fundamentally constrained by the relevance-diversity trade-off. Large language models (LLMs) offer a unified generative paradigm for inferring user intent and producing semantically coherent candidates, yet industrial deployment faces two critical challenges: (i) scarce supervision for modeling novelty transitions and diversity-aware list construction, and (ii) reward granularity mismatch, where standard RL assigns coarse sequence-level rewards that fail to capture item-level redundancy and complementarity. We present BALANCE, a hierarchical reasoning-and-generation framework that decomposes novelty recommendation into three structured stages: generating a Novelty Tag for exploration direction, refining an Interest Topic for intent specification, and constructing a Recommendation List for facet coverage. We address data scarcity through a self-reflection pipeline that synthesizes high-quality supervision by integrating real behavior logs with structured rationales. We resolve granularity mismatch through Sequence-Item Policy Optimization (SIPO), which jointly optimizes sequence- and item-level objectives via granularity-aware advantage fusion. Extensive offline experiments and online A/B test on the JD.com recommender system, validate the performance of our method, highlighting its superior novelty and diversity without compromising relevance.
Ying Sun 0026, Yanyan Zou 0003, Xiao Wang 0097, Hanchuan Xu, Xuanhua Yang, Sulong Xu, Junbo Qi, Shengjie Li 0001
SIGIR2
2026 GenRec: A Preference-Oriented Generative Framework for Large-Scale Recommendation
abstract
Generative Retrieval (GR) offers a promising paradigm for recommendation through next-token prediction (NTP). However, scaling it to large-scale industrial systems introduces three challenges: (i) within a single request, the identical model inputs may produce inconsistent outputs due to the pagination request mechanism; (ii) the prohibitive cost of encoding long user behavior sequences with multi-token item representations based on semantic IDs, and (iii) aligning the generative policy with nuanced user preference signals. We present GenRec, a preference-oriented generative framework deployed on the JD App https://www.jd.com that addresses above challenges within a single decoder-only architecture. For training objective, we propose Page-wise NTP task, which supervises over an entire interaction page rather than each interacted item individually, providing denser gradient signal and resolving the one-to-many ambiguity of point-wise training. On the prefilling side, an asymmetric linear Token Merger compresses multi-token Semantic IDs in the prompt while preserving full-resolution decoding, reducing input length by ~2× with negligible accuracy loss. To further align outputs with user satisfaction, we introduce GRPO-SR, a reinforcement learning method that pairs Group Relative Policy Optimization with NLL regularization for training stability, and employs Hybrid Rewards combining a dense reward model with a relevance gate to mitigate reward hacking. In month-long online A/B tests serving production traffic, GenRec achieves 9.5% improvement in click count and 8.7% in transaction count over the existing pipeline.
Yanyan Zou 0003, Junbo Qi, Lunsong Huang, Kewei Xu, Jiahao Gao, Binglei Zhao 0002, Xuanhua Yang, Sulong Xu, Shengjie Li 0001
SIGIR1
2025 WiLife: Long-Term Daily Status Monitoring and Habit Mining of the Elderly Leveraging Ubiquitous Wi-Fi Signals
abstract
The global aging demographic underscores the imperative for continuous in-home monitoring of the empty-nest elderly, ensuring their safety and well-being. The widespread deployment of Wi-Fi infrastructure has paved the way to monitor the elderly in a non-intrusive and privacy-preserving manner. Numerous studies have explored the potential of utilizing Wi-Fi signals to address urgent life safety concerns such as fall detection and vital sign monitoring. However, apart from these acute safety issues, the early detection of potential disease symptoms and managing the progression of chronic diseases are also crucial for elderly care, which calls for long-term and continuous monitoring of the elderly’s daily routines. Unfortunately, challenges like continuous activity segmentation and location/orientation dependencies have hindered the implementation of a long-term, around-the-clock activity monitoring system for the elderly. This work introduces “WiLife,” a cutting-edge Wi-Fi-based framework for continuous monitoring of the elderly’s spatio-temporal daily status information. Specifically, WiLife adopts a strategy of partitioning living spaces into functional areas and categorizing daily activities into atomic states . By encapsulating daily life status into a unique series of triple unit format: \(\left\langle\textit{Time, Area, State}\right\rangle\) , WiLife is able to offer valuable insights into when, where, and how activities occur. Field implementations spanning 1,080 hours (45 days \(\times\) 24 hours) in real-world home environments highlight WiLife’s exceptional capability in understanding individual living habits and timely detection of irregularities.
Shengjie Li 0001, Zhaopeng Liu, Qin Lv, Yanyan Zou 0003, Daqing Zhang 0001
ACM Trans. Comput. Heal.4
2023 An Industrial Framework for Personalized Serendipitous Recommendation in E-commerce
abstract
Classical recommendation methods typically face the filter bubble problem where users likely receive recommendations of their familiar items, making them bored and dissatisfied. To alleviate such an issue, this applied paper introduces a novel framework for personalized serendipitous recommendation in an e-commerce platform (i.e., JD.com), which allows to present user unexpected and satisfying items deviating from user’s prior behaviors, considering both accuracy and novelty. To achieve such a goal, it is crucial yet challenging to recognize when a user is willing to receive serendipitous items and how many novel items are expected. To address above two challenges, a two-stage framework is designed. Firstly, a DNN-based scorer is deployed to quantify the novelty degree of a product category based on user behavior history. Then, we resort to a potential outcome framework to decide the optimal timing to recommend a user serendipitous items and the novelty degree of the recommendation. Online A/B test on the e-commerce recommender platform in JD.com demonstrates that our model achieves significant gains on various metrics, 0.54% relative increase of impressive depth, 0.8% of average user click count, 3.23% and 1.38% of number of novel impressive and clicked items individually.
Zongyi Wang, Yanyan Zou 0003, Anyu Dai, Linfang Hou, Nan Qiao 0011, Luobao Zou, Mian Ma, Zhuoye Ding, Sulong Xu
RecSys2
2022 Automatic Product Copywriting for E-commerce
abstract
Product copywriting is a critical component of e-commerce recommendation platforms. It aims to attract users' interest and improve user experience by highlighting product characteristics with textual descriptions. In this paper, we report our experience deploying the proposed Automatic Product Copywriting Generation (APCG) system into the JD.com e-commerce product recommendation platform. It consists of two main components: 1) natural language generation, which is built from a transformer-pointer network and a pre-trained sequence-to-sequence model based on millions of training data from our in-house platform; and 2) copywriting quality control, which is based on both automatic evaluation and human screening. For selected domains, the models are trained and updated daily with the updated training data. In addition, the model is also used as a real-time writing assistant tool on our live broadcast platform. The APCG system has been deployed in JD.com since Feb 2021. By Sep 2021, it has generated 2.53 million product descriptions, and improved the overall averaged click-through rate (CTR) and the Conversion Rate (CVR) by 4.22% and 3.61%, compared to baselines, respectively on a year-on-year basis. The accumulated Gross Merchandise Volume (GMV) made by our system is improved by 213.42%, compared to the number in Feb 2021.
Yanyan Zou 0003, Hainan Zhang 0001, Shiliang Diao, Zhuoye Ding, Xueqi He, Bo Long, Han Yu 0001, Lingfei Wu 0001
AAAI2
2022 Summarizing Dialogues with Negative Cues
abstract
Abstractive dialogue summarization aims to convert a long dialogue content into its short form where the salient information is preserved while the redundant pieces are ignored. Different from the well-structured text, such as news and scientific articles, dialogues often consist of utterances coming from two or more interlocutors, where the conversations are often informal, verbose, and repetitive, sprinkled with false-starts, backchanneling, reconfirmations, hesitations, speaker interruptions and the salient information is often scattered across the whole chat. The above properties of conversations make it difficult to directly concentrate on scattered outstanding utterances and thus present new challenges of summarizing dialogues. In this work, rather than directly forcing a summarization system to merely pay more attention to the salient pieces, we propose to explicitly have the model perceive the redundant parts of an input dialogue history during the training phase. To be specific, we design two strategies to construct examples without salient pieces as negative cues. Then, the sequence-to-sequence likelihood loss is cooperated with the unlikelihood objective to drive the model to focus less on the unimportant information and also pay more attention to the salient pieces. Extensive experiments on the benchmark dataset demonstrate that our simple method significantly outperforms the baselines with regard to both semantic matching and factual consistent based metrics. The human evaluation also proves the performance gains.
Junpeng Liu 0001, Yanyan Zou 0003, Yuxuan Xi, Shengjie Li 0001, Mian Ma, Zhuoye Ding
COLING2
2022 Negative Guided Abstractive Dialogue Summarization
Junpeng Liu 0001, Yanyan Zou 0003, Yuxuan Xi, Shengjie Li 0001, Mian Ma, Zhuoye Ding, Bo Long
INTERSPEECH2
2021 Adaptive Bridge between Training and Inference for Dialogue Generation
abstract
Although exposure bias has been widely studied in some NLP tasks, it faces its unique challenges in dialogue response generation, the representative one-to-various generation scenario.In real human dialogue, there are many appropriate responses for the same context, not only with different expressions, but also with different topics.Therefore, due to the much bigger gap between various ground-truth responses and the generated synthetic response, exposure bias is more challenging in dialogue generation task.What's more, as MLE encourages the model to only learn the common words among different ground-truth responses, but ignores the interesting and specific parts, exposure bias may further lead to the common response generation problem, such as "I don't know" and "HaHa?"In this paper, we propose a novel adaptive switching mechanism, which learns to automatically transit between ground-truth learning and generated learning regarding the word-level matching score, such as the cosine similarity.Experimental results on both Chinese STC dataset and English Reddit dataset, show that our adaptive method achieves a significant improvement in terms of metric-based evaluation and human evaluation, as compared with the state-of-the-art exposure bias approaches.Further analysis on NMT task also shows that our model can achieve a significant improvement.
Hainan Zhang 0001, Yanyan Zou 0003, Hongshen Chen, Zhuoye Ding, Yanyan Lan
EMNLP (1)3