Can Xu 0004

dblp:33/965-4 · DBLP profile ↗
← Back
7ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0002-4254-8678ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4Information Retrieval & Web Search · 3
YearPublicationVenuePosition
2025 Reward Shaping for User Satisfaction
Konstantina Christakopoulou, Can Xu 0004, Sriraj Badam, Trevor Potter, Xinyang Yi, Ya Le, Chris Berg, Eric Bencomo Dixon, Ed H. Chi, Minmin Chen
ECML/PKDD (6)2
2022 Surrogate for Long-Term User Experience in Recommender Systems
abstract
Over the years we have seen recommender systems shifting focus from optimizing short-term engagement toward improving long-term user experience on the platforms. While defining good long-term user experience is still an active research area, we focus on one specific aspect of improved long-term user experience here, which is user revisiting the platform. These long term outcomes however are much harder to optimize due to the sparsity in observing these events and low signal-to-noise ratio (weak connection) between these long-term outcomes and a single recommendation. To address these challenges, we propose to establish the association between these long-term outcomes and a set of more immediate term user behavior signals that can serve as surrogates for optimization.
Mohit Sharma 0002, Can Xu 0004, Sriraj Badam, Qian Sun 0005, Lee Richardson, Lisa Chung, Ed H. Chi, Minmin Chen
KDD3
2022 Off-Policy Actor-critic for Recommender Systems
abstract
Industrial recommendation platforms are increasingly concerned with how to make recommendations that cause users to enjoy their long term experience on the platform. Reinforcement learning emerged naturally as an appealing approach for its promise in 1) combating feedback loop effect resulted from myopic system behaviors; and 2) sequential planning to optimize long term outcome. Scaling RL algorithms to production recommender systems serving billions of users and contents, however remain challenging. Sample inefficiency and instability of online RL hinder its widespread adoption in production. Offline RL enables usage of off-policy data and batch learning. It on the other hand faces significant challenges in learning due to the distribution shift.
Minmin Chen, Can Xu 0004, Vince Gatto, Devanshu Jain, Aviral Kumar, Ed H. Chi
RecSys2
2021 Values of User Exploration in Recommender Systems
abstract
Reinforcement Learning (RL) has been sought after to bring next-generation recommender systems to further improve user experience on recommendation platforms. While the exploration-exploitation tradeoff is the foundation of RL research, the value of exploration in (RL-based) recommender systems is less well understood. Exploration, commonly seen as a tool to reduce model uncertainty in regions of sparse user interaction/feedback, is believed to cost user experience in the short term, while the indirect benefit of better model quality arrives at a later time. We focus on another aspect of exploration, which we refer to as user exploration to help discover new user interests, and argue it can improve user experience even in the more imminent term.
Minmin Chen, Can Xu 0004, Ya Le, Mohit Sharma 0002, Lee Richardson, Su-Lin Wu, Ed H. Chi
RecSys3
2021 User Response Models to Improve a REINFORCE Recommender System
abstract
Reinforcement Learning (RL) techniques have been sought after as the next-generation tools to further advance the field of recommendation research. Different from classic applications of RL, recommender agents, especially those deployed on commercial recommendation platforms, have to operate in extremely large state and action spaces, serving a dynamic user base in the order of billions, and a long-tail item corpus in the order of millions or billions. The (positive) user feedback available to train such agents is extremely scarce in retrospect. Improving the sample efficiency of RL algorithms is thus of paramount importance when developing RL agents for recommender systems. In this work, we present a general framework to augment the training of model-free RL agents with auxiliary tasks for improved sample efficiency. More specifically, we opt to add additional tasks that predict users' immediate responses (positive or negative) toward recommendations, i.e., user response modeling, to enhance the learning of the state and action representations for the recommender agents. We also introduce a tool based on gradient correlation analysis to guide the model design. We showcase the efficacy of our method in offline experiments, learning and evaluating agent policies over hundreds of millions of user trajectories. We also conduct live experiments on an industrial recommendation platform serving billions of users and tens of millions of items to verify its benefit.
Minmin Chen, Bo Chang 0002, Can Xu 0004, Ed H. Chi
WSDM3
2019 Towards Neural Mixture Recommender for Long Range Dependent User Sequences
abstract
Understanding temporal dynamics has proved to be highly valuable for accurate recommendation. Sequential recommenders have been successful in modeling the dynamics of users and items over time. However, while different model architectures excel at capturing various temporal ranges or dynamics, distinct application contexts require adapting to diverse behaviors.
Jiaxi Tang, Francois Belletti, Sagar Jain, Minmin Chen, Alex Beutel, Can Xu 0004, Ed H. Chi
WWW6
2018 Latent Cross: Making Use of Context in Recurrent Recommender Systems
abstract
The success of recommender systems often depends on their ability to understand and make use of the context of the recommendation request. Significant research has focused on how time, location, interfaces, and a plethora of other contextual features affect recommendations. However, in using deep neural networks for recommender systems, researchers often ignore these contexts or incorporate them as ordinary features in the model.
Alex Beutel, Paul Covington, Sagar Jain, Can Xu 0004, Vince Gatto, Ed H. Chi
WSDM4