VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Shi 0001
dblp:26/8377-1
· DBLP profile ↗
30ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-4267-7795ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Fairness-aware Interactive Recommendation: Item Lifecycle as a Control KnobabstractThis paper revisits fairness-aware interactive recommendation (e.g., TikTok, KuaiShou) by introducing a novel control knob, i.e., the lifecycle of items. We make threefold contributions. First, we conduct a comprehensive empirical analysis and uncover that item lifecycles in short-video platforms follow a compressed three-phase pattern, i.e., rapid growth, transient stability, and sharp decay, which significantly deviates from the classical four-stage model (introduction, growth, maturity, decline). Second, we introduce LHRL, a lifecycle-aware hierarchical reinforcement learning framework that dynamically harmonizes fairness and accuracy by leveraging phase-specific exposure dynamics. LHRL consists of two key components: (1) PhaseFormer, a lightweight encoder combining STL decomposition and attention mechanisms for robust phase detection; (2) a two-level HRL agent, where the high-level policy imposes phase-aware fairness constraints, and the low-level policy optimizes immediate user engagement. This decoupled optimization allows for effective reconciliation between long-term equity and short-term utility. Third, experiments on multiple real-world interactive recommendation datasets demonstrate that LHRL significantly improves both fairness and user engagement. Furthermore, the integration of lifecycle-aware rewards into existing RL-based models consistently yields performance gains, highlighting the generalizability and practical value of our approach. Xiaoyu Shi 0001, Hong Xie 0004, Chongjun Xia, Zhenhui Gong, Mingsheng Shang 0001 |
AAAI | 2 |
| 2026 | Dual debiasing learning for single-source federated domain generalization under data imbalance
Yan Wang 0147, Xiaoyu Shi 0001, Hong Xie 0004, Jinyang He, Mingsheng Shang 0001 |
Neurocomputing | 2 |
| 2026 | DPRO-GNN: Bridging differential privacy and advanced optimization for privacy-preserving graph learning
Yanan Bai, Liji Xiao, Hongbo Zhao 0011, Xiaoyu Shi 0001 |
Inf. Sci. | 4 |
| 2026 | BGV-TCF: A trust-guided homomorphic protocol for scalable privacy-preserving collaborative filtering
Yanan Bai, Hongbo Zhao 0011, Xiaoyu Shi 0001, Liji Xiao, Zepeng Gong, Yanglei Lu |
Knowl. Based Syst. | 3 |
| 2026 | DEL4CW: Deep Expansion Learning for Cloud Workloads PredictionabstractCloud Workload Prediction (CWP) is a critical task in cloud computing, essential for resource scheduling, performance optimization, and cost management. However, existing time series prediction methods struggle with instability and inefficiency when applied directly to cloud workloads due to their high variability and frequent fluctuations. To address these challenges, we propose DEL4CW, a novel D eep E xpansion L earning framework specifically designed for CWP . DEL4CW introduces a unique self-decoupling mechanism to disentangle the complex dependencies present in highly variable cloud workloads, leading to more accurate predictions of job arrival rates. The core contribution of DEL4CW lies in its ability to decouple cloud workload signals into three key components—trend, periodicity, and residuals—by treating these as hidden variables. This enables the model to better manage both short-term fluctuations and long-term workload trends. DEL4CW employs a deep expansion learning framework structured as stacked blocks, where each block includes dedicated modules for trend, periodicity, and compensation. Specifically, the trend module utilizes multi-layer fully connected networks to capture evolving trends at multiple granularities, while the periodicity module leverages multi-head attention to identify diverse periodic patterns. The compensation module addresses unpredictable, localized fluctuations, improving the model’s robustness to noise. In addition to its predictive accuracy, DEL4CW provides interpretable insights through its hierarchical design, allowing for layer-by-layer aggregation of meaningful partial predictions. This interpretability stems from the doubly residual learning pipeline, which ensures that each prediction block contributes progressively refined predictions. Extensive experiments on real-world cloud workload traces demonstrate that DEL4CW significantly outperforms existing baselines, with error reductions reaching up to 27.74% in certain scenarios. Xiaoyu Shi 0001, Qiuyue Lv, Bingchao Wang, Hong Xie 0004, Mingsheng Shang 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2026 | Beyond Trade-offs: Leveraging Spatiotemporal Heterogeneity of User Preference for Long-term Fairness and Accuracy in Interactive RecommendationabstractAs recommender systems are essential to various web domains such as e-commerce and web content sharing, providing equitable item exposure regardless of popularity becomes an imperative requirement. However, traditional fairness-aware approaches typically aim to achieve a better tradeoff between recommendation accuracy and fairness, and focus on improving the exposure rate of the long-tail items on static settings, evaluating fairness on one-shot recommendation decisions using logged data. Such methods overlook the dynamic nature of user preferences in real-world interactive environments. In contrast, our work seeks a win-win solution that simultaneously enhances recommendation accuracy and fairness over the long term, rather than merely trading off one against the other. To achieve this goal, we empirically demonstrate and analyze the spatiotemporal heterogeneity of user popularity preference. Our findings reveal complementary characteristics that, when fully exploited, can guide personalized strategies for long-term fairness. Building on this insight, we propose HER4IF, a novel hierarchical reinforcement learning framework designed for interactive recommendation. HER4IF decomposes the recommendation process into two key tasks: dynamic fairness control and item recommendation. The high-level agent continuously learns adaptive fairness constraints from evolving user popularity preferences, while the low-level agent refines recommendation policies under these personalized constraints. Extensive experiments on three real-world datasets and the interactive recommendation platform KuaiSim demonstrate that HER4IF significantly outperforms state-of-the-art methods, achieving substantial improvements in both fairness and recommendation accuracy. Our code is available at: https://github.com/1163710212/HER4IF . Chongjun Xia, Xiaoyu Shi 0001, Hong Xie 0004, Mingsheng Shang 0001 |
ACM Trans. Web | 2 |
| 2025 | Cross-Domain Semantic Transfer for Domain GeneralizationabstractData augmentation is a kind of mainstream domain generalization method aimed at enhancing the model’s ability to learn from out-of-distribution data. Most existing data augmentation methods fail to simultaneously preserve the semantic consistency and ensure the domain diversity in the augmented samples, which hinders further improvements in the generalization capacity of the model. To cope with this issue, we propose a novel cross-domain semantic transfer (CDST)-based data augmentation method, which improves generalizability from a novel perspective of exploring the diversity of semantic directions. Specifically, to ensure semantic consistency, an adjacent domain center interpolation module is proposed to find new generalization centers far from the classification boundary. To ensure data diversity, a semantic sample reproduction module is proposed to synthesize new semantic directions by transferring the cross-domain semantic information and reproduce new samples along new semantic directions around new feature centers. Furthermore, a category-preserving regularization is introduced to further constrain the category invariance of the augmented samples. Extensive experiments are implemented to verify the superiority, effectiveness, and transferability of CDST on the Digit-DG, PACS, VLCS, and OfficeHome datasets. Yan Wang 0147, Hong Xie 0004, Jinyang He, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Hierarchical Reinforcement Learning for Long-term Fairness in Interactive RecommendationabstractWith the growing influence of Recommender System (RS) in people’s daily lives, the issue of fairness in recommendations has become increasingly crucial. Previous fairness-aware methods focus on static or one-shot recommendation settings, where the recommendation model provides static fairness solutions via solving a fixed fairness-constrained optimization. However, these approaches face challenges in maintaining the delicate balance between recommendation accuracy and fairness in dynamic environments, as they fail to account for the evolving nature of RS, including changes in user preferences and item popularity over time.In this paper, we explore the problem of long-term fairness in interactive recommendation and accomplish the problem through dynamic fairness-constrained decisions. We focus on maintaining the exposure fairness of items across different groups. We argue that fairness does not always in conflict with recommendation performance, especially when considering the spatiotemporal heterogeneity of user preference on item popularity. To achieve this, we propose HER4IF, a dynamic fairness-aware interactive recommendation method based on a hierarchical reinforcement learning framework. Its main idea is first to aggregate the interacted item popularity with the time-forgetting model in state representation to capture user popularity preference. It then introduces a high-level agent to generate a dynamic fairness constraint based on the user’s current state, while a low-level agent generates recommendations under this constraint. Experiments on two datasets and an authentic Reinforcement Learning environment (KuaiSim) show the effectiveness and superiority of the proposed framework in terms of recommendation accuracy and fairness. It demonstrates a win-win for fairness and accuracy in a dynamic recommendation setting, when considering the dynamic nature of RS and incorporating spatiotemporal heterogeneity in user preference on item popularity. Chongjun Xia, Xiaoyu Shi 0001, Hong Xie 0004, Quanliang Liu, Mingsheng Shang 0001 |
ICWS | 2 |
| 2024 | Robust and efficient algorithms for conversational contextual bandit
Haoran Gu, Yunni Xia, Hong Xie 0004, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
Inf. Sci. | 4 |
| 2024 | Asynchronous SGD with stale gradient dynamic adjustment for deep learning training
Tao Tan 0008, Hong Xie 0004, Yunni Xia, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
Inf. Sci. | 4 |
| 2024 | Adaptive moving average Q-learning
Tao Tan 0008, Hong Xie 0004, Yunni Xia, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
Knowl. Inf. Syst. | 4 |
| 2024 | A Meta-Learning Approach to Mitigating the Estimation Bias of Q-LearningabstractIt is a longstanding problem that Q-learning suffers from the overestimation bias. This issue originates from the fact that Q-learning uses the expectation of maximum Q-value to approximate the maximum expected Q-value. A number of algorithms, such as Double Q-learning, were proposed to address this problem by reducing the estimation of maximum Q-value, but this may lead to an underestimation bias. Note that this underestimation bias may have a larger performance penalty than the overestimation bias. Different from previous algorithms, this article studies this issue from a fresh perspective, i.e., meta-learning view, which leads to our Meta-Debias Q-learning. The main idea is to extract the maximum expected Q-value with meta-learning over multiple tasks to remove the estimation bias of maximum Q-value and help the agent choose the optimal action more accurately. However, there are two challenges: (1) How to automatically select suitable training tasks? (2) How to positively transfer the meta-knowledge from selected tasks to remove the estimation bias of maximum Q-value? To address the two challenges mentioned above, we quantify the similarity between the training tasks and the test task. This similarity enables us to select appropriate “partial” training tasks and helps the agent extract the maximum expected Q-value to remove the estimation bias. Extensive experiment results show that our Meta-Debias Q-learning outperforms SOTA baselines drastically in three evaluation indicators, i.e., maximum Q-value, policy, and reward. More specifically, our Meta-Debias Q-learning only underestimates \(1.2*10^{-3}\) than the maximum expected Q-value in the multi-armed bandit environment and only differs \(5.04\%-5\%=0.04\%\) than the optimal policy in the two states MDP environment. In addition, we compare the uniform weight and our similarity weight. Experiment results reveal fundamental insights into why our proposed algorithm outperforms in the maximum Q-value, policy, and reward. Tao Tan 0008, Hong Xie 0004, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Probabilistic Modeling of Assimilate-Contrast Effects in Online Rating SystemsabstractOnline rating system serves as an indispensable building block for many web applications. Previous studies showed that due to assimilate-contrast effects, historical ratings could significantly distort users' ratings, leading to low accuracy of product quality estimation and recommendation. To understand assimilate-contrast effects, an “accurate” model is still missing as previous models do not capture important factors like rating recency, selection bias, etc. Furthermore, an analytical framework to characterize product estimation accuracy under assimilate-contrast effects is also missing. This paper aims to fill in this gap. We propose a probabilistic model to quantify the aforementioned important factors on assimilate-contrast effects. We apply stochastic approximation theory to show that when the rating bias satisfies mild contraction conditions, the aggregate rating converges under aggregate opinion heterogeneity. We also apply non-stationary Markov chain theory to show that when the strength of assimilate-contrast satisfies mild stable conditions, the aggregate rating converges under rating recency. We also derive an equation to characterize the converged aggregate ratings. These conditions reveal important insights on how the aforementioned factors influence the convergence and guide the online rating system operator to design appropriate rating aggregation rules and rating displaying strategies. We apply it to rating prediction tasks and product recommendation tasks. Experiment results on four public datasets show that our model can improve the rating prediction and recommendation accuracy over previous models significantly, under various metrics like RMSE, NDCG, etc. We also demonstrate the flexibility of our model by showing that it can be applied to enhance other rating behavior models. Hong Xie 0004, Mingze Zhong, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Relieving Popularity Bias in Interactive Recommendation: A Diversity-Novelty-Aware Reinforcement Learning ApproachabstractWhile personalization increases the utility of item recommendation, it also suffers from the issue of popularity bias. However, previous methods emphasize adopting supervised learning models to relieve popularity bias in the static recommendation, ignoring the dynamic transfer of user preference and amplification effects of the feedback loop in the recommender system (RS). In this paper, we focus on studying this issue in the interactive recommendation. We argue that diversification and novelty are both equally crucial for improving user satisfaction of IRS in the aforementioned setting. To achieve this goal, we propose a D iversity- N ovelty- a ware I nteractive R ecommendation framework (DNaIR) that augments offline reinforcement learning (RL) to increase the exposure rate of long-tail items with high quality. Its main idea is first to aggregate the item similarity, popularity, and quality into the reward model to help the planning of RL policy. It then designs a diversity-aware stochastic action generator to achieve an efficient and lightweight DNaIR algorithm. Extensive experiments are conducted on the three real-world datasets and an authentic RL environment (Virtual-Taobao). The experiments show that our model can better and full use of the long-tail items to improve recommendation satisfaction, especially those low popularity items with high-quality ones, thus achieving state-of-the-art performance. Xiaoyu Shi 0001, Quanliang Liu, Hong Xie 0004, Di Wu 0056, Bo Peng 0039, Mingsheng Shang 0001, Defu Lian |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Maximum Entropy Policy for Long-Term Fairness in Interactive Recommender SystemsabstractThis article considers the problem of maintaining the long-term fairness of item exposure in interactive recommender systems under the dynamic setting that user preference and item popularity evolve over time. The challenge is that the evolving dynamics of user preference and item popularity in the feedback loop amplify the long-term “unfairness” of item exposure. To address this challenge, we first formulate a constrained Markov Decision Process (MDP) to capture the evolving dynamics of user preference. The proposed constrained MDP imposes long-term fairness requirements via maximum entropy techniques. Moreover, to illuminate the “unfairness” amplifying effect caused by the evolving dynamic of item popularity in the feedback loop, we design a debiased reward function to eliminate popularity bias in the training data. To this end, the proposed framework can maintain acceptable recommendation accuracy while exposing items as randomly as possible, ensuring long-term benefits for users. To address the data sparsity issue, the proposed framework can easily integrate self-supervised learning methods to enhance state representation. Experiments on three datasets and an authentic Reinforcement Learning environment (Virtual-Taobao) demonstrate the effectiveness and superiority of the proposed framework in terms of recommendation accuracy and fairness, and show the robustness against data sparsity and noise. Xiaoyu Shi 0001, Quanliang Liu, Hong Xie 0004, Yanan Bai, Mingsheng Shang 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | A Self-decoupled Interpretable Prediction Framework for Highly-Variable Cloud Workloads
Bingchao Wang, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
DASFAA (1) | 2 |
| 2023 | Towards Long-term Fairness in Interactive Recommendation: A Maximum Entropy Reinforcement Learning ApproachabstractThis paper considers the problem of maintaining the long-term fairness of item exposure in interactive recommendation systems under the dynamic setting that user preference and item popularity evolve over time. The challenge is that the evolving dynamics of user preference and item popularity in the feedback loop amplify the long-term “unfairness” of item exposure. To address this challenge, we first formulate a constrained Markov Decision Process (MDP) to capture evolving dynamics of user preference. The proposed constrained MDP imposes long-term fairness requirements via maximum entropy techniques. Moreover, to illuminate the “unfairness” amplifying effect caused by the evolving dynamic of item popularity in the feedback loop, we design a debiased reward function to eliminate popularity bias in the training data. To this end, the proposed framework can maintain acceptable recommendation accuracy while exposing items as randomly as possible, ensuring long-term benefits for users. Experiments on three datasets demonstrate the effectiveness and superiority of our proposed framework in terms of recommendation performance and fairness. Xiaoyu Shi 0001, Quanliang Liu, Hong Xie 0004, Mingsheng Shang 0001 |
ICWS | 1 |
| 2023 | CGCCMR: An Enhanced Multi-gate Model for Cross-Market RecommendationabstractE-commerce applications such as Amazon and Tmall often provide services in multiple countries, forming multiple markets around the world. Generally, different markets over-lap in some items and differ in users. To make effective use of data from different markets, Cross-Market Recommendation (CMR) has been proposed to improve recommendation performance in data-scarce markets(i.e. target markets) using knowledge learned from other data-rich markets(i.e. source markets). Previous works on CMR do not capture the differences between markets well and still do not provide good recommendations to users in the target market. To address this limitation, we propose the CGCCMR model to leverage data from multiple markets and to improve the recommendation performance in the target market. CGCCMR first pre-trains the embedding layer by Light Graph Convolution(LGC) and performs Mask Data Augmentation on the samples from target markets. Then, it extracts the user interest representation related to the target item by User Interest Attention Network(UIAN) and adopts Cross-Market Customized Gate Control(CM-CGC) to effectively capture the similarities and differences between markets. Finally, it adopts an auxiliary network and a market-specific tower network to make final predictions. The experimental results from real-world datasets validate the superiority of the proposed CGCCMR model. Jinyu Mo, Hong Xie 0004, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
IJCNN | 3 |
| 2023 | Graph neural networks via contrast between separation and aggregation for self and neighborhood
Xiaoyu Shi 0001, Mingsheng Shang 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Neighbor importance-aware graph collaborative filtering for item recommendation
Qing-Xian Wang 0001, Suqiang Wu, Yanan Bai, Quanliang Liu, Xiaoyu Shi 0001 |
Neurocomputing | 5 |
| 2022 | Performance and Cost-Aware Task Scheduling via Deep Reinforcement Learning in Cloud Environment
Zihui Zhao, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
ICSOC | 2 |
| 2022 | Aspect-aware Asymmetric Representation Learning Network for Review-based RecommendationabstractRecently, user-provided reviews have been identified as an essential resource to improve user and item representation in recommender systems. Previous methods focus on the review-based recommender typically leverages symmetric networks to process user and item reviews. However, in reality, these two sets of reviews are markedly different: a user's reviews reflect the experience of buying diverse items and show their heterogeneous interests. In contrast, an item's reviews emphasize the quality of the specific item. Thus an item's reviews are usually homogeneous. This paper seeks to explore the aspect of review difference in the review-based recommendation framework. We propose a novel asymmetric neural network model that accurately learns the user and item representation by identifying this critical difference. We focus on capturing the dynamic change of user interest for the user-aspect reviews via modeling the temporal information into the conventional neural network(CNN). On the other side, we try to identify a specific item's essential yet essential features by utilizing the self-attention neural network. Finally, a factorization machine (FM) is adopted to finish the rating prediction task, where the user and item IDs are encoded as supplementary review embedding. We conduct comprehensive experiments on four Amazon datasets, and the experimental results show that our proposed model consistently outperforms several state-of-the-art methods. Hezhe Qiao, Xiaoyu Shi 0001, Mingsheng Shang 0001 |
IJCNN | 3 |
| 2022 | Large-Scale and Scalable Latent Factor Analysis via Distributed Alternative Stochastic Gradient Descent for Recommender SystemsabstractLatent factor analysis (LFA) via stochastic gradient descent (SGD) is highly efficient in discovering user and item patterns from high-dimensional and sparse (HiDS) matrices from recommender systems. However, most LFA-based recommender systems adopt a standard SGD algorithm, which suffers limited scalability when addressing big data. On the other hand, most existing parallel SGD solvers are either under the memory-sharing framework designed for a bare machine or suffering high communicational costs, which also greatly limits their applications in large-scale systems. To address the above issues, this article proposes a distributed alternative stochastic gradient descent (DASGD) solver for an LFA-based recommender. Its training-dependences among latent features are decoupled via alternatively fixing one-half of the features to learn the other half following the principle of SGD but in parallel. It's distribution mechanism consists of efficient data partition, allocation and task parallelization strategies, which greatly reduces its communicational cost for high scalability. Experimental results on three large-scale HiDS matrices generated by real-world applications demonstrate that the proposed DASGD algorithm outperforms state-of-the-art distributed SGD solvers for recommender systems in terms of prediction accuracy as well as scalability. Hence, it is highly useful for training LFA-based recommenders on large scale HiDS matrices with the help of cloud computing facilities. Xiaoyu Shi 0001, Qiang He 0001, Xin Luo 0001, Yanan Bai, Mingsheng Shang 0001 |
IEEE Trans. Big Data | 1 |
| 2021 | Joint Modeling Dynamic Preferences of Users and Items Using Reviews for Sequential Recommendation
Tianqi Shang, Xiaoyu Shi 0001, Qing-Xian Wang 0001 |
PAKDD (2) | 3 |
| 2020 | A Clustering-Based Collaborative Filtering Recommendation Algorithm via Deep Learning User Side Information
Chonghao Zhao, Xiaoyu Shi 0001, Mingsheng Shang 0001, Yiqiu Fang |
WISE (2) | 2 |
| 2019 | Elastic-net regularized latent factor analysis-based models for recommender systems
Dexian Wang 0004, Yanbin Chen, Junxiao Guo, Xiaoyu Shi 0001, Xin Luo 0001, Huaqiang Yuan |
Neurocomputing | 4 |
| 2018 | Incremental Slope-one recommenders
Qing-Xian Wang 0001, Xin Luo 0001, Xiaoyu Shi 0001, Liang Gu, Mingsheng Shang 0001 |
Neurocomputing | 4 |
| 2018 | On minimizing total energy consumption in the scheduling of virtual machine reservations
Wenhong Tian, Majun He, Wenxia Guo, Wenqiang Huang, Xiaoyu Shi 0001, Mingsheng Shang 0001, Adel Nadjaran Toosi, Rajkumar Buyya |
J. Netw. Comput. Appl. | 5 |
| 2017 | Long-term performance of collaborative filtering based recommenders in temporally evolving systems
Xiaoyu Shi 0001, Xin Luo 0001, Mingsheng Shang 0001, Liang Gu |
Neurocomputing | 1 |
| 2016 | PAPMSC: Power-Aware Performance Management Approach for Virtualized Web Servers via Stochastic Control
Xiaoyu Shi 0001, Jin Dong 0001, Seddik M. Djouadi, Xiao Ma 0008, Yefu Wang |
J. Grid Comput. | 1 |