EDBT 2026 Demo / reviewers in the wild / expert
Xiaoqiang Gui
dblp:303/9201
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Recommendation via Stochastic Aggregation and Consistency InferenceabstractWith growing concerns over user privacy, federated recommendation (FedRec) has emerged as a mainstream solution for personalized recommendation services. FedRec trains user-private parameters on local clients while collaboratively updating global parameters on a centralized server. However, despite advances in optimizing these local and global parameters, existing methods overlook two key challenges: tradeoff training and distribution discrepancy . Tradeoff training balances timely local updates with diverse global parameters, limiting the model’s learning ability. Distribution discrepancy arises from the divergence between locally trained global parameters and those aggregated by the server, corrupting inference performance. To fill in the gap, we propose FedSC , a principled federated recommendation framework that boosts FedRec’s training and inference processes with minimal yet nontrivial efforts. During training, FedSC employs a stochastic aggregation strategy where all users participate in every round, while only a random subset is selected for aggregation, preserving the diversity of global parameters and ensuring timely local updates. During inference, FedSC makes recommendations with a consistency inference mechanism that uses the most recent locally trained global parameters of each user to improve the model’s understanding of user preferences. Extensive experiments on multiple benchmark datasets demonstrate the superiority of FedSC, achieving up to a 20% improvement in most evaluation scenarios. Xiaoqiang Gui, Qiaoyu Tan, Jun Wang 0035, Yongqing Zheng, Qingzhong Li, Li-Zhen Cui 0001, Guoxian Yu |
ACM Trans. Inf. Syst. | 1 |
| 2025 | Sophon: Byzantine-Robust Federated Learning via Dual Trust MechanismabstractFederated Learning is a data privacy-protected distributed machine learning framework, but malicious clients can damage it. Byzantine-robust federated learning aims to learn an accurate global model despite the presence of malicious clients. Most current defenses assume that clients have identically and independently distributed (i.i.d.) data and thus cannot perform well in canonical non-i.i.d. scenarios. Several non-i.i.d. statistic-based defenses have been recently proposed to identify malicious clients through gradient statistics without any auxiliary techniques. They thus can only ensure robustness under certain attacks with specific characteristics. We propose Sophon, a comprehensive defense using auxiliary data to combat an arbitrary number of malicious clients in both i.i.d. and non-i.i.d. cases. Specifically, Sophon first normalizes received client gradients to reduce the dominance of malicious gradients. Then it introduces a dual trust mechanism to assign the aggregation weight for each gradient. The dual trust mechanism estimates the consistency-based and diversity-based trust scores of client gradients and integrates the two scores as the aggregation weight to effectively suppress the impact of malicious gradients. Extensive experimental results on three datasets from different domains, with diverse models and FL scenarios, show that Sophon is robust in maintaining the overall accuracy of the training model. Xiaoqiang Gui, Guoxian Yu, Jun Wang 0035, Zhongmin Yan, Wei Wang 0012, Carlotta Domeniconi, Li-Zhen Cui 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Interaction Privacy Vulnerability in Federated Recommendation and Lossless CountermeasureabstractFederated Recommendation (FedRec) systems are recognized as privacy-preserving solutions for collaboratively training recommender models without sharing users’ private data. However, recent studies have revealed that FedRec systems are vulnerable to interaction-level membership inference attacks. In such attacks, a semi-honest server can employ crafted methods to infer users’ interacted items. In this article, we identify that user preference information is predominantly stored in the user-uploaded parameters rather than in the local parameters after local training. Leveraging this insight, we expose a new interaction vulnerability and introduce the PubPara attack. Our experiments show that PubPara improves the inference performance by at least 40% over existing attacks, while requiring minimal inference time and remaining robust against current defense methods. To safeguard user privacy without compromising recommender performance, we propose MultiVerse, a novel countermeasure. MultiVerse utilizes untrained items outside the user’s local training data to obfuscate the server’s inference of interacted items. It includes a four-step strategy (training, optimization, refinement, and denoising) to achieve robust defense. Extensive experiments on three representative FedRec models (F-NCF, F-LightGCN, and FedRAP) across three real-world datasets validate that MultiVerse significantly degrades the attack’s inference performance to near the level of random guess while maintaining lossless recommender performance. Xiaoqiang Gui, Guoxian Yu, Jun Wang 0035, Shuguang Han, Qingzhong Li, Yongqing Zheng, Wei Wang 0012 |
ACM Trans. Inf. Syst. | 1 |
| 2024 | Calibration-compatible Listwise Distillation of Privileged Features for CTR PredictionabstractIn machine learning systems, privileged features refer to the features that are available during offline training but inaccessible for online serving. Previous studies have recognized the importance of privileged features and explored ways to tackle online-offline discrepancies. A typical practice is privileged features distillation (PFD): train a teacher model using all features (including privileged ones) and then distill the knowledge from the teacher model using a student model (excluding the privileged features), which is then employed for online serving. In practice, the pointwise cross-entropy loss is often adopted for PFD. However, this loss is insufficient to distill the ranking ability for CTR prediction. First, it does not consider the non-i.i.d. characteristic of the data distribution, i.e., other items on the same page significantly impact the click probability of the candidate item. Second, it fails to consider the relative item order ranked by the teacher model's predictions, which is essential to distill the ranking ability. To address these issues, we first extend the pointwise-based PFD to the listwise-based PFD. We then define the calibration-compatible property of distillation loss and show that commonly used listwise losses do not satisfy this property when employed as distillation loss, thus compromising the model's calibration ability, which is another important measure for CTR prediction. To tackle this dilemma, we propose Calibration-compatible LIstwise Distillation (CLID), which employs carefully-designed listwise distillation loss to achieve better ranking ability than the pointwise-based PFD while preserving the model's calibration ability. We theoretically prove it is calibration-compatible. Extensive experiments on public datasets and a production dataset collected from the display advertising system of Alibaba further demonstrate the effectiveness of CLID. Xiaoqiang Gui, Yueyao Cheng, Xiang-Rong Sheng, Guoxian Yu, Shuguang Han, Yuning Jiang 0001, Jian Xu 0015, Bo Zheng 0007 |
WSDM | 1 |
| 2023 | Entire Space Cascade Delayed Feedback Modeling for Effective Conversion Rate PredictionabstractConversion rate (CVR) prediction is an essential task for e-commerce platforms. However, refunds frequently occur after conversion in online shopping systems, which drives us to pay attention to effective conversion for building healthier services. This paper defines the probability of item purchasing without any subsequent refund as an effective conversion rate (ECVR). A simple paradigm for ECVR prediction is to decompose it into two sub-tasks: CVR prediction and post-conversion refund rate (RFR) prediction. However, RFR prediction suffers from data sparsity (DS) and sample selection bias (SSB) issues, as refund behaviors are only available after user purchase. Furthermore, there is delayed feedback in both sequentially dependent conversion and refund events, named cascade delayed feedback (CDF). Previous studies mainly focus on tackling DS and SSB or delayed feedback for a single event. To jointly tackle these issues in ECVR prediction, we propose an Entire space CAscade Delayed feedback modeling (ECAD) method. Specifically, ECAD deals with DS and SSB by constructing two tasks including CVR and conversion&refund rate (CVRFR) predictions using the entire space modeling framework. In addition, it carefully schedules auxiliary tasks to leverage both conversion and refund time within data to alleviate CDF. Experiments on the offline industrial dataset and online A/B testing demonstrate the effectiveness of ECAD. ECAD has been deployed in the Xianyu recommender system of Alibaba, contributing to a significant improvement of ECVR. Xiaoqiang Gui, Shuguang Han, Xiang-Rong Sheng, Guoxian Yu, Jufeng Chen, Bo Zheng 0007 |
CIKM | 3 |
| 2021 | Cost-effective Batch-mode Multi-label Active Learning
Xiaoqiang Gui, Xudong Lu 0001, Guoxian Yu |
Neurocomputing | 1 |