EDBT 2026 Demo / reviewers in the wild / expert
Sanchuan Guo
dblp:119/2649
· DBLP profile ↗
6ranked-venue papers in the field
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4 (1 first)Information Retrieval & Web Search · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A multi-expert adaptive framework for test-time personalization in federated learning
Sanchuan Guo, Zongyi Chen, Chaozhuo Li, Xi Zhang 0008 |
Inf. Process. Manag. | 1 |
| 2026 | Discovering new intents via spatio-temporal pseudo-label denoising
Yuming Shang, Wei Huang 0013, Sanchuan Guo, Jinhu Chen, Xi Zhang 0008, Philip S. Yu |
Inf. Process. Manag. | 4 |
| 2023 | Improving Knowledge Distillation for Federated Learning on Non-IID DataabstractFederated learning (FL) leverages knowledge from decentralized clients to train a global model in a privacy-preserving manner. One of the critical challenges in FL is the heterogeneity of client local data (i.e. non-IID data), which can lead to significant performance degradation. Federated Distillation (FD) can mitigate this issue by distilling the client predictions on the unlabeled public data into a global student model. However, a simple averaging of the client predictions without considering model confidences may lead to erroneous predictions. Moreover, it is challenging to rectify the predicted pseudo labels as no human labels are available in the public data. To address these issues, we propose a novel FD framework named FedUSL, aiming to distill a powerful global model on non-IID data. Specifically, we estimate the uncertainty of each client’s prediction on the unlabeled data and adopt a weighted ensemble of their predictions in consideration of uncertainties. We further rectify the global model predictions by a self-label reassigning method, without the requirement of manual labels. Extensive experiments on image and text tasks show that our proposal can achieve superior performance than state-of-the-art methods, and incurs no extra computing burden on the client side. The code is available at https://anonymous.4open.science/r/FedUSL-6338/. Zongyi Chen, Sanchuan Guo, Liyan Shen, Xi Zhang 0008, Zhuonan Chang |
IEEE Big Data | 2 |
| 2023 | Clean-label Poisoning Attack against Fake News Detection ModelsabstractResearching data poisoning attacks against fake news detection models is crucial for bolstering their robustness and curbing the dissemination of fake news. Existing textual data poisoning attacks necessitate control over both the content and labels of news samples, making them impractical for real attack scenarios. In this paper, we propose COMCP, a novel clean-label poisoning attack model aimed at fake news detection models. Diverging from existing methods, COMCP ensures the poison samples are accurately labeled, while crafting stealthy poison comments without modifying the headlines or content, thereby enhancing the feasibility of the attack. Furthermore, COMCP generates poison comments by appending stealthy characters to ensure the stealthiness of the attack. Comprehensive experimental evaluations on three benchmark datasets illustrate that our proposal outperforms SOTA baselines in terms of attack success rate and text quality, while maintaining the accuracy of detecting clean samples. Jiayi Liang, Xi Zhang 0008, Yuming Shang, Sanchuan Guo, Chaozhuo Li |
IEEE Big Data | 4 |
| 2022 | Abusive Language Detection with Graph based Multi-task LearningabstractTo counter the online abusive language in social media, it is desirable to develop automated detection methods. Previous research has primarily formulated this problem as a sentence-level classification task, ignoring the crucial role of abusive lexicons that can strengthen the model explainability and enable more faithful predictions. Although a few methods have introduced the abusive lexicons for detection, the lexicons they use are either externally provided or labeled by human annotators, suffering from two limitations: (1) lack adaptability to diverse and evolving offensive scenarios; (2) require large human efforts to annotate the words.This paper overcomes the limitations of prior work with a multi-task abusive language detection framework. It combines sentence-level and word-level classification tasks, based on dependency tree based graph attention networks (GAT). With the two tasks, it is encouraged to capture both global and local data properties to produce better sentence representations. It is also advantageous in automatic lexicon construction during the learning process, without human annotations. Extensive experiments on two public datasets exhibit that our proposal can outperform the state-of-the-art baselines. Case studies show that the model explainability can be strengthened with the abusive parts identified by our framework. Our code is released to public.1 Chunyun Zhang, Xi Zhang 0008, Quan Wang 0002, Jiayi Liang, Sanchuan Guo, Wenyu Zang, Yongdong Zhang 0001 |
IEEE Big Data | 6 |
| 2014 | Linear Rate Estimation Model for HEVC RDO Using Binary Classification Based RegressionabstractRate-Distortion Optimization in High Efficiency Video Coding promotes the coding efficiency, but also imposes intensive computation to the encoder, because the complex Syntax-based context-adaptive Binary Arithmetic Coding is performed for each candidate coding configuration. We develop the classification based regression method to derive the rate models, which fast estimate the bit cost of quantization coefficient block from its distribution features. Experiments demonstrate that, our method reduces the averaged 28.4% computation time in rate cost estimation, while the coding efficiency degradation is 0.0428dB. Sanchuan Guo, Zhenyu Liu 0001, Dongsheng Wang 0002, Qingrui Han, Yang Song 0002 |
DCC | 1 |