EDBT 2026 Demo / reviewers in the wild / expert
Xuemei Cao 0001
dblp:235/1471-1
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-9083-8998ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhanced Federated Deep Multi-View Clustering Under Uncertainty ScenarioabstractTraditional Federated Multi-View Clustering assumes uniform views across clients, yet practical deployments reveal heterogeneous view completeness with prevalent incomplete, redundant, or corrupted data. While recent approaches model view heterogeneity, they neglect semantic conflicts from dynamic view combinations, failing to address dual uncertainties: view uncertainty (semantic inconsistency from arbitrary view pairings) and aggregation uncertainty (divergent client updates with imbalanced contributions). To address these, we propose a novel Enhanced Federated Deep Multi-View Clustering framework: first align local semantics, hierarchical contrastive fusion within clients resolves view uncertainty by eliminating semantic conflicts; a view adaptive drift module mitigates aggregation uncertainty through global-local prototype contrast that dynamically corrects parameter deviations; and a balanced aggregation mechanism coordinates client updates. Experimental results demonstrate that EFDMVC achieves superior robustness against heterogeneous uncertain views across multiple benchmark datasets, consistently outperforming all state-of-the-art baselines in comprehensive evaluations. Bingjun Wei, Xuemei Cao 0001, Jiafen Liu, Haoyang Liang, Xin Yang 0012 |
AAAI | 2 |
| 2026 | G-GBC: A Gaussian mixture-based granular ball computing method for robust CNN classification under noisy labels
Junxiao Yang, Qiang Liu 0021, Xuemei Cao 0001, Xin Yang 0012 |
Knowl. Based Syst. | 3 |
| 2026 | Continual deep multi-view clustering via contrastive knowledge replay
Haoyang Liang, Bingjun Wei, Xuemei Cao 0001, Jiafen Liu, Xiaocao Ouyang, Jiangtao Qiu, Hao Wang 0068, Xin Yang 0012, Tianrui Li 0001 |
Pattern Recognit. | 3 |
| 2026 | Lifelong multi-view clustering with anchor-prototype collaboration
Yonghao Li, Xuemei Cao 0001, Hao Yu 0023, Jiafen Liu, Xin Yang 0012 |
Pattern Recognit. | 3 |
| 2025 | ErrorEraser: Unlearning Data Bias for Improved Continual LearningabstractContinual Learning (CL) primarily aims to retain knowledge to prevent catastrophic forgetting and transfer knowledge to facilitate learning new tasks. Unlike traditional methods, we propose a novel perspective: CL not only needs to prevent forgetting, but also requires intentional forgetting.This arises from existing CL methods ignoring biases in real-world data, leading the model to learn spurious correlations that transfer and amplify across tasks. From feature extraction and prediction results, we find that data biases simultaneously reduce CL's ability to retain and transfer knowledge. To address this, we propose ErrorEraser, a universal plugin that removes erroneous memories caused by biases in CL, enhancing performance in both new and old tasks. ErrorEraser consists of two modules: Error Identification and Error Erasure. The former learns the probability density distribution of task data in the feature space without prior knowledge, enabling accurate identification of potentially biased samples. The latter ensures only erroneous knowledge is erased by shifting the decision space of representative outlier samples. Additionally, an incremental feature distribution learning strategy is designed to reduce the resource overhead during error identification in downstream tasks. Extensive experimental results show that ErrorEraser significantly mitigates the negative impact of data biases, achieving higher accuracy and lower forgetting rates across three types of CL methods. The code is available at https://github.com/diadai/ErrorEraser. Xuemei Cao 0001, Hanlin Gu, Xin Yang 0012, Bingjun Wei, Haoyang Liang, Xiangkun Wang, Tianrui Li 0001 |
KDD (2) | 1 |
| 2025 | RACE: Robust adaptive and clustering elimination for noisy labels in continual learning
Guannan Lai, Dan Meng 0004, Xuemei Cao 0001, Xin Yang 0012 |
Knowl. Based Syst. | 4 |
| 2025 | Ten Challenging Problems in Federated Foundation ModelsabstractFederated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large foundation models and the small local domain models at the remote clients to learn from each other in a teacher-student learning setting. This paper provides a comprehensive summary of the ten challenging problems inherent in FedFMs, encompassing foundational theory, utilization of private data, continual learning, unlearning, Non-IID and graph data, bidirectional knowledge transfer, incentive mechanism design, game mechanism design, model watermarking, and efficiency. The ten challenging problems manifest in five pivotal aspects: “Foundational Theory,” which aims to establish a coherent and unifying theoretical framework for FedFMs. “Data,” addressing the difficulties in leveraging domain-specific knowledge from private data while maintaining privacy; “Heterogeneity,” examining variations in data, model, and computational resources across clients; “Security and Privacy,” focusing on defenses against malicious attacks and model theft; and “Efficiency,” highlighting the need for improvements in training, communication, and parameter efficiency. For each problem, we offer a clear mathematical definition on the objective function, analyze existing methods, and discuss the key challenges and potential solutions. This in-depth exploration aims to advance the theoretical foundations of FedFMs, guide practical implementations, and inspire future research to overcome these obstacles, thereby enabling the robust, efficient, and privacy-preserving FedFMs in various real-world applications. Tao Fan 0002, Hanlin Gu, Xuemei Cao 0001, Chee Seng Chan, Qian Chen 0023, Yiqiang Chen 0001, Yihui Feng, Yang Gu 0001, Jiaxiang Geng, Bing Luo 0002, Shuoling Liu, WinKent Ong, Chao Ren 0006, Jiaqi Shao, Xiaoli Tang 0001, Hong Xi Tae, Yongxin Tong, Shuyue Wei 0001, Fan Wu 0006, Wei Xi 0003, Mingcong Xu, Xin Yang 0012, Jiangpeng Yan, Hao Yu 0023, Han Yu 0001, Xiaojin Zhang 0002, Zhenzhe Zheng 0001, Lixin Fan, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Open Continual Feature Selection via Granular-Ball Knowledge TransferabstractThis paper presents a novel framework for continual feature selection (CFS) in data preprocessing, particularly in the context of an open and dynamic environment where unknown classes may emerge. CFS encounters two primary challenges: the discovery of unknown knowledge and the transfer of known knowledge. To this end, we propose a GBCFS method, which combines the strengths of continual learning (CL) with granular-ball computing (GBC). The GBCFS method focuses on constructing a granular-ball knowledge base to detect unknown classes and facilitate the transfer of previously learned knowledge for further feature selection. GBCFS consists of two stages: initial learning and open learning. The former aims to establish an initial knowledge base through multi-granularity representation using granular balls. The latter utilizes prior granular-ball knowledge to identify unknowns, updates the knowledge base for granular-ball knowledge transfer, reinforces old knowledge, and integrates new knowledge. Subsequently, we devise an optimal feature subset mechanism that incorporates minimal new features into the existing optimal subset, often yielding superior results during each period. Extensive experimental results on public benchmark datasets demonstrate our method's superiority in terms of both effectiveness and efficiency compared to state-of-the-art feature selection methods. Xuemei Cao 0001, Xin Yang 0012, Shuyin Xia, Guoyin Wang 0001, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Interpreting vulnerabilities of multi-instance learning to adversarial perturbations
Xuemei Cao 0001, Zhengchun Zhou, Mei Yang 0002, Avik Ranjan Adhikary |
Pattern Recognit. | 3 |
| 2022 | Hypersphere Neighborhood Rough Set for Rapid Attribute Reduction
Yu Fang 0009, Xuemei Cao 0001, Xin Wang 0064, Fan Min 0001 |
PAKDD (2) | 2 |
| 2022 | Three-way sampling for rapid attribute reductionabstractAs data dimensions and volume rapidly increase, attribute reduction using the original data becomes computationally infeasible. Large data frequently contain various redundant attributes and types of noise. This leads to the problems of overfitting and inefficiency in data processing. To address these problems, this paper proposes a general sampling method for attribute reduction by introducing three-way decisions, namely, three-way sampling (3WS), which is the first sampling method that describes the decision boundary accurately while improving the data quality significantly. To improve the effectiveness and efficiency of attribute reduction, we designed a rapid attribute reduction method based on three-way sampling (3WS-RAR). The 3WS-RAR method consists of three main steps: data sampling, attribute reduction, and model effectiveness evaluation. For data sampling, we define the three regions of the 3WS using support vectors to describe the data and use the boundary region as the sampling results. For the attribute reduction, we compute the neighborhood self-information for each attribute while considering the upper and lower approximations. For the effectiveness evaluation, we conducted experiments on 15 relatively large-scale datasets and analysed the influence of parameters. The experimental results reveal that, compared with state-of-the-art attribute reduction models, 3WS-RAR performs better on public benchmark datasets. Yu Fang 0009, Xuemei Cao 0001, Xin Wang 0064, Fan Min 0001 |
Inf. Sci. | 2 |