EDBT 2026 Demo / reviewers in the wild / expert
Xichun Sheng
dblp:393/2349
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0002-3590-0060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedAU2: Attribute Unlearning for User-Level Federated Recommender Systems with Adaptive and Robust Adversarial TrainingabstractFederated Recommender Systems (FedRecs) leverage federated learning to protect user privacy by retaining data locally. However, user embeddings in FedRecs often encode sensitive attribute information, rendering them vulnerable to attribute inference attacks. Attribute unlearning has emerged as a promising approach to mitigate this issue. In this paper, we focus on user-level FedRecs, which is a more practical yet challenging setting compared to group-level FedRecs. Adversarial training emerges as the most feasible approach within this context. We identify two key challenges in implementing adversarial training-based attribute unlearning for user-level FedRecs: i) mitigating training instability caused by user data heterogeneity, and ii) preventing attribute information leakage through gradients. To address these challenges, we propose FedAU2, an attribute unlearning method for user-level FedRecs. For CH1, we propose a adaptive adversarial training strategy, where the training dynamics are adjusted in response to local optimization behavior. For CH2, we propose a dual-stochastic variational autoencoder to perturb the adversarial model, effectively preventing gradient-based information leakage. Extensive experiments on three real-world datasets demonstrate that our proposed FedAU2 achieves superior performance in unlearning effectiveness and recommendation performance compared to existing baselines. Yuyuan Li 0001, Junjie Fang, Fengyuan Yu 0001, Xichun Sheng, Tianyu Du, Xuyang Teng, Shaowei Jiang, Linbo Jiang, Jianan Lin 0003, Chaochao Chen 0001 |
AAAI | 4 |
| 2026 | Forgetting Knowledge Localization and Isolation for Continual Forgetting of Pre-trained Vision ModelsabstractContinual forgetting task aims to continuously remove multiple target knowledge subsets from pre-trained models while maintaining the integrity of remaining knowledge. Existing methods suffer from both incomplete forgetting of target knowledge and unintended forgetting of indistinguishable remaining knowledge. To address these challenges, we propose the forgetting knowledge localization and isolation for continual forgetting in pre-trained vision models which precisely forgets target knowledge while reducing over-forgetting of remaining knowledge. To achieve precise forgetting, we first propose the forgetting knowledge layer localization to explore layers in the model which are more related to forgetting knowledge. Then, we design the forgetting knowledge parameter isolation to isolate the parameters sensitive to forgetting knowledge in these selected layers, mitigating over-forgetting of remaining knowledge. Finally, we fine-tune these isolated parameters and freeze the remaining parameters to achieve efficient forgetting while maintaining high performance on retained datasets. Extensive experimental results demonstrate that our method achieves superior performance over state-of-the-art methods across multiple continual forgetting tasks. Zhiwen Yang 0003, Chenggang Yan 0001, Zongpeng Li, Xichun Sheng, Liang Li 0003 |
AAAI | 6 |
| 2026 | Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalabstractExisting text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing it presents two key challenges: (1) Temporally dense scene text disrupts sync with sparse video frames, obstructing video understanding;(2) Redundant scene text and irrelevant video frames hinder the learning of discriminative temporal clues for retrieval. To address them, we propose a temporal scene-text calibrating and distilling (TCD) network for textvideo retrieval. Specifically, we first design a window-OCR captioner that aggregates dense scene text into OCR captions to facilitate feature interaction. Next, we devise a heterogeneous semantics calibration module that leverages scene text as a self-supervised signal to temporally align window-level OCR captions and frame-level video features. Further, we introduce a context-guided temporal clue distillation module to learn the complementary and relevant details between scene text and video modalities, thereby obtaining discriminative temporal clues for retrieval. Extensive experiments show that our TCD achieves state-of-the-art performance on three scene-text related benchmarks. Zhiqian Zhao, Liang Li 0003, Xichun Sheng, Yaoqi Sun, Fang Kang, Chenggang Yan 0001 |
AAAI | 4 |
| 2026 | Lightweight multi-scale weight pruning network for salient object detectionabstractSalient object detection (SOD) is fundamental to computer vision, yet deep learning approaches often suffer from high computational costs, limiting deployment on resource-constrained devices. We propose a Lightweight Multi-scale Weight Pruning Network (LMWP-Net) to balance high performance with low complexity. LMWP-Net employs an encoder–decoder architecture featuring two key components: a Multi-scale Weight Pruning Module (MWPM) for efficient feature extraction and redundancy reduction, and a Multi-scale Attention Fusion Module (MAFM) for effective integration via attention mechanisms. Extensive experiments on public datasets demonstrate that LMWP-Net consistently outperforms existing lightweight methods and achieves competitive accuracy against state-of-the-art models. Remarkably, compared to the prominent BANet, LMWP-Net achieves a 94.6% reduction in parameters and a 99.5% reduction in FLOPs, validating its superior efficiency and effectiveness for real-time applications. The implemented code is publicly available at https://github.com/IMOP-lab/LMWP-Net . Xichun Sheng, Yaoqi Sun, Gaopeng Huang, Ya-Hong Chen, Jin Liu 0025, Xiaoshuai Zhang, Xingru Huang |
J. Vis. Commun. Image Represent. | 1 |
| 2026 | Event-aware temporal modeling and semantic alignment for long-form video question answering
Xichun Sheng, Haibo Gong, Liang Li 0003, Chenggang Yan 0001, Tao Tan 0002 |
Pattern Recognit. | 1 |
| 2026 | Empirical Study on Fusion Strategy in RGB-T Salient Object DetectionabstractIn the research field of RGB-Thermal saliency object detection (RGB-T SOD), the effective exploitation of the complementary characteristics of the two modalities represents a major challenge for enhancing detection performance. Current fusion methodologies can be roughly classified into early fusion and middle fusion strategies, with prevalent techniques primarily encompassing concatenation, summation, and multiplication of the two modalities. To in depth assess the efficacy of these fusion strategies, we took an empirical investigation on them. Our findings demonstrate that the concatenation of middle features constitutes a more advantageous fusion strategy, yielding superior performance and demonstrating enhanced stability. Furthermore, observing the unique properties of thermal (T) images, we introduced gamma correction as a novel data augmentation methodology to RGB-T SOD. We subsequently evaluated the responses across varying correction parameter ranges, revealing that while the response to this data augmentation technique differs across various models, data augmentation is found to be effective in general. Building upon these findings, we proposed the Gamma Correction Network (GaCNet). Specifically, we also integrated image pyramid mechanism in a lightweight manner, which facilitates a more effective recovery of fine-grained image details. Significant improvement was achieved on commonly used RGB-T testing datasets, especially in VT821 dataset, manifesting the effectiveness of our method. Shuai Wang 0003, Qiang Zhao 0005, Junbo Ma, Xichun Sheng, Yaoqi Sun, Hongfa Wen, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Prompt Learning With Knowledge Regularization for Pre-Trained Vision-Language ModelsabstractPrompt learning is an effective way to adapt pre-trained models to downstream tasks by training a small number of additional learnable prompts. Recent studies address several early challenges by combining generalized knowledge from frozen pre-trained VL models with task-specific knowledge from training data as guidance for prompt learning. However, existing methods still struggle with the generalization-adaptation (GA) trade-off dilemma: excessive reliance on generalized knowledge hinders adaptation to downstream tasks, while overemphasis on task-specific knowledge undermines the inherent generalization capabilities of pre-trained models. To address this issue, we propose a novel prompt learning method called Prompt Learning with Knowledge Regularization (PLKR). PLKR effectively mitigates the GA trade-off dilemma by offering greater flexibility in adapting to task-specific knowledge while minimizing the disruption of pre-trained knowledge. Specifically, we propose category-invariant and topology-invariant knowledge regularization to preserve generalized knowledge: the former enhances category-level discriminative capabilities while allowing flexible task-specific learning, and the latter maintains global topological stability during adaptation to new tasks. Through the proposed regularization, PLKR improves the performance on both base and new tasks. We evaluate the effectiveness of our approach on four representative tasks over 11 datasets. Experimental results show our method outperforms existing SOTA methods by a large margin. Boyang Guo, Liang Li 0003, Yaoqi Sun, Chenggang Yan 0001, Xichun Sheng |
IEEE Trans. Multim. | 6 |
| 2025 | Heterogeneous Prompt-Guided Entity Inferring and Distilling for Scene-Text Aware Cross-Modal RetrievalabstractIn cross-modal retrieval, comprehensive image understanding is vital while the scene text in images can provide fine-grained information to understand visual semantics. Current methods fail to make full use of scene text. They suffer from the semantic ambiguity of independent scene text and overlook the heterogeneous concepts in image-caption pairs. In this paper, we propose a heterogeneous prompt-guided entity inferring and distilling (HOPID) network to explore the nature connection of scene text in images and captions and learn a property-centric scene text representation. Specifically, we propose to align scene text in images and captions via heterogeneous prompt, which consists of visual and text prompt. For text prompt, we introduce the discriminative entity inferring module to reason key scene text words from captions, while visual prompt highlights the corresponding scene text in images. Furthermore, to secure a robust scene text representation, we design a perceptive entity distilling module that distills the beneficial information of scene text at a fine-grained level. Extensive experiments show that the proposed method significantly outperforms existing approaches on two public cross-modal retrieval benchmarks. Zhiqian Zhao, Liang Li 0003, Yaoqi Sun, Xichun Sheng, Haibing Yin, Shaowei Jiang |
AAAI | 5 |
| 2025 | Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain AdaptationabstractThis paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-class representation learning and the shocks of new target knowledge to old ones. To address them, we propose the Multi-Granularity Class Prototype Topology Distillation (GROTO) algorithm, which effectively transfers the source knowledge to the class-incremental target domain. Concretely, we design the multi-granularity class prototype self-organization module and the prototype topology distillation module. First, we mine the positive classes by modeling accumulation distributions. Next, we introduce multi-granularity class prototypes to generate reliable pseudo-labels, and exploit them to promote the positive-class target feature self-organization. Second, the positive-class prototypes are leveraged to construct the topological structures of source and target feature spaces. Then, we perform the topology distillation to continually mitigate the shocks of new target knowledge to old ones. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on three public datasets. Peihua Deng, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Ying Fu 0001, Liang Li 0003 |
CVPR | 3 |
| 2025 | Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning
Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Xichun Sheng |
ICCV | 6 |
| 2025 | Hypergraph-Guided Federated Distillation Learning for Efficient and Robust Multi-center fMRI Data Analysis
Yidan Xu, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Xiangmin Han, Yue Gao 0002 |
MICCAI (11) | 4 |
| 2025 | Improving Integrated Satellite-Terrestrial Cell-Free Massive MIMO Systems by Rate-Splitting Multiple AccessabstractWe investigate the spectral and energy efficiencies of the uplink in an integrated satellite-terrestrial cell-free massive multiple-input multiple-output (IST-CF-mMIMO) system assisted by rate-splitting multiple access (RSMA). In the IST-CF-mMIMO system, the terrestrial users employ RSMA to transmit a message as a superposition of two parts with different power to the terrestrial access points and low-Earth-orbit satellite. Taking realistic conditions such as the spatially correlated Ricean fading channels, imperfect channel knowledge, and successive interference cancellation into account, we derive rigorous closed-form expressions for uplink achievable spectral and energy efficiencies and evaluate these performance metrics across a range of system configurations. Additionally, to enhance the system energy efficiency, we formulate the design of users’ power control coefficients as an energy efficiency optimization problem and design an efficient algorithm based on Lagrangian dual transformation and quadratic transformation techniques to solve it. Comprehensive simulations validate our theoretical propositions and evaluate the efficacy of the proposed energy efficiency maximization algorithm. Yao Zhang 0016, Jintao Shen, Yaoqi Sun, Xichun Sheng, Haitao Zhao 0004, Hongbo Zhu 0002 |
IEEE Internet Things J. | 6 |