EDBT 2026 Demo / reviewers in the wild / expert
Xianxun Zhu
dblp:259/6633
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-3958-7040ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Injecting image text structure and edge priors into segment anything for scene text segmentation
Qian Shao, Libo Weng, Yanjing Lei, Xianxun Zhu, Hui Chen 0026 |
Image Vis. Comput. | 5 |
| 2026 | WiVi-UF: Unified feature learning in cross-modal transformers with WiFi and vision data fusion for enhanced human activity recognition
Xinhang Lin, Xianxun Zhu, Erik Cambria |
Knowl. Based Syst. | 2 |
| 2026 | DMSAA-SLAM: RGB-D SLAM for dynamic scenes via diffusion self-attentionabstract• Design a self-attention aggregation module using a pre-trained diffusion model. • Integrate high-precision masks into RGB-D SLAM for robust dynamic tracking. • Validate superior accuracy and efficiency on dynamic simulation datasets. In dynamic environments, performing RGB-D SLAM (Simultaneous Localization and Mapping) faces significant challenges primarily due to the presence of moving objects. The motion of these objects can introduce tracking errors and inaccuracies in map construction, thereby compromising the stability and overall performance of the system. To maintain high-precision localization and mapping under such conditions, a SLAM system must effectively detect and handle dynamic objects. To address these challenges, this paper presents a novel RGB-D SLAM method, referred to as DMSAA-SLAM (Dynamic Scene SLAM Based on Diffusion Model Self-Attention Aggregation). The core idea is to leverage a pre-trained stable diffusion model, particularly its self-attention layers, to handle the complexity of dynamic scenes. By employing a multi-resolution aggregation approach, combined with iterative merging and nonmaximum suppression, the proposed method generates high-precision segmentation masks. These masks enable fine-grained segmentation of moving objects and effectively eliminate dynamic feature points, thereby mitigating the impact of dynamic elements on the SLAM process and ensuring efficient and accurate tracking and mapping. Hui Chen 0026, Xianxun Zhu, Ling Fan |
Pattern Recognit. | 5 |
| 2026 | Advancing federated domain generalization in ophthalmology: Vision enhancement and consistency assurance for multicenter fundus image segmentation
Yang Zhao 0019, Xianxun Zhu, Jun Wang 0121, Yan Liu 0052 |
Pattern Recognit. | 4 |
| 2026 | AOSNet-Sec: Aperture-orientation-spectrum fusion with statistical Markov repair for trustworthy super-resolution
Zhengnan Yin, Luwei Xiao, Xuan Feng 0002, Yiwei Chen 0002, Xianxun Zhu, Cai Luo, Faten S. Alamri, Rui Mao 0010, Erik Cambria |
Pattern Recognit. | 5 |
| 2026 | Exploring personalized federated learning from a distribution-based perspectiveabstractPersonalized federated learning (PFL) is a promising technique for tackling data heterogeneity in federated learning systems. Recently, Bayesian neural networks (BNNs) have been introduced into the PFL framework to enable uncertainty quantification and improve performance in data-scarce settings. Despite these advantages, existing BNN-based PFL methods face two key challenges in practical applications. First, in real-world scenarios, client heterogeneity often arises in the form of group-wise variation, which cannot be adequately captured by a single shared distribution as assumed in prior work. Second, existing methods rely on deterministic or stochastic approximation techniques for posterior inference, which lead to substantial computational and memory overhead, hindering their scalability and deployment. To address these limitations, we propose DBFed, a novel BNN-based PFL framework from a distribution-based perspective. DBFed introduces group-specific distributions to better model the structural heterogeneity commonly observed in federated settings. Moreover, DBFed employs a rank-1 parameterization technique to map uncertainty from the weight space to a low-dimensional subspace, significantly reducing the computational and memory overhead. Theoretically, we establish the effectiveness of the rank-1 parameterization approach. Empirically, extensive experiments on diverse datasets demonstrate that DBFed consistently outperforms alternative PFL baselines in a heterogeneous setting. Tianhao Yu, Kheng Cher Yeo, Sami Azam, Xiaohan Yu 0001, Hui Chen 0026, Xianxun Zhu |
Pattern Recognit. | 6 |
| 2026 | Uncertainty-aware multimodal affective data fusion for personalized mental health dialogueabstractPersonalized affective dialogue systems are critical for mental health applications, where responses must be emotionally appropriate and tailored to individual users. However, most existing large language model (LLM) based approaches rely on deterministic personalization, ignore uncertainty in affective understanding, and are difficult to deploy under privacy constraints. In this paper, we propose PALLM, a personalized affective large language modeling framework designed for privacy-preserving mental health dialogue. PALLM decouples personalization into two complementary components: a deterministic personalization layer that captures stable user preferences, and a Bayesian affective representation layer that models dynamic emotional states and uncertainty. By restricting uncertainty modeling to affective representations rather than full LLM parameters, PALLM achieves efficient and scalable personalization under federated learning. Extensive experiments on EmpatheticDialogues and a real-world mental health conversation dataset show that PALLM improves affective alignment and robustness compared with non-personalized and partially personalized baselines. Xianxun Zhu, Erik Cambria, Hui Chen 0026 |
Pattern Recognit. | 1 |
| 2026 | FedBayesMamba: Uncertainty-aware federated learning for multimodal and audio-visual sequential modeling with selective state space modelsabstractFederated learning has emerged as an effective paradigm for training machine learning models across distributed clients without sharing raw data. In many real-world applications, sequential data are inherently multimodal, involving heterogeneous streams such as audio, visual, and temporal signals. However, most existing federated approaches rely on deterministic neural networks, which often struggle to capture predictive uncertainty under heterogeneous data distributions, cross-modal inconsistencies, and dynamic client participation. In this paper, we propose FedBayesMamba , a Bayesian federated learning framework for multimodal sequential data modeling based on selective state space models. The proposed approach introduces Bayesian parameterization into the Mamba architecture to enable uncertainty-aware sequence modeling while preserving the computational efficiency of state space models. To effectively integrate uncertainty across distributed clients, we further develop a posterior aggregation strategy that combines client-level posterior distributions in a principled probabilistic manner. Extensive experiments on multiple benchmark datasets demonstrate that the proposed framework achieves competitive predictive performance and improved uncertainty estimation under Non-IID federated settings. The results also indicate that FedBayesMamba exhibits strong robustness and stability in challenging federated scenarios. These findings highlight the potential of combining Bayesian learning with state space models for multimodal temporal modeling, particularly in audio-visual perception and cross-modal sequence understanding tasks. Xianxun Zhu, Xiaosong E, Michele Nappi, Imad Rida, Hui Chen 0026 |
Pattern Recognit. | 1 |
| 2026 | PBDD: A Prompt-Based Learning Approach for Few-Shot Social Media Depression DetectionabstractAutomated detection of depressive moods from social media holds great promise for early mental health intervention, yet existing multimodal approaches typically require large quantities of annotated data and extensive feature engineering, impeding their deployment in real‐world settings where labels are scarce. To address this challenge, we propose prompt‐based depression detection (PBDD), a novel prompt‐based few‐shot learning framework that leverages frozen pretrained language and vision models to identify depression indicators from paired text‐image posts without fine‐tuning. Our method begins with rigorous data cleaning and sampling to construct a high‐quality few‐shot dataset, then encodes text via a masked language model and images via a self‐supervised rotation‐prediction task to capture deep semantic cues. Multimodal representations are seamlessly fused into a unified prompt template containing a [MASK] token, enabling the pre‐trained model to infer depressive states by language completion. Extensive experiments on both large‐scale and 1 % few‐shot subsets demonstrate that PBDD consistently outperforms state‐of‐the‐art baselines, achieving significant gains in accuracy and Macro‐F1. These results validate the effectiveness and scalability of our framework for depression detection under severe label scarcity, offering a practical solution for real‐time mental health monitoring in social media environments. Rui Wang 0034, Heyang Feng, Erik Cambria, Kaize Shi, Xiaohan Yu 0001, Xuhui Fan 0001, Xianxun Zhu |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2026 | CIME: Contextual Interaction-Based Multimodal Emotion Analysis With Enhanced Semantic InformationabstractMultimodal emotion analysis is pivotal in decoding complex human affect by integrating diverse data sources such as text, audio, and visual signals. In this article, we introduce contextual interaction-based multimodal emotion analysis with enhanced semantic information (CIME), a novel spatio-temporal interaction network that significantly improves emotion recognition accuracy and robustness. CIME employs a text-centric cross-modal attention mechanism to refine semantic representations, while simultaneously leveraging a graph convolutional network to model contextual dialog information by capturing both intraspeaker and interspeaker relationships. This dual approach enables the effective fusion of modality-specific cues and the mining of latent emotional associations across modalities. Extensive experiments conducted on benchmark datasets—including IEMOCAP and MOSEI—demonstrate that CIME consistently outperforms existing state-of-the-art methods in terms of overall classification accuracy and weighted F1-scores. Furthermore, detailed ablation studies underscore the critical contributions of both the cross-modal attention and graph-based contextual modules. Rui Wang 0034, Chaopeng Guo, Mohammad Shabaz, Imad Rida, Erik Cambria, Xianxun Zhu |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Dynamic Spectral Graph Anomaly DetectionabstractGraph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distribution of graph spectral energy. Such spectrum-based methods rely on two steps: graph wavelet extraction and feature fusion. However, both steps are hand-designed, capturing incomprehensive anomaly information of wavelet-specific features and resulting in their inconsistent feature fusion. To address these problems, we propose a dynamic spectral graph anomaly detection framework DSGAD to adaptively capture comprehensive anomaly information and perform consistent feature fusion. DSGAD introduces dynamic wavelets, consisting of trainable wavelets to adaptively learn anomalous patterns and capture wavelet-specific features with comprehensive anomaly information. Furthermore, the consistent fusion of wavelet-specific features achieves dynamic fusion by combining wavelet-specific feature extraction with energy difference and channel convolution fusion using location correlation. Experimental results on four datasets substantiate the efficacy of our DSGAD method, surpassing state-of-the-art methods in both homogeneous and heterogeneous graphs. Jianbo Zheng, Chao Yang 0015, Tairui Zhang, Longbing Cao, Bin Jiang 0006, Xuhui Fan 0001, Xiao-Ming Wu 0002, Xianxun Zhu |
AAAI | 8 |
| 2025 | DNLN: Image super-resolution with Deformable Non-Local attention and Multi-Branch Weighted Feature Fusion
Dong Xing, Mohammad Shabaz, Yongpei Zhu, Xianxun Zhu |
Image Vis. Comput. | 6 |
| 2025 | Towards trustworthy image super-resolution via symmetrical and recursive artificial neural network
Mingliang Gao 0001, Jianhao Sun, Qilei Li, Muhammad Attique Khan, Jianrun Shang, Xianxun Zhu, Gwanggil Jeon |
Image Vis. Comput. | 6 |
| 2025 | Generalizable deepfake detection via Spatial Kernel Selection and Halo Attention Network
Siyou Guo, Qilei Li, Mingliang Gao 0001, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 4 |
| 2025 | Pyramidal attention with progressive multi-stage iterative feature refinement for salient object segmentation
Rahim Khan, Nada Alzaben, Yousef Ibrahim Daradkeh, Xianxun Zhu, Inam Ullah 0001 |
Image Vis. Comput. | 4 |
| 2025 | Deepfake detection via Feature Refinement and Enhancement Network
Weicheng Song, Siyou Guo, Mingliang Gao 0001, Qilei Li, Xianxun Zhu, Imad Rida |
Image Vis. Comput. | 5 |
| 2025 | HMPFormer: Hierarchical vision transformer with multi-perspective feature learning for precise polyp segmentation
Muhammad Talha Usman, Habib Khan, Haseeb Khan, Imad Rida, Xianxun Zhu, Jakeoung Koo |
Image Vis. Comput. | 5 |
| 2025 | A Dual-branch Progressive Network with spatial-frequency constraint for image fusion
Zenghui Wang 0011, Xuening Xing, Lina Liu 0009, Xianxun Zhu, Mingliang Gao 0001 |
Image Vis. Comput. | 5 |
| 2025 | A Geometric algebra-enhanced network for skin lesion detection with diagnostic prior
Ming Ju, Xianxun Zhu, Chunhua Qian, Rui Wang 0034 |
J. Supercomput. | 3 |
| 2024 | Emotion recognition based on brain-like multimodal hierarchical perception
Xianxun Zhu, Xiangyang Wang 0003, Rui Wang 0034 |
Multim. Tools Appl. | 1 |