Xianxun Zhu

dblp:259/6633 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-3958-7040ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Injecting image text structure and edge priors into segment anything for scene text segmentation
Qian Shao, Libo Weng, Yanjing Lei, Xianxun Zhu, Hui Chen 0026
Image Vis. Comput.5
2026 WiVi-UF: Unified feature learning in cross-modal transformers with WiFi and vision data fusion for enhanced human activity recognition
Xinhang Lin, Xianxun Zhu, Erik Cambria
Knowl. Based Syst.2
2026 DMSAA-SLAM: RGB-D SLAM for dynamic scenes via diffusion self-attention
abstract
• Design a self-attention aggregation module using a pre-trained diffusion model. • Integrate high-precision masks into RGB-D SLAM for robust dynamic tracking. • Validate superior accuracy and efficiency on dynamic simulation datasets. In dynamic environments, performing RGB-D SLAM (Simultaneous Localization and Mapping) faces significant challenges primarily due to the presence of moving objects. The motion of these objects can introduce tracking errors and inaccuracies in map construction, thereby compromising the stability and overall performance of the system. To maintain high-precision localization and mapping under such conditions, a SLAM system must effectively detect and handle dynamic objects. To address these challenges, this paper presents a novel RGB-D SLAM method, referred to as DMSAA-SLAM (Dynamic Scene SLAM Based on Diffusion Model Self-Attention Aggregation). The core idea is to leverage a pre-trained stable diffusion model, particularly its self-attention layers, to handle the complexity of dynamic scenes. By employing a multi-resolution aggregation approach, combined with iterative merging and nonmaximum suppression, the proposed method generates high-precision segmentation masks. These masks enable fine-grained segmentation of moving objects and effectively eliminate dynamic feature points, thereby mitigating the impact of dynamic elements on the SLAM process and ensuring efficient and accurate tracking and mapping.
Hui Chen 0026, Xianxun Zhu, Ling Fan
Pattern Recognit.5
2026 Advancing federated domain generalization in ophthalmology: Vision enhancement and consistency assurance for multicenter fundus image segmentation
Yang Zhao 0019, Xianxun Zhu, Jun Wang 0121, Yan Liu 0052
Pattern Recognit.4
2026 AOSNet-Sec: Aperture-orientation-spectrum fusion with statistical Markov repair for trustworthy super-resolution
Zhengnan Yin, Luwei Xiao, Xuan Feng 0002, Yiwei Chen 0002, Xianxun Zhu, Cai Luo, Faten S. Alamri, Rui Mao 0010, Erik Cambria
Pattern Recognit.5
2026 Exploring personalized federated learning from a distribution-based perspective
abstract
Personalized federated learning (PFL) is a promising technique for tackling data heterogeneity in federated learning systems. Recently, Bayesian neural networks (BNNs) have been introduced into the PFL framework to enable uncertainty quantification and improve performance in data-scarce settings. Despite these advantages, existing BNN-based PFL methods face two key challenges in practical applications. First, in real-world scenarios, client heterogeneity often arises in the form of group-wise variation, which cannot be adequately captured by a single shared distribution as assumed in prior work. Second, existing methods rely on deterministic or stochastic approximation techniques for posterior inference, which lead to substantial computational and memory overhead, hindering their scalability and deployment. To address these limitations, we propose DBFed, a novel BNN-based PFL framework from a distribution-based perspective. DBFed introduces group-specific distributions to better model the structural heterogeneity commonly observed in federated settings. Moreover, DBFed employs a rank-1 parameterization technique to map uncertainty from the weight space to a low-dimensional subspace, significantly reducing the computational and memory overhead. Theoretically, we establish the effectiveness of the rank-1 parameterization approach. Empirically, extensive experiments on diverse datasets demonstrate that DBFed consistently outperforms alternative PFL baselines in a heterogeneous setting.
Tianhao Yu, Kheng Cher Yeo, Sami Azam, Xiaohan Yu 0001, Hui Chen 0026, Xianxun Zhu
Pattern Recognit.6
2026 Uncertainty-aware multimodal affective data fusion for personalized mental health dialogue
abstract
Personalized affective dialogue systems are critical for mental health applications, where responses must be emotionally appropriate and tailored to individual users. However, most existing large language model (LLM) based approaches rely on deterministic personalization, ignore uncertainty in affective understanding, and are difficult to deploy under privacy constraints. In this paper, we propose PALLM, a personalized affective large language modeling framework designed for privacy-preserving mental health dialogue. PALLM decouples personalization into two complementary components: a deterministic personalization layer that captures stable user preferences, and a Bayesian affective representation layer that models dynamic emotional states and uncertainty. By restricting uncertainty modeling to affective representations rather than full LLM parameters, PALLM achieves efficient and scalable personalization under federated learning. Extensive experiments on EmpatheticDialogues and a real-world mental health conversation dataset show that PALLM improves affective alignment and robustness compared with non-personalized and partially personalized baselines.
Xianxun Zhu, Erik Cambria, Hui Chen 0026
Pattern Recognit.1
2026 FedBayesMamba: Uncertainty-aware federated learning for multimodal and audio-visual sequential modeling with selective state space models
abstract
Federated learning has emerged as an effective paradigm for training machine learning models across distributed clients without sharing raw data. In many real-world applications, sequential data are inherently multimodal, involving heterogeneous streams such as audio, visual, and temporal signals. However, most existing federated approaches rely on deterministic neural networks, which often struggle to capture predictive uncertainty under heterogeneous data distributions, cross-modal inconsistencies, and dynamic client participation. In this paper, we propose FedBayesMamba , a Bayesian federated learning framework for multimodal sequential data modeling based on selective state space models. The proposed approach introduces Bayesian parameterization into the Mamba architecture to enable uncertainty-aware sequence modeling while preserving the computational efficiency of state space models. To effectively integrate uncertainty across distributed clients, we further develop a posterior aggregation strategy that combines client-level posterior distributions in a principled probabilistic manner. Extensive experiments on multiple benchmark datasets demonstrate that the proposed framework achieves competitive predictive performance and improved uncertainty estimation under Non-IID federated settings. The results also indicate that FedBayesMamba exhibits strong robustness and stability in challenging federated scenarios. These findings highlight the potential of combining Bayesian learning with state space models for multimodal temporal modeling, particularly in audio-visual perception and cross-modal sequence understanding tasks.
Xianxun Zhu, Xiaosong E, Michele Nappi, Imad Rida, Hui Chen 0026
Pattern Recognit.1
2026 PBDD: A Prompt-Based Learning Approach for Few-Shot Social Media Depression Detection
abstract
Automated detection of depressive moods from social media holds great promise for early mental health intervention, yet existing multimodal approaches typically require large quantities of annotated data and extensive feature engineering, impeding their deployment in real‐world settings where labels are scarce. To address this challenge, we propose prompt‐based depression detection (PBDD), a novel prompt‐based few‐shot learning framework that leverages frozen pretrained language and vision models to identify depression indicators from paired text‐image posts without fine‐tuning. Our method begins with rigorous data cleaning and sampling to construct a high‐quality few‐shot dataset, then encodes text via a masked language model and images via a self‐supervised rotation‐prediction task to capture deep semantic cues. Multimodal representations are seamlessly fused into a unified prompt template containing a [MASK] token, enabling the pre‐trained model to infer depressive states by language completion. Extensive experiments on both large‐scale and 1 % few‐shot subsets demonstrate that PBDD consistently outperforms state‐of‐the‐art baselines, achieving significant gains in accuracy and Macro‐F1. These results validate the effectiveness and scalability of our framework for depression detection under severe label scarcity, offering a practical solution for real‐time mental health monitoring in social media environments.
Rui Wang 0034, Heyang Feng, Erik Cambria, Kaize Shi, Xiaohan Yu 0001, Xuhui Fan 0001, Xianxun Zhu
IEEE Trans. Comput. Soc. Syst.8
2026 CIME: Contextual Interaction-Based Multimodal Emotion Analysis With Enhanced Semantic Information
abstract
Multimodal emotion analysis is pivotal in decoding complex human affect by integrating diverse data sources such as text, audio, and visual signals. In this article, we introduce contextual interaction-based multimodal emotion analysis with enhanced semantic information (CIME), a novel spatio-temporal interaction network that significantly improves emotion recognition accuracy and robustness. CIME employs a text-centric cross-modal attention mechanism to refine semantic representations, while simultaneously leveraging a graph convolutional network to model contextual dialog information by capturing both intraspeaker and interspeaker relationships. This dual approach enables the effective fusion of modality-specific cues and the mining of latent emotional associations across modalities. Extensive experiments conducted on benchmark datasets—including IEMOCAP and MOSEI—demonstrate that CIME consistently outperforms existing state-of-the-art methods in terms of overall classification accuracy and weighted F1-scores. Furthermore, detailed ablation studies underscore the critical contributions of both the cross-modal attention and graph-based contextual modules.
Rui Wang 0034, Chaopeng Guo, Mohammad Shabaz, Imad Rida, Erik Cambria, Xianxun Zhu
IEEE Trans. Comput. Soc. Syst.6
2025 Dynamic Spectral Graph Anomaly Detection
abstract
Graph anomaly detection is crucial for identifying anomalous nodes within graphs and addressing applications like financial fraud detection and social spam detection. Recent spectral graph neural network methods advance graph anomaly detection by focusing on anomalies that notably affect the distribution of graph spectral energy. Such spectrum-based methods rely on two steps: graph wavelet extraction and feature fusion. However, both steps are hand-designed, capturing incomprehensive anomaly information of wavelet-specific features and resulting in their inconsistent feature fusion. To address these problems, we propose a dynamic spectral graph anomaly detection framework DSGAD to adaptively capture comprehensive anomaly information and perform consistent feature fusion. DSGAD introduces dynamic wavelets, consisting of trainable wavelets to adaptively learn anomalous patterns and capture wavelet-specific features with comprehensive anomaly information. Furthermore, the consistent fusion of wavelet-specific features achieves dynamic fusion by combining wavelet-specific feature extraction with energy difference and channel convolution fusion using location correlation. Experimental results on four datasets substantiate the efficacy of our DSGAD method, surpassing state-of-the-art methods in both homogeneous and heterogeneous graphs.
Jianbo Zheng, Chao Yang 0015, Tairui Zhang, Longbing Cao, Bin Jiang 0006, Xuhui Fan 0001, Xiao-Ming Wu 0002, Xianxun Zhu
AAAI8
2025 DNLN: Image super-resolution with Deformable Non-Local attention and Multi-Branch Weighted Feature Fusion
Dong Xing, Mohammad Shabaz, Yongpei Zhu, Xianxun Zhu
Image Vis. Comput.6
2025 Towards trustworthy image super-resolution via symmetrical and recursive artificial neural network
Mingliang Gao 0001, Jianhao Sun, Qilei Li, Muhammad Attique Khan, Jianrun Shang, Xianxun Zhu, Gwanggil Jeon
Image Vis. Comput.6
2025 Generalizable deepfake detection via Spatial Kernel Selection and Halo Attention Network
Siyou Guo, Qilei Li, Mingliang Gao 0001, Xianxun Zhu, Imad Rida
Image Vis. Comput.4
2025 Pyramidal attention with progressive multi-stage iterative feature refinement for salient object segmentation
Rahim Khan, Nada Alzaben, Yousef Ibrahim Daradkeh, Xianxun Zhu, Inam Ullah 0001
Image Vis. Comput.4
2025 Deepfake detection via Feature Refinement and Enhancement Network
Weicheng Song, Siyou Guo, Mingliang Gao 0001, Qilei Li, Xianxun Zhu, Imad Rida
Image Vis. Comput.5
2025 HMPFormer: Hierarchical vision transformer with multi-perspective feature learning for precise polyp segmentation
Muhammad Talha Usman, Habib Khan, Haseeb Khan, Imad Rida, Xianxun Zhu, Jakeoung Koo
Image Vis. Comput.5
2025 A Dual-branch Progressive Network with spatial-frequency constraint for image fusion
Zenghui Wang 0011, Xuening Xing, Lina Liu 0009, Xianxun Zhu, Mingliang Gao 0001
Image Vis. Comput.5
2025 A Geometric algebra-enhanced network for skin lesion detection with diagnostic prior
Ming Ju, Xianxun Zhu, Chunhua Qian, Rui Wang 0034
J. Supercomput.3
2024 Emotion recognition based on brain-like multimodal hierarchical perception
Xianxun Zhu, Xiangyang Wang 0003, Rui Wang 0034
Multim. Tools Appl.1