Xianbing Zhao

dblp:321/6605 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-5482-3895ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CCAF: Coarse-to-fine Cross-Modal Alignment and Fusion for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) has witnessed remarkable advancements in recent years. Existing MSA methods focus primarily on learning coarse-grained representations from different modalities to perform global cross-modal alignment or fusion. However, these approaches often neglect fine-grained valuable sentimental clues derived from local cross-modal interactions. Furthermore, the cross-modal alignment and fusion of complex global and local cross-modal information pose significant challenges in MSA tasks. To address this issue, we propose a novel MSA framework that simultaneously captures coarse-grained and fine-grained cross-modal sentiment cues through global and local cross-modal alignment and fusion. Our approach consists of three key components: i) optimal transport-based global and local cross-modal alignment, which separately aligns valuable global and local sentiment clues across modalities, ii) global and local cross-modal gated attention, which respectively fuse the aligned global and local cross-modal representations, and iii) prototype-informed information bottleneck, which utilizes learnable sentiment prototypes and contrastive prototype match to eliminate redundant cross-modal information at both global and local levels. Extensive experiments conducted on two publicly available MSA datasets demonstrate the effectiveness and superiority of our proposed model.
Xianbing Zhao, Shengzun Yang, Buzhou Tang
WWW1
2026 AdaSAM-AD: Boosting SAM2 for fine-grained pixel-level anomaly detection via spatial-channel calibration and deformable cascades
Xianbing Zhao, Xinyang Yang, Wai Keung Wong, Chengliang Liu 0003
Pattern Recognit.1
2026 Toward Multimodal Sentiment Analysis via Contrastive Cross-Modal Retrieval Augmentation and Hierachical Prompts
abstract
Multimodal Sentiment Analysis (MSA) is a fundamental problem in the field of affective computing. Although significant progress has been made in cross-modal interaction, it remains a challenge due to the insufficient reference context in cross-modal interactions. Current cross-modal approaches primarily focus on leveraging modality-level reference context within a individual sample for cross-modal feature enhancement, neglecting the potential cross-sample relationships that can serve as sample-level reference context to enhance the cross-modal features. To address this issue, we propose a novel multimodal retrieval-augmented framework to simultaneously incorporate cross-sample modality-level reference context and cross-sample sample-level reference context to enhance the multimodal features. In particular, we first design a contrastive cross-modal retrieval module to retrieve semantic similar samples and enhance anchor modality. To endow the model to capture both cross-sample and intra-sample information, we integrate two different types of prompts, modality-level prompts and sample-level prompts, to generate modality-level and sample-level reference contexts, respectively. Finally, we design a cross-modal retrieval-augmented encoder that simultaneously leverages modality-level and sample-level reference contexts to enhance the anchor modality. Extensive experiments demonstrate the effectiveness and superiority of our model on two publicly available datasets.
Xianbing Zhao, Shengzun Yang, Buzhou Tang, Ronghuan Jiang
IEEE Trans. Affect. Comput.1
2026 MoDE: Improving Mixture of Depression Experts With Mutual Information Estimator for Depression Detection
abstract
The clinical interview dialogues is a critical approach in diagnosing depression. Existing methods have achieved impressive results on clinical depression interview datasets. However, they heavily rely on neural networks to automatically discover crucial question-answer pairs within clinical dialogues, lacking explicit modeling of depression factors present in clinical interviews. To fill this gap, we propose a novel mutual information-based mixture of depression experts, which explicitly analyzes depression factors within clinical interview dialogues and identify the contribution of individual and composite depression factors. Specifically, we first identify depression factors, such as social abilities, mental state, and medication history from a causal perspective. We design a Mixture of Depression Experts, consisting of multiple depression expert networks, each specialized in handling either individual or composite depression factors. In addition, we propose a mutual information-based gating function to enable dynamic depression diagnosis decisions conditioned on either individual or composite depression factors. Experiments conducted on publicly available datasets demonstrate the superiority and interpretability of our model.
Xianbing Zhao, Di Wang 0011, Buzhou Tang, Yefeng Zheng 0001
IEEE Trans. Affect. Comput.1
2025 Toward Robust Multimodal Sentiment Analysis using multimodal foundational models
Xianbing Zhao, Soujanya Poria, Buzhou Tang
Expert Syst. Appl.1
2025 Exploiting Prior Tacit Knowledge to Enhance Alignment and Verification in zero-shot video grounding
Jing Wang 0169, Xianbing Zhao, Xiaojie Wang 0006, Fangxiang Feng
Neurocomputing2
2024 Bidirectional Multimodal Block-Recurrent Transformers for Depression Detection
abstract
Depression, as a prevalent and severe psychological disorder, has become a burden to individuals, families and societies all over the world. Recently, some deep learning methods have been introduced for depression detection and achieved promising performance on a number of public datasets. Most of them rely on Long Short-Term Memory (LSTM) or Transformer to model multimodal time series data used for depression detection which fail in filtering noisy information within multiple modalities. Motivated by block-recurrent transformers, which has a strong ability to filter noisy information among single-modal time series, we propose novel block-recurrent transformers, called Bidirectional Multimodal Block-Recurrent Transformers (BMBRT), for multimodal data analysis and apply it to depression detection. BMBRT is extended from the block-recurrent transformers by introducing multimodal data enhancement module to obtain complementary information across modalities and designing a new multi-block transformer module for noise filtering. Experiments on three publicly available depression detection datasets show that our proposed method significantly outperforms current state-of-the-art methods.
Xiangyu Jia, Xianbing Zhao, Buzhou Tang, Ronghuan Jiang
BIBM2
2024 DKINet: Medication Recommendation via Domain Knowledge Informed Deep Learning
abstract
Medication recommendation is a fundamental yet crucial branch of healthcare that presents opportunities to assist physicians in making more accurate medication prescriptions for patients with complex health conditions. Previous studies have primarily concentrated on deriving patient representations from electronic health records (EHRs) to recommend medications, often overlooking the effective integration of domain-specific prior knowledge. However, integrating domain knowledge with the patient’s clinical manifestations can be challenging, particularly when dealing with complex clinical manifestations. Therefore, in this paper, we first identify comprehensive domain-specific prior knowledge, namely the Unified Medical Language System (UMLS), which is a comprehensive repository of biomedical vocabularies and standards, for knowledge extraction. Subsequently, we propose a knowledge injection module that addresses the effective integration of domain knowledge with complex clinical manifestations, enabling an effective characterization of the health conditions of the patient. Moreover, acknowledging the influence of historical medications on patients’ current treatments, we propose a historical medication-aware patient representation module to capture the longitudinal influence of historical medication information on the representation of current patients. Extensive experiments on three publicly benchmark datasets verify the superiority of our proposed method, which outperformed other methods by a significant margin. The code is available at: https://github.com/sherry6247/DKINet.
Sicen Liu, Xiaolong Wang 0001, Xianbing Zhao, Hao Chen 0011
BIBM3
2024 Learning in Order! A Sequential Strategy to Learn Invariant Features for Multimodal Sentiment Analysis
Xianbing Zhao, Lizhen Qu, Tao Feng 0013, Jianfei Cai 0001, Buzhou Tang
ACM Multimedia1
2023 TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment Analysis
abstract
Existing methods for Multimodal Sentiment Analysis (MSA) mainly focus on integrating multimodal data effectively on limited multimodal data. Learning more informative multimodal representation often relies on large-scale labeled datasets, which are difficult and unrealistic to obtain. To learn informative multimodal representation on limited labeled datasets as more as possible, we proposed TMMDA for MSA, a new Token Mixup Multimodal Data Augmentation, which first generates new virtual modalities from the mixed token-level representation of raw modalities, and then enhances the representation of raw modalities by utilizing the representation of the generated virtual modalities. To preserve semantics during virtual modality generation, we propose a novel cross-modal token mixup strategy based on the generative adversarial network. Extensive experiments on two benchmark datasets, i.e., CMU-MOSI and CMU-MOSEI, verify the superiority of our model compared with several state-of-the-art baselines. The code is available at https://github.com/xiaobaicaihhh/TMMDA.
Xianbing Zhao, Sicen Liu, Xuan Zang, Yang Xiang 0003, Buzhou Tang
WWW1
2023 Shared-Private Memory Networks For Multimodal Sentiment Analysis
abstract
Text, visual, and acoustic are usually complementary in the Multimodal Sentiment Analysis (MSA) task. However, current methods primarily concern shared representations while neglecting the critical private aspects of data within individual modalities. In this work, we propose shared-private memory networks based on the recent advances in the attention mechanism, called SPMN, to decouple multimodal representation from shared and private perspectives. It contains three components: a) a shared memory to learn the shared representations of multimodal data; b) three private memories to learn the private representations of individual modalities, respectively; c) and adaptive fusion gates to fuse multimodal private and shared representations. To evaluate the effectiveness of SPMN, we integrate it into different pre-trained language representation models, such as BERT and XLNET, and conduct experiments on two public datasets, CMU-MOSI and CMU-MOSEI. Experimental results indicate that the performances of pre-trained language representation models are significantly improved because of SPMN and demonstrate the superiority of our model compared to the state-of-the-art methods. SPMN's source code is publicly available at:https://github.com/xiaobaicaihhh/SPMN.
Xianbing Zhao, Yinxin Chen, Sicen Liu, Buzhou Tang
IEEE Trans. Affect. Comput.1
2023 SHAPE: A Sample-Adaptive Hierarchical Prediction Network for Medication Recommendation
abstract
Effectively medication recommendation with complex multimorbidity conditions is a critical yet challenging task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the encoding format of intra-visit medical events are serialized and information transmitted patterns of learning longitudinal sequence data are stable. However, the following conditions may have been ignored: 1) A more compact encoder for intra-relationship in the intra-visit medical event is urgent; 2) Strategies for learning accurate representations of the variable longitudinal sequences of patients are different. In this article, we proposed a novel Sample-adaptive Hierarchical medicAtion Prediction nEtwork, termed SHAPE, to tackle the above challenges in the medication recommendation task. Specifically, we design a compact intra-visit set encoder to encode the relationship in the medical event for obtaining visit-level representation and then develop an inter-visit longitudinal encoder to learn the patient-level longitudinal representation efficiently. To endow the model with the capability of modeling the variable visit length, we introduce a soft curriculum learning method to assign the difficulty of each sample automatically by the visit length. Extensive experiments on a benchmark dataset verify the superiority of our model compared with several state-of-the-art baselines.
Sicen Liu, Xiaolong Wang 0001, Jingcheng Du, Yongshuai Hou, Xianbing Zhao, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang
IEEE J. Biomed. Health Informatics5
2022 MAG+: An Extended Multimodal Adaptation Gate for Multimodal Sentiment Analysis
abstract
Human multimodal sentiment analysis is a challenging task that devotes to extract and integrate information from multiple resources, such as language, acoustic and visual information. Recently, multimodal adaptation gate (MAG), an attachment to transformer-based pre-trained language representation models, such as BERT and XLNet, has shown state-of-the-art performance on multimodal sentiment analysis. MAG only uses a 1-layer network to fuse multimodal information directly, and does not pay attention to relationships among different modalities. In this paper, we propose an extended MAG, called MAG+, to reinforce multimodal fusion. MAG+ contains two modules: multi-layer MAGs with modality reinforcement (M3R) and Adaptive Layer Aggregation (ALA). In the MAG with modality reinforcement of M3R, each modality is reinforced by all other modalities via crossmodal attention at first, and then all modalities are fused via MAG. The ALA module leverages the multimodal representations at low and high levels as the final multimodal representation. Similar to MAG, MAG+ is also attached to BERT and XLNet. Experimental results on two widely used datasets demonstrate the efficacy of our proposed MAG+.
Xianbing Zhao, Lei Gao 0007, Buzhou Tang
ICASSP1
2022 HMAI-BERT: Hierarchical Multimodal Alignment and Interaction Network-Enhanced BERT for Multimodal Sentiment Analysis
abstract
Human language is multimodal, including textual, visual and acoustic information. The task of multimodal sentiment analysis is to use human multimodal information for sentiment recognition. Among the three modalities, text contains richer information than other modalities. With the development of pre-trained representation models on text, most of multimodal sentiment analysis methods use text as primary information and the other modalities as supplementary information. The existing methods suffer from the following limitations: 1) inherent heterogeneity of multimodal data, which makes multimodal fusion difficult as different modalities reside in different feature spaces; 2) asynchronism caused by the inconsistent sampling rates of the time series data of different modalities. To alleviate the heterogeneity and asynchronism of multimodal data, we propose HMAI-BERT, a hierarchical multimodal alignment and interaction network-enhanced BERT. In HMAI-BERT, to improve the efficiency of multimodal interaction, we introduce a memory network to align the different multimodal representations before fusion. After multimodal alignment, we propose a modal update method to address the problem of asynchronism, where each modality is reinforced by interacting with other modalities. In addition, we introduce a fusion module to integrate the three reinforced modalities, and a sentiment enhanced memory to enhance multimodal representation. Our experiments on two public datasets show that the proposed HMAI-BERT outperforms the state-of-the-art methods.
Xianbing Zhao, Yiting Chen 0010, Sicen Liu, Buzhou Tang
ICME1
2022 Boosting lesion annotation via aggregating explicit relations in external medical knowledge graph
Xianbing Zhao, Buzhou Tang
Artif. Intell. Medicine2