VLDB 2026 Research / reviewers in the wild / expert
Xiaoming Zhao 0002
dblp:64/3046-2
· DBLP profile ↗
22ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4708-4171ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-stage multiple instance learning networks with attention-based hybrid aggregation for speech emotion recognition
Shiqing Zhang, Xiaoming Zhao 0002 |
Comput. Speech Lang. | 5 |
| 2025 | Distribution-Aware Multi-Attention Tri-branch Networks with Feedforward Differential Features for semi-supervised medical image segmentation
Peilian Shi, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002 |
Expert Syst. Appl. | 6 |
| 2025 | PACMR: Progressive Adaptive Crossmodal Reinforcement for Multimodal Apparent Personality Traits AnalysisabstractMultimodal apparent personality traits analysis is a challenging issue due to the asynchrony among modalities. To address this issue, this paper proposes a Progressive Adaptive Crossmodal Reinforcement (PACMR) approach for multimodal apparent personality traits analysis. PACMR adopts a progressive reinforcement strategy to provide a multi-level information exchange among different modalities for crossmodal interactions, resulting in reinforcing the source and target modalities simultaneously. Specifically, PACMR introduces an Adaptive Modality Reinforcement Unit (AMRU) to adaptively adjust the weights of self-attention and crossmodal attention for capturing reliable contextual dependencies of multimodal sequence data. Experiment results on the public First Impressions dataset demonstrate the effectiveness of the proposed method. Shiqing Zhang, Xiaoming Zhao 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | MDKAT: Multimodal Decoupling With Knowledge Aggregation and Transfer for Video Emotion RecognitionabstractMultimodal Emotion Recognition (MER) leverages multiple input signals to identify the expressed emotions in user-generated data. Currently, effectively addressing both modality heterogeneity and homogeneity on MER tasks is a challenging issue due to the diversity of multimodal inputs in videos. To address this issue, this work proposes an efficient Multimodal Decoupling Method with Knowledge Aggregation and Transfer (MDKAT) for robust multimodal feature learning in emotional videos. MDKAT is consisted of three key steps: modality-independent feature extraction, modality-specific feature extraction, and multi-loss integration for decoupling. In these three steps, four crucial modules are individually designed to improve different aspects of multimodal learning on MER tasks, including a Cross-modal Feature Fusion (CFF) module for enhancing modality-independent features, an Adaptive Masked Self-Attention (AMSA) module for feature refinement, a Knowledge Aggregation (KA) module for ensuring the semantic similarity of modality-independent features, and a Knowledge Transfer (KT) module for balancing the strengths of different modalities. Experimental results on the typical CMU-MOSI and CMU-MOSEI datasets show that MDKAT obtains superior performance over state-of-the-art methods, demonstrating the effectiveness of MDKAT on MER tasks. Jian Wang 0066, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Yaowei Wang 0001, Yi Yang 0001, Siwei Ma 0001, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Multi-Scale Group Agent Attention-Based Graph Convolutional Decoding Networks for 2D Medical Image SegmentationabstractAutomated medical image segmentation plays a crucial role in assisting doctors in diagnosing diseases. Feature decoding is a critical yet challenging issue for medical image segmentation. To address this issue, this work proposes a novel feature decoding network, called multi-scale group agent attention-based graph convolutional decoding networks (MSGAA-GCDN), to learn local-global features in graph structures for 2D medical image segmentation. The proposed MSGAA-GCDN combines graph convolutional network (GCN) and a lightweight multi-scale group agent attention (MSGAA) mechanism to represent features globally and locally within a graph structure. Moreover, in skip connections a simple yet efficient attention-based upsampling convolution fusion (AUCF) module is designed to enhance encoder-decoder feature fusion in both channel and spatial dimensions. Extensive experiments are conducted on three typical medical image segmentation tasks, namely Synapse abdominal multi-organs, Cardiac organs, and Polyp lesions. Experimental results demonstrate that the proposed MSGAA-GCDN outperforms the state-of-the-art methods, and the designed MSGAA is a lightweight yet effective attention architecture. The proposed MSGAA-GCDN can be easily taken as a plug-and-play decoder cascaded with other encoders for general medical image segmentation tasks. Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002, Qi Tian 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Integrating Representation Subspace Mapping with Unimodal Auxiliary Loss for Attention-based Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) aims to identify emotions by utilizing affective information from multiple modalities. Due to the inherent disparities among these heterogeneous modalities, there is a large modality gap in their representations, leading to the challenge of fusing multiple modalities for MER. To address this issue, this work proposes a novel attention-based MER framework by integrating representation subspace mapping with unimodal auxiliary loss for enhancing multimodal fusion capabilities. Initially, a representation subspace mapping module is proposed to map each modality into two distinct subspaces. One is modality-public, enabling the acquisition of common representations and reducing the discrepancies across modalities. The other is modality-unique, retaining the unique characteristics of each modality while eliminating redundant inter-modal attributes. Then, a cross-modality attention is leveraged to bridge the modality gap in unique representations and facilitate modality adaptation. Additionally, our method designs an unimodal auxiliary loss to remove the noise unrelated to emotion classification, resulting in robust and meaningful representations for MER. Comprehensive experiments are conducted on the IEMOCAP and MSP-Improv datasets, and experiment results show that our method achieves superior performance to state-of-the-art MER methods. Keywords: Multimodal emotion recognition, representation subspace mapping, cross-modality attention, unimodal auxiliary loss, fusion Xulong Du, Xingnan Zhang, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Liangliang Lou |
LREC/COLING | 7 |
| 2024 | Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects
Shiqing Zhang, Yijiao Yang, Xingnan Zhang, Qingming Leng, Xiaoming Zhao 0002 |
Expert Syst. Appl. | 6 |
| 2024 | Wireless-Sensing-Based Human-Vehicle Classification Method via Deep Learning: Analysis and ImplementationabstractWireless sensing methods for human-vehicle classification (WHVC) offer cost-effective advantages and enhance the detection efficiency of traffic parameters in intelligent transportation systems (ITSs). Existing WHVC methods primarily utilize channel state information (CSI) or received signal strength (RSS) features extracted from the surrounding wireless signals. Although CSI data provides more detailed and accurate channel information compared to RSS data, extracting and processing CSI is more challenging than RSS. Moreover, for applications that do not require fine-grained human-vehicle classification, such as intelligent street lighting systems, RSS-based WHVC has the advantages of easy implementation and low cost. Therefore, investigating the performance of CSI-and RSS-based WHVC methods in different application scenarios could provide valuable insights for the WHVC domain. To address this issue, this paper proposes a deep learning-based WHVC method, which employs deep learning as a tool to evaluate the performance of RSS and CSI methods in various classification tasks. Specifically, this paper collects CSI and RSS data for seven different classification tasks in real traffic road scenarios and evaluates these tasks using a convolutional neural network-based deep learning model designed in this paper. Experimental results demonstrate that for road user categories less than four, RSS-based WHVC achieves higher accuracy than CSI-based WHVC. However, as the number of categories increases, CSI-based WHVC exhibits superior accuracy compared to RSS-based WHVC. Additionally, the developed dataset is publicly available at https://github.com/TZ-mx/mixeddataset. Liangliang Lou, Mingxin Song, Xiaoming Zhao 0002, Shiqing Zhang, Ming Zhan |
IEEE Internet Things J. | 4 |
| 2024 | MTDAN: A Lightweight Multi-Scale Temporal Difference Attention Networks for Automated Video Depression DetectionabstractDeep learning based video depression analysis has been recently an interesting and challenging topic. Most of existing works focus on learning single-scale facial dynamics of participants for depression detection. Besides, they usually adopt expensive deep learning models with high computational complexity, resulting in difficulty in real-time clinical applications. To address these two issues, this work proposes a lightweight Multi-scale Temporal Difference Attention Networks (MTDAN) integrating the temporal difference and attention mechanism to model both short-term and long-term temporal facial behaviors for automated video depression detection. Initially, two simple yet effective sub-branches, i.e., a Short-term Temporal Difference Attention Network (ST-TDAN), and a Long-term Temporal Difference Attention Network (LT-TDAN), are designed to perform individually short-term and long-term depressive behavior modeling. Then, a simple Interactive Multi-head Attention Fusion (IMHAF) strategy is employed for integrating short-term and long-term spatiotemporal features, followed by a linear fully-collected layer for depression score prediction. Experiments on two public AVEC2013 and AVEC2014 datasets show that our proposed method not only achieves highly competitive performance to state-of-the-art methods, but also has much smaller computational complexity than them on video depression detection tasks. Shiqing Zhang, Xingnan Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Mingyue Niu, Ziping Zhao 0001, Jun Yu 0002, Qi Tian 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Optimized Wireless Sensing and Deep Learning for Enhanced Human-Vehicle RecognitionabstractIn the realm of traffic parameter measurement, wireless sensing-based human-vehicle recognition methods have been pivotal due to their low cost and non-invasive nature. Traditionally, these methods have relied on the 2.4 GHz frequency band, often neglecting the rich potential of the sub-GHz bands. Furthermore, the energy attributes of wireless signals are influenced by both antenna height and carrier frequency, yet few studies have explored their impact on human-vehicle recognition performance. Addressing this critical research gap, this study introduces an innovative convolutional neural network-based method that leverages both sub-GHz bands and variable antenna heights. Specifically, this paper focuses on two key aspects: received signal strength signal-to-noise ratio analysis and wireless sensing-based human-vehicle recognition performance analysis. Experimental results demonstrate that the optimal human-vehicle recognition performance is achieved with 2.4 GHz wireless signals and an antenna height of 0.8 m, resulting in an average vehicle recognition accuracy of 95.96%. Besides, the dataset with various carrier frequencies and antenna heights has been publicly available athttps://github.com/TZ-mx/mixed-RSS-dataset. Liangliang Lou, Mingxin Song, Xiaoming Zhao 0002, Shiqing Zhang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Deep learning-based panoptic segmentation: Recent advances and perspectivesabstractAbstract In recent years, panoptic segmentation has drawn increasing amounts of attention, leading to the rapid emergence of numerous related algorithms. A variety of deep neural networks have been used more frequently for panoptic segmentation, which is motivated by the significant success of deep learning methods in other tasks. This article presents a comprehensive exploration of panoptic segmentation, focusing on the analysis and understanding of RGB image data. Initially, the authors introduce the background of panoptic segmentation, including deep learning models and image segmentation. Then, the authors thoroughly cover a variety of panoptic segmentation‐related topics, such as datasets connected to the field, evaluation metrics, panoptic segmentation models, and derived subfields based on panoptic segmentation. Finally, the authors examine the difficulties and possibilities in this area and identify its future paths. Yuelong Chuang, Shiqing Zhang, Xiaoming Zhao 0002 |
IET Image Process. | 3 |
| 2023 | Learning inter-class optical flow difference using generative adversarial networks for facial expression recognitionabstractAbstract Facial expression recognition is a fine-grained task because different emotions have subtle facial movements. This paper proposes to learn inter-class optical flow difference using generative adversarial networks (GANs) for facial expression recognition. Initially, the proposed method employs a GAN to produce inter-class optical flow images from the difference between the static fully expressive samples and neutral expression samples. Such inter-class optical flow difference is used to highlight the displacement of facial parts between the neutral facial images and fully expressive facial images, which can avoid the disadvantage that the optical flow change between adjacent frames of the same video expression image is not obvious. Then, the proposed method designs four-channel convolutional neural networks (CNNs) to learn high-level optical flow features from the produced inter-class optical flow images, and high-level static appearance features from the fully expressive facial images, respectively. Finally, a decision-level fusion strategy is adopted to implement facial expression classification. The proposed method is validated on two public facial expression databases, BAUM_1a, SAMM and AFEW5.0, demonstrating its promising performance. Wenping Guo, Xiaoming Zhao 0002, Shiqing Zhang, Xianzhang Pan |
Multim. Tools Appl. | 2 |
| 2022 | Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion RecognitionabstractThis paper focuses on an interesting task, i.e., unsupervised cross-corpus Speech Emotion Recognition (SER), in which the labelled training (source) corpus and the unlabelled testing (target) corpus have different feature distributions, resulting in the discrepancy between the source and target domains. To address this issue, this paper proposes an unsupervised domain adaptation method integrating Transformers and Mutual Information (MI) for cross-corpus SER. Initially, our method employs encoder layers of Transformers to capture long-term temporal dynamics in an utterance from the extracted segment-level log-Mel spectrogram features, thereby producing the corresponding utterance-level features for each utterance in two domains. Then, we propose an unsupervised feature decomposition method with a hybrid Max-Min MI strategy to separately learn domain-invariant features and domain-specific features from the extracted mixed utterance-level features, in which the discrepancy between two domains is eliminated as much as possible and meanwhile their individual characteristic is preserved. Finally, an interactive Multi-Head attention fusion strategy is designed to learn the complementarity between domain-invariant features and domain-specific features so that they can be interactively fused for SER. Extensive experiments on the IEMOCAP and MSP-Improv datasets demonstrate the effectiveness of our proposed method on unsupervised cross-corpus SER tasks, outperforming state-of-the-art unsupervised cross-corpus SER methods. Shiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao 0002, Jun Yu 0002 |
ACM Multimedia | 4 |
| 2022 | Spontaneous Speech Emotion Recognition Using Multiscale Deep Convolutional LSTMabstractRecently, emotion recognition in real sceneries such as in the wild has attracted extensive attention in affective computing, because existing spontaneous emotions in real sceneries are more challenging and difficult to identify than other emotions. Motivated by the diverse effects of different lengths of audio spectrograms on emotion identification, this paper proposes a multiscale deep convolutional long short-term memory (LSTM) framework for spontaneous speech emotion recognition. Initially, a deep convolutional neural network (CNN) model is used to learn deep segment-level features on the basis of the created image-like three channels of spectrograms. Then, a deep LSTM model is adopted on the basis of the learned segment-level CNN features to capture the temporal dependency among all divided segments in an utterance for utterance-level emotion recognition. Finally, different emotion recognition results, obtained by combining CNN with LSTM at multiple lengths of segment-level spectrograms, are integrated by using a score-level fusion strategy. Experimental results on two challenging spontaneous emotional datasets, i.e., the AFEW5.0 and BAUM-1s databases, demonstrate the promising performance of the proposed method, outperforming state-of-the-art methods. Shiqing Zhang, Xiaoming Zhao 0002, Qi Tian 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Learning deep multimodal affective features for spontaneous speech emotion recognition
Shiqing Zhang, Yuelong Chuang, Xiaoming Zhao 0002 |
Speech Commun. | 4 |
| 2016 | Biologically Inspired Pattern Recognition for E-nose Sensors
Sanad Al-Maskari, Wenping Guo, Xiaoming Zhao 0002 |
ADMA | 3 |
| 2016 | Graph Regularized Sparsity Discriminant Analysis for face recognition
Songjiang Lou, Xiaoming Zhao 0002, Yuelong Chuang, Shiqing Zhang |
Neurocomputing | 2 |
| 2015 | Spoken emotion recognition via locality-constrained kernel sparse representation
Xiaoming Zhao 0002, Shiqing Zhang |
Neural Comput. Appl. | 1 |
| 2014 | Locality-sensitive kernel sparse representation classification for face recognition
Shiqing Zhang, Xiaoming Zhao 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Robust emotion recognition in noisy speech via sparse representation
Xiaoming Zhao 0002, Shiqing Zhang, Bicheng Lei |
Neural Comput. Appl. | 1 |
| 2013 | Dimensionality reduction-based spoken emotion recognition
Shiqing Zhang, Xiaoming Zhao 0002 |
Multim. Tools Appl. | 2 |
| 2012 | Phoneme recognition using an adaptive supervised manifold learning algorithm
Xiaoming Zhao 0002, Shiqing Zhang |
Neural Comput. Appl. | 1 |