VLDB 2026 Research / reviewers in the wild / expert
Xiaohua Wang 0002
dblp:51/893-2
· DBLP profile ↗
21ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1751-2291ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EaNet: Enhanced Multimodal Awareness Alignment Network for Multimodal Aspect-Based Sentiment Analysis
Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Yan Xing 0002, Yiming Tang 0001, Jiaoyun Yang, Ning An 0001, Fuji Ren |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Beneficial Noise Learning for Robust Multimodal FusionabstractMultimodal fusion aims to integrate complementary information from different modalities, yet it is often hindered by data distribution imbalance and cross-modal conflicts. In this paper, we propose Beneficial Noise Learning for Robust Multimodal Fusion (BRMF), a novel framework that reframes these challenges as beneficial perturbations signals for robust training. Specifically, BRMF first learns a stable and consistent joint representation within a Gaussian latent space under distributional constraints. Based on this joint representation, it then constructs diverse distributional environments and further encourages invariance by extending invariant risk minimization (IRM) to multimodal settings. Extensive experiments across multiple benchmarks demonstrate that BRMF exhibits superior robustness to noise and establishes new state-of-the-art performance. Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Fuji Ren |
IEEE Trans. Multim. | 3 |
| 2025 | Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete DataabstractMultimodal Sentiment Analysis (MSA) with incomplete data has gained significant attention recently.Existing studies focus on optimizing model structures to handle modality missingness, but models still face challenges in robustness when dealing with uncertain missingness.To this end, we propose a data-centric robust multimodal sentiment analysis method, Proxy-Driven Robust Multimodal Fusion (P-RMF).First, we map unimodal data to the latent space of Gaussian distributions to capture core features and structure, thereby learn stable modality representation.Then, we combine the quantified modality intrinsic uncertainty to learn stable multimodal joint representation (i.e., proxy modality), which is further enhanced through multi-layer dynamic cross-modal injection to increase its diversity.Extensive experimental results show that P-RMF outperforms existing models in noise resistance and achieves state-of-the-art performance on multiple benchmark datasets. Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Jiaoyun Yang, Yiming Tang 0001, Ning An 0001 |
ACL (1) | 3 |
| 2025 | REKGC: Learning Re-coupled Representations for Knowledge Graph Completion
Min Hu 0010, Wenlong Fei, Xiaohua Wang 0002 |
DASFAA (3) | 4 |
| 2024 | Automatic Depression Detection Network Based on Facial Images and Facial Landmarks Feature FusionabstractArtificial intelligence methods offer objectivity and convenience in automatic depression detection, however, current research often neglects the critical role of facial landmarks. This oversight results in insufficient spatial structure information and a lack of detailed local representation, which fails to capture the nuanced semantic information crucial for identifying depression-related clues. To address these issues, we introduce a novel dual-branch network model comprising the Landmark-Image-Landmark Net (LIL Net) and the Global Context Vision Transformer Net (GCVit Net). Through a dual-stream, multiscale, and cross-fusion strategy, LIL Net is designed to extract original facial image features alongside landmark features, prioritizing the detailed semantic information of potential depression clues. LIL Net employs an innovative LIL Attention approach to jointly learn multiscale features from facial landmarks and images, thereby enhancing the model’s ability to capture fine-grained depression-related cues. Furthermore, the Multi-scale Feature Fusion (MSFF) module fuses the obtained multiscale features, augmenting the semantic expression of potential depression clues within facial landmarks via attention mechanisms. Meanwhile, the GCVit Net branch network supplements global information by extracting global facial features. Finally, the features from both branches are concatenated to enhance the accuracy of depression degree predictions. Experimental results demonstrate that our model has superior performance in detecting depression compared to existing methods. We release our code at https://github.com/xlx777/LIL-Net. Min Hu 0010, Lingxiang Xu, Xiaohua Wang 0002, Jiaoyun Yang |
BIBM | 4 |
| 2024 | MEVTR: A Multilingual Model Enhanced with Visual Text RepresentationsabstractThe goal of multilingual modelling is to generate multilingual text representations for various downstream tasks in different languages. However, some state-of-the-art pre-trained multilingual models perform poorly on many low-resource languages due to the lack of representation space and model capacity. To alleviate this issue, we propose a Multilingual model Enhanced with Visual Text Representations (MEVTR), which complements textual representations and extends the multilingual representation space with visual text representations. First, the visual encoder focuses on the glyphs and structure of the text to obtain visual text representations, and the textual encoder obtains textual representations. Then, multilingual representations are enhanced by aligning and fusing visual text representations and textual representations. Moreover, we propose similarity constraint, a self-supervised task to prompt the visual encoder to focus on more additional information. Prefix alignment and multi-head bilinear module are designed to acquire an improved integration effect of visual text representations and textual representations. Experimental results indicate that MEVTR benefits from visual text representations and achieves significant performance gains in downstream tasks. In particular, in the zero-shot cross-lingual transfer task, MEVTR achieves results that outperform the state-of-the-art adapter-based framework without the target language adapter. Xiaohua Wang 0002, Wenlong Fei, Min Hu 0010, Aoqiang Zhu |
LREC/COLING | 1 |
| 2024 | MTLS: Making Texts into Linguistic SymbolsabstractIn linguistics, all languages can be considered as symbolic systems, with each language relying on symbolic processes to associate specific symbols with meanings.In the same language, there is a fixed correspondence between linguistic symbol and meaning.In different languages, universal meanings follow varying rules of symbolization in one-to-one correspondence with symbols.Most work overlooks the properties of languages as symbol systems.In this paper, we shift the focus to the symbolic properties and introduce MTLS: a pre-training method to improve the multilingual capability of models by Making Texts into Linguistic Symbols.Initially, we replace the vocabulary in pre-trained language models by mapping relations between linguistic symbols and semantics.Subsequently, universal semantics within the symbolic system serve as bridges, linking symbols from different languages to the embedding space of the model, thereby enabling the model to process linguistic symbols.To evaluate the effectiveness of MTLS, we conducted experiments on multilingual tasks using BERT and RoBERTa, respectively, as the backbone.The results indicate that despite having just over 12,000 pieces of English data in pre-training, the improvement that MTLS brings to multilingual capabilities is remarkably significant. Wenlong Fei, Xiaohua Wang 0002, Min Hu 0010 |
EMNLP | 2 |
| 2024 | KEBR: Knowledge Enhanced Self-Supervised Balanced Representation for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) aims to integrate multiple modalities of information to better understand human sentiment. The current research mainly focuses on conducting multimodal fusion, which neglects the under-optimized modal representations generated by the imbalance of unimodal performances in joint learning. Moreover, the size of labeled datasets limits the generalization ability of existing supervised models. To address the above issues, this paper proposes a knowledge-enhanced self-supervised balanced representation approach (KEBR). First, a text-based cross-modal fusion method (TCMF) is constructed, which injects the non-verbal information from the videos into the semantic representation of text to enhance the multimodal representation of text. Then, a multimodal cosine constrained loss (MCC) is designed to constrain the fusion of non-verbal information in joint learning to balance the representation. Finally, with the help of sentiment knowledge and non-verbal information, KEBR conducts sentiment word masking and sentiment intensity prediction. Experimental results show that KEBR outperforms the baseline. Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Jiaoyun Yang, Yiming Tang 0001, Fuji Ren |
ACM Multimedia | 3 |
| 2024 | Dual-task enhanced global-local temporal-spatial network for depression recognition from facial videosabstractSummary In previous studies on facial video depression recognition, although convolutional neural network (CNN) has become a mainstream method, its performance still has room for improvement due to the insufficient extraction of global and local information and the neglect of the correlation of temporal and spatial information. This paper proposes a novel dual‐task enhanced global–local temporal–spatial network (DTE‐GLTS) to enhance the extraction capability of global and local features and deepen the analysis of temporal–spatial information correlation. We design a dual‐task learning mode that utilizes the data‐efficient image transformer (Deit) as the main body to learn the global features of video sequences and guides Deit to learn local features with the pre‐trained temporal–spatial fusion network (TSF). In addition, we propose the TSF mechanism to more effectively fuse temporal–spatial information in video sequences, strengthen the correlation between frames and pixels, and embed it in Resnet to form the TSF network. To the best of our knowledge, this is the first application of Deit and dual‐task learning mode in the field of facial video depression recognition. The experimental results on AVEC 2013 and AVEC 2014 show that our method achieves competitive performance, with mean absolute error/root mean square error (MAE/RMSE) scores of 6.06/7.73 and 5.91/7.68, respectively, while significantly reducing the number of parameters. Jinjie Shen, Yan Xing 0002, Min Hu 0010, Xiaohua Wang 0002, Daolun Li, Wen-shu Zha |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | Parallel Multiscale Bridge Fusion Network for Audio-Visual Automatic Depression AssessmentabstractDepression is a prevalent and severe mental illness that significantly impacts patients’ physical health and daily life. Recent studies have focused on multimodal depression assessment, aiming to objectively and conveniently evaluate depression using multimodal data. However, existing methods based on audio–visual modalities struggle to capture the dynamic variations in depression clues and cannot fully explore multimodal data over a long time. In addition, they rely heavily on insufficient single-stage multimodal fusion, which limits the accuracy of depression assessment. To address these limitations, we propose a novel parallel multiscale bridge fusion network (PMBFN) for audio–visual depression assessment. PMBFN comprehensively captures subtle multilevel dynamic changes in depression expression through parallel multiscale dynamic convolutions and long short-term memories (LSTMs) and effectively solves the problem of long-term audio–visual sequence information loss by using spatiotemporal attention pooling modules. Furthermore, the multimodal bridge fusion module is proposed in PMBFN to achieve multistage interactive recursive multimodal fusion, enhancing the expressive capacity of multimodal depression-related features to improve the accuracy of assessment. Extensive experiments on the DAIC-WOZ and E-DAIC datasets demonstrate that our method outperforms current state-of-the-art methods and clearly shows our method's effectiveness eventually. Min Hu 0010, Xiaohua Wang 0002, Yiming Tang 0001, Jiaoyun Yang, Ning An 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | A Multiscale Dynamic Temporal Convolution Network For Continuous Dimensional Emotion RecognitionabstractVideo-based dimensional emotion recognition, especially long sequence modeling, is a challenging task. This paper proposes a Multiscale Dynamic Temporal Convolution Network(MDTCN) for long sequence modeling of dimensional emotion recognition. Although Temporal Convolution Network(TCN) has performed well in sequence modeling, there are still some limitations that the feature scale of each layer is single and the receptive field of the upper layer is too large to capture the short-term dependence. Therefore, Dynamic Dilated Block(DDB) is constructed to better capture the temporal dependence by using dilated convolutions of different kernel size. In order to obtain both long-term and short-term temporal features, Multiscale TCN Structure is proposed. Extensive experiments are conducted on the AFFWILD2 and SEWA Dataset and the competitive result is shown compared to state-of-the-art methods. Min Hu 0010, Jialu Sun, Xiaohua Wang 0002, Ning An 0001 |
IJCNN | 3 |
| 2023 | Location Attention Knowledge Embedding Model for Image-Text Matching
Min Hu 0010, Xiaohua Wang 0002, Jiaoyun Yang, Nan Li 0065 |
PRCV (1) | 3 |
| 2023 | EERCA-ViT: Enhanced Effective Region and Context-Aware Vision Transformers for image sentiment analysis
Xiaohua Wang 0002, Min Hu 0010, Fuji Ren |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | A multi-scale framework based on jigsaw patches and focused label smoothing for bone age assessment
Xiaohua Wang 0002, Min Hu 0010, Fuji Ren |
Vis. Comput. | 1 |
| 2022 | A spatio-temporal integrated model based on local and global features for video expression recognition
Min Hu 0010, Peng Ge, Xiaohua Wang 0002, Fuji Ren |
Vis. Comput. | 3 |
| 2021 | A Two-Stage Spatiotemporal Attention Convolution Network for Continuous Dimensional Emotion Recognition From Facial VideoabstractContinuous dimensional emotion recognition for facial video sequence is a crucial and challenging task in Affective Computing and Human-Computer Intelligent Interaction. The key of this task is to effectively extract and discriminate spatial-temporal features in a more fine-grained way. In this paper, a Two-Stage Spatiotemporal Attention Temporal Convolution Network (TS-SATCN) is designed for continuous dimensional emotion recognition of facial videos. The first stage generates an initial recognition result that is later fed into the second for correction. In each stage, the introduced spatiotemporal attention branch helps the network learn different attention levels and focuses on the informative spatial-temporal features adaptively. The network is trained by a proposed smooth loss function which can further improve the predictions' quality. Extensive experiments are performed on two datasets, RECOLA and AFEW-VA, which shows that the proposed method achieves significant improvement over state-of-the-art methods. Min Hu 0010, Qian Chu, Xiaohua Wang 0002, Lei He 0002, Fuji Ren |
IEEE Signal Process. Lett. | 3 |
| 2019 | Two-Level Attention with Multi-task Learning for Facial Emotion Estimation
Xiaohua Wang 0002, Muzi Peng, Lijuan Pan, Min Hu 0010, Fuji Ren |
MMM (1) | 1 |
| 2019 | Video facial emotion recognition based on local enhanced motion history image and CNN-CTSLSTM networks
Min Hu 0010, Xiaohua Wang 0002, Juan Yang 0001, Ronggui Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Two-level attention with two-stage multi-task learning for facial emotion recognition
Xiaohua Wang 0002, Muzi Peng, Lijuan Pan, Min Hu 0010, Fuji Ren |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Facial Expression Recognition Based on Adaptive Weighted Fusion Histograms
Min Hu 0010, Yanxia Xu, Liangfeng Xu, Xiaohua Wang 0002 |
ICIC (2) | 4 |
| 2013 | Single Sample Face Recognition Based on Multiple Features and Twice Classification
Xiaohua Wang 0002, Min Hu 0010, Liangfeng Xu |
ICIC (1) | 1 |