VLDB 2026 Research / reviewers in the wild / expert
Kai Gao 0006
dblp:12/4000-6
· DBLP profile ↗
26ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent RecognitionabstractMultimodal Intent Recognition (MIR) aims to understand complex user intentions by leveraging text, video, and audio signals. However, existing approaches face two key challenges: (1) overlooking intricate cross-modal interactions for distinguishing consistent and inconsistent cues, and (2) ineffectively modeling multimodal conflicts, leading to semantic cancellation. To address these, we propose a novel Cognitive Dual-Pathway Reasoning (CDPR) framework, which constructs a stable semantic foundation via the intuition pathway and mitigates high-level semantic conflicts through the reasoning pathway, cooperatively establishing deep semantic relations. Specifically, we first employ a representation disentanglement strategy to extract modality-invariant and specific features. Subsequently, the intuition pathway aggregates cross-modal consensus using shared features for solid global representations. The reasoning pathway introduces an inconsistency perception mechanism, combining semantic prototype matching with statistical probability calibration to precisely quantify conflict severity, and dynamically adjusting the weights between both pathways. Furthermore, a multi-view loss function is adopted to alleviate modality laziness and learn structured features at different stages. Extensive experiments on two benchmarks show that CDPR achieves SOTA performance and superior robustness in mitigating multimodal inconsistency. The code is available at https://github.com/Hebust-NLP/CDPR. Peiwu Wang, Yunxian Chi, Zhinan Gou, Kai Gao 0006 |
ICMR | 5 |
| 2025 | TG-ERC: Utilizing three generation models to handle emotion recognition in conversation tasks
Zhinan Gou, Yuchen Long, Jieli Sun, Kai Gao 0006 |
Expert Syst. Appl. | 4 |
| 2025 | Multimodal Classification and Out-of-Distribution Detection for Multimodal Intent Understanding
Hanlei Zhang, Qianrui Zhou, Hua Xu 0003, Jianhua Su, Roberto Evans, Kai Gao 0006 |
IEEE Trans. Multim. | 6 |
| 2024 | Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionabstractMultimodal intent recognition aims to leverage diverse modalities such as expressions, body movements and tone of speech to comprehend user's intent, constituting a critical task for understanding human language and behavior in real-world multimodal scenarios. Nevertheless, the majority of existing methods ignore potential correlations among different modalities and own limitations in effectively learning semantic features from nonverbal modalities. In this paper, we introduce a token-level contrastive learning method with modality-aware prompting (TCL-MAP) to address the above challenges. To establish an optimal multimodal semantic environment for text modality, we develop a modality-aware prompting module (MAP), which effectively aligns and fuses features from text, video and audio modalities with similarity-based modality alignment and cross-modality attention mechanism. Based on the modality-aware prompt and ground truth labels, the proposed token-level contrastive learning framework (TCL) constructs augmented samples and employs NT-Xent loss on the label token. Specifically, TCL capitalizes on the optimal textual semantic insights derived from intent labels to guide the learning processes of other modalities in return. Extensive experiments show that our method achieves remarkable improvements compared to state-of-the-art methods. Additionally, ablation analyses demonstrate the superiority of the modality-aware prompt over the handcrafted prompt, which holds substantial significance for multimodal prompt learning. The codes are released at https://github.com/thuiar/TCL-MAP. Qianrui Zhou, Hua Xu 0003, Hanlei Zhang, Kai Gao 0006 |
AAAI | 7 |
| 2024 | Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal UtterancesabstractDiscovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.Existing methods manifest limitations in leveraging nonverbal information for discerning complex semantics in unsupervised scenarios.This paper introduces a novel unsupervised multimodal clustering method (UMC), making a pioneering contribution to this field.UMC introduces a unique approach to constructing augmentation views for multimodal data, which are then used to perform pre-training to establish well-initialized representations for subsequent clustering.An innovative strategy is proposed to dynamically select high-quality samples as guidance for representation learning, gauged by the density of each sample's nearest neighbors.Besides, it is equipped to automatically determine the optimal value for the top-K parameter in each cluster to refine sample selection.Finally, both high-and low-quality samples are used to learn representations conducive to effective clustering.We build baselines on benchmark multimodal intent and dialogue act datasets.UMC shows remarkable improvements of 2-6% scores in clustering metrics over state-of-the-art methods, marking the first successful endeavor in this domain.The complete code and data are available at https://github.com/thuiar/UMC. Hanlei Zhang, Hua Xu 0003, Xin Wang 0220, Kai Gao 0006 |
ACL (1) | 5 |
| 2024 | MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in ConversationsabstractMultimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. However, most existing multimodal intent benchmark datasets are limited in scale and suffer from difficulties in handling out-of-scope samples that arise in multi-turn conversational interactions. In this paper, we introduce MIntRec2.0, a large-scale benchmark dataset for multimodal intent recognition in multi-party conversations. It contains 1,245 high-quality dialogues with 15,040 samples, each annotated within a new intent taxonomy of 30 fine-grained classes, across text, video, and audio modalities. In addition to more than 9,300 in-scope samples, it also includes over 5,700 out-of-scope samples appearing in multi-turn contexts, which naturally occur in real-world open scenarios, enhancing its practical applicability. Furthermore, we provide comprehensive information on the speakers in each utterance, enriching its utility for multi-party conversational research. We establish a general framework supporting the organization of single-turn and multi-turn dialogue data, modality feature extraction, multimodal fusion, as well as in-scope classification and out-of-scope detection. Evaluation benchmarks are built using classic multimodal fusion methods, ChatGPT, and human evaluators. While existing methods incorporating nonverbal information yield improvements, effectively leveraging context information and detecting out-of-scope samples remains a substantial challenge. Notably, powerful large language models exhibit a significant performance gap compared to humans, highlighting the limitations of machine learning methods in the advanced cognitive intent understanding task. We believe that MIntRec2.0 will serve as a valuable resource, providing a pioneering foundation for research in human-machine conversational interactions, and significantly facilitating related applications.
The full dataset and codes are available for use at https://github.com/thuiar/MIntRec2.0. Hanlei Zhang, Xin Wang 0220, Hua Xu 0003, Qianrui Zhou, Kai Gao 0006, Jianhua Su, Jinyue Zhao |
ICLR | 5 |
| 2024 | Multimodal Consistency-Based Teacher for Semi-Supervised Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis holds significant importance within the realm of human-computer interaction. Due to the ease of collecting unlabeled online resources compared to the high costs associated with annotation, it becomes imperative for researchers to develop semi-supervised methods that leverage unlabeled data to enhance model performance. Existing semi-supervised approaches, particularly those applied to trivial image classification tasks, are not suitable for multimodal regression tasks due to their reliance on task-specific augmentation and thresholds designed for classification tasks. To address this limitation, we propose the Multimodal Consistency-based Teacher (MC-Teacher), which incorporates consistency-based pseudo-label technique into semi-supervised multimodal sentiment analysis. In our approach, we first propose synergistic consistency assumption which focus on the consistency among bimodal representation. Building upon this assumption, we develop a learnable filter network that autonomously learns how to identify misleading instances instead of threshold-based methods. This is achieved by leveraging both the implicit discriminant consistency on unlabeled instances and the explicit guidance on constructed training data with labeled instances. Additionally, we design the self-adaptive exponential moving average strategy to decouple the student and teacher networks, utilizing a heuristic momentum coefficient. Through both quantitative and qualitative experiments on two benchmark datasets, we demonstrate the outstanding performances of the proposed MC-Teacher approach. Furthermore, detailed analysis experiments and case studies are provided for each crucial component to intuitively elucidate the inner mechanism and further validate their effectiveness. Jingliang Fang, Hua Xu 0003, Kai Gao 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | A Clustering Framework for Unsupervised and Semi-Supervised New Intent DiscoveryabstractNew intent discovery is of great value to natural language processing, allowing for a better understanding of user needs and providing friendly services. However, most existing methods struggle to capture the complicated semantics of discrete text representations when limited or no prior knowledge of labeled data is available. To tackle this problem, we propose a novel clustering framework, USNID, forunsupervised andsemi-supervisednewintentdiscovery, which has three key technologies. First, it fully utilizes of unsupervised or semi-supervised data to mine shallow semantic similarity relations and provide well-initialized representations for clustering. Second, it designs a centroid-guided clustering mechanism to address the issue of cluster allocation inconsistency and provide high-quality self-supervised targets for representation learning. Third, it captures high-level semantics in unsupervised or semi-supervised data to discover fine-grained intent-wise clusters by optimizing both cluster-level and instance-level objectives. We also propose an effective method for estimating the cluster number in open-world scenarios without knowing the number of new intents beforehand. USNID performs exceptionally well on several benchmark intent datasets, achieving new state-of-the-art results in unsupervised and semi-supervised new intent discovery and demonstrating robust performance with different cluster numbers. Hanlei Zhang, Hua Xu 0003, Xin Wang 0220, Kai Gao 0006 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Noise Imitation Based Adversarial Training for Robust Multimodal Sentiment AnalysisabstractAs an inevitable phenomenon in real-world applications, data imperfection has emerged as one of the most critical challenges for multimodal sentiment analysis. However, existing approaches tend to overly focus on a specific type of imperfection, leading to performance degradation in real-world scenarios where multiple types of noise exist simultaneously. In this work, we formulate the imperfection with the modality feature missing at the training period and propose the noise intimation based adversarial training framework to improve the robustness against various potential imperfections at the inference period. Specifically, the proposed method first uses temporal feature erasing as the augmentation for noisy instances construction and exploits the modality interactions through the self-attention mechanism to learn multimodal representation for original-noisy instance pairs. Then, based on paired intermediate representation, a novel adversarial training strategy with semantic reconstruction supervision is proposed to learn unified joint representation between noisy and perfect data. For experiments, the proposed method is first verified with the modality feature missing, the same type of imperfection as the training period, and shows impressive performance. Moreover, we show that our approach is capable of achieving outstanding results for other types of imperfection, including modality missing, automation speech recognition error and attacks on text, highlighting the generalizability of our model. Finally, we conduct case studies on general additive distribution, which introduce background noise and blur into raw video clips, further revealing the capability of our proposed method for real-world applications. Hua Xu 0003, Kai Gao 0006 |
IEEE Trans. Multim. | 4 |
| 2024 | Meta Noise Adaption Framework for Multimodal Sentiment Analysis With Feature NoiseabstractImproving the robustness of models against feature noise has emerged as one of the most crucial research topics in the field of multimodal sentiment analysis. Recent studies assume that the training instances are free of noise and develop either translation or reconstruction based method under the guidance of perfect training data for robust testing time performance. However, such an ideal assumption neglects the potential presence of the feature noise in training instances and inevitably results in degradation for the scenario where high-quality training instances are unavailable. In order to achieve robust training with noisy instances, we propose the Meta Noise Adaption (Meta-NA) learning strategy, a meta learning method accumulating the experience of dealing with various types of feature noise. Specifically, we first formulate the tasks distribution where each task is corresponding to one specific pattern of noise, and propose the feature adaption module adding on the unimodal encoder in late fusion based architecture. Through an nested online optimization between the auxiliary feature adaption module and the late fusion backbone modules, the proposed method can leverage shared knowledge across different noisy source tasks and learn how to learn from the noisy instances for robust testing performances. Extensive experiments are conducted on two benchmark multimodal sentiment analysis datasets, namely MOSI and CH-SIMS v2. The results demonstrate that our proposed method can rapidly adapt to various unseen types of feature noise and outperforms all baseline methods, particularly when the training instances are limited. Baozheng Zhang, Hua Xu 0003, Kai Gao 0006 |
IEEE Trans. Multim. | 4 |
| 2024 | Crossmodal Translation Based Meta Weight Adaption for Robust Image-Text Sentiment AnalysisabstractImage-Text Sentiment Analysis task has garnered increased attention in recent years due to the surge in user-generated content on social media platforms. Previous research efforts have made noteworthy progress by leveraging the affective concepts shared between vision and text modalities. However, emotional cues may reside exclusively within one of the prevailing modalities, owing to modality independent nature and the potential absence of certain modalities. In this study, we aim to emphasize the significance of modality-independent emotional behaviors, in addition to the modality-invariant behaviors. To achieve this, we propose a novel approach called Crossmodal Translation-Based Meta Weight Adaption (CTMWA). Specifically, our approach involves the construction of the crossmodal translation network, which serves as the encoder. This architecture captures the shared concepts between vision content and text, empowering the model to effectively handle scenarios where either the vision or textual modality is missing. Building upon the translation-based framework, we introduce the strategy of unimodal weight adaption. Leveraging the meta-learning paradigm, our proposed strategy gradually learns to acquire unimodal weights for individual instances from a few hand-crafted meta instances with unimodal annotations. This enables us to modulate the gradients of each modality encoder based on the discrepancy between modalities during model training. Extensive experiments are conducted on three benchmark image-text sentiment analysis datasets, namely MVSA-Single, MVSA-Multiple, and TumEmo. The empirical results demonstrate that our proposed approach achieves the highest performance across all conventional image-text databases. Furthermore, experiments under modality missing settings and case study for reliable sentiment prediction are also conducted further exhibiting superior robustness as well as reliability of the propose approach. Baozheng Zhang, Hua Xu 0003, Kai Gao 0006 |
IEEE Trans. Multim. | 4 |
| 2023 | Topic model for personalized end-to-end task-oriented dialogue
Zhinan Gou, Yuanzhen Liu, Kai Gao 0006 |
Expert Syst. Appl. | 4 |
| 2022 | An End-to-End Traditional Chinese Medicine Constitution Assessment System Based on Multimodal Clinical Feature Representation and FusionabstractTraditional Chinese Medicine (TCM) constitution is a fundamental concept in TCM theory. It is determined by multimodal TCM clinical features which, in turn, are obtained from TCM clinical information of image (face, tongue, etc.), audio (pulse and voice), and text (inquiry) modality. The auto assessment of TCM constitution is faced with two major challenges: (1) learning discriminative TCM clinical feature representations; (2) jointly processing the features using multimodal fusion techniques. The TCM Constitution Assessment System (TCM-CAS) is proposed to provide an end-to-end solution to this task, along with auxiliary functions to aid TCM researchers. To improve the results of TCM constitution prediction, the system combines multiple machine learning algorithms such as facial landmark detection, image segmentation, graph neural networks and multimodal fusion. Extensive experiments are conducted on a four-category multimodal TCM constitution dataset, and the proposed method achieves state-of-the-art accuracy. Provided with datasets containing annotations of diseases, the system can also perform automatic disease diagnosis from a TCM perspective. Huisheng Mao, Baozheng Zhang, Hua Xu 0003, Kai Gao 0006 |
AAAI | 4 |
| 2022 | An In-depth Interactive and Visualized Platform for Evaluating and Analyzing MRC ModelsabstractMachine Reading Comprehension (MRC) has made leaps and bounds when focusing on answering questions. However, since the existing accuracy-based evaluation metrics are agnostic to the nuances of neural networks, the true understanding and inferencing abilities of MRC models remain largely unknown. To address the above limitations, InDepth-Eva-MRC, an interactive and visualized platform, is proposed to provide analysis from cognitive fine-grained for MRC models. Concretely, the platform makes post-hoc systems to explain the behavior of MRC models. On the one hand, it analyzes the linguistic bias via performances with different linguistic properties. On the other hand, it performs skill-based analysis methods based on the modified test samples and semi-automatically generated test samples. Furthermore, through its detailed and interactive visualizations, the platform offers in-depth results analysis and model comparison from cognitive fine-grained. A screencast video and additional external material are available on https://github.com/thuiar/InDepth-Eva-MRC. Zhijing Wu 0002, Jingliang Fang, Hua Xu 0003, Kai Gao 0006 |
CIKM | 4 |
| 2022 | Make Acoustic and Visual Cues Matter: CH-SIMS v2.0 Dataset and AV-Mixup Consistent ModuleabstractMultimodal sentiment analysis (MSA), which supposes to improve text-based sentiment analysis with associated acoustic and visual modalities, is an emerging research area due to its potential applications in Human-Computer Interaction (HCI). However, existing researches observe that the acoustic and visual modalities contribute much less than the textual modality, termed as text-predominant. Under such circumstances, in this work, we emphasize making non-verbal cues matter for the MSA task. Firstly, from the resource perspective, we present the CH-SIMS v2.0 dataset, an extension and enhancement of the CH-SIMS. Compared with the original dataset, the CH-SIMS v2.0 doubles its size with another 2121 refined video segments containing both unimodal and multimodal annotations and collects 10161 unlabelled raw video segments with rich acoustic and visual emotion-bearing context to highlight non-verbal cues for sentiment prediction. Secondly, from the model perspective, benefiting from the unimodal annotations and the unsupervised data in the CH-SIMS v2.0, the Acoustic Visual Mixup Consistent (AV-MC) framework is proposed. The designed modality mixup module can be regarded as an augmentation, which mixes the acoustic and visual modalities from different videos. Through drawing unobserved multimodal context along with the text, the model can learn to be aware of different non-verbal contexts for sentiment prediction. Our evaluations demonstrate that both CH-SIMS v2.0 and AV-MC framework enable further research for discovering emotion-bearing acoustic and visual cues and pave the path to interpretable end-to-end HCI applications for real-world scenarios. The full dataset and code are available for use at https://github.com/thuiar/ch-sims-v2. Huisheng Mao, Zhiyun Liang, Wanqiuyue Yang, Yuanzhe Qiu, Tie Cheng, Xiaoteng Li, Hua Xu 0003, Kai Gao 0006 |
ICMI | 10 |
| 2022 | GAR-Net: A Graph Attention Reasoning Network for conversation understanding
Hua Xu 0003, Yunfeng Xu, Jiyun Zou, Kai Gao 0006 |
Knowl. Based Syst. | 6 |
| 2021 | Representation iterative fusion based on heterogeneous graph neural network for joint entity and relation extraction
Hua Xu 0003, Xiaoteng Li, Kai Gao 0006 |
Knowl. Based Syst. | 5 |
| 2020 | Multi-Channel Convolutional Neural Networks with Adversarial Training for Few-Shot Relation Classification (Student Abstract)abstractThe distant supervised (DS) method has improved the performance of relation classification (RC) by means of extending the dataset. However, DS also brings the problem of wrong labeling. Contrary to DS, the few-shot method relies on few supervised data to predict the unseen classes. In this paper, we use word embedding and position embedding to construct multi-channel vector representation and use the multi-channel convolutional method to extract features of sentences. Moreover, in order to alleviate few-shot learning to be sensitive to overfitting, we introduce adversarial learning for training a robust model. Experiments on the FewRel dataset show that our model achieves significant and consistent improvements on few-shot RC as compared with baselines. Yuxiang Xie, Hua Xu 0003, Congcong Yang, Kai Gao 0006 |
AAAI | 4 |
| 2020 | Dynamic Prototype Selection by Fusing Attention Mechanism for Few-Shot Relation Classification
Linfang Wu, Huaping Zhang, Yaofei Yang, Xin Liu 0141, Kai Gao 0006 |
ACIIDS (1) | 5 |
| 2020 | CM-BERT: Cross-Modal BERT for Text-Audio Sentiment AnalysisabstractMultimodal sentiment analysis is an emerging research field that aims to enable machines to recognize, interpret, and express emotion. Through the cross-modal interaction, we can get more comprehensive emotional characteristics of the speaker. Bidirectional Encoder Representations from Transformers (BERT) is an efficient pre-trained language representation model. Fine-tuning it has obtained new state-of-the-art results on eleven natural language processing tasks like question answering and natural language inference. However, most previous works fine-tune BERT only base on text data, how to learn a better representation by introducing the multimodal information is still worth exploring. In this paper, we propose the Cross-Modal BERT (CM-BERT), which relies on the interaction of text and audio modality to fine-tune the pre-trained BERT model. As the core unit of the CM-BERT, masked multimodal attention is designed to dynamically adjust the weight of words by combining the information of text and audio modality. We evaluate our method on the public multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results show that it has significantly improved the performance on all the metrics over previous baselines and text-only finetuning of BERT. Besides, we visualize the masked multimodal attention and proves that it can reasonably adjust the weight of words by introducing audio modality information. Kaicheng Yang 0005, Hua Xu 0003, Kai Gao 0006 |
ACM Multimedia | 3 |
| 2020 | Heterogeneous graph neural networks for noisy few-shot relation classification
Yuxiang Xie, Hua Xu 0003, Jiaoe Li, Congcong Yang, Kai Gao 0006 |
Knowl. Based Syst. | 5 |
| 2018 | Two-Stage Attention Network for Aspect-Level Sentiment Classification
Kai Gao 0006, Hua Xu 0003, Chengliang Gao, Xiaomin Sun 0001, Junhui Deng |
ICONIP (4) | 1 |
| 2018 | Attention-Based BiLSTM Network with Lexical Feature for Emotion ClassificationabstractEmotion classification is an important task for identifying users' emotional expressions in text. Though a variety of neural models have been proposed nowadays, these models mainly focus on modeling the content of words or characters without fully employing the emotional features in lexical features, especially the features of part-of -speech (POS). In this paper, we reveal that the information of POS as well as that of words is important for identifying the type of emotion in a given text. We propose two simple models to fully learn the emotional features of the POS of words. Every model consists of the long short-term memory (LSTM) network as input encoders and the component of attention mechanism. One model is to concatenate the POS tags of vectors into the hidden states of representations generated by LSTM as raw feature representations and put them into the component of attention mechanism to generate the text representation toward a special emotion. The other is to use both LSTM and attention mechanism to model the context representation of words and those of POS tags respectively and concatenate these context representations as the text representation toward a special emotion. We conduct some experiments on datasets for evaluation and demonstrate the effectiveness of our model, where the datasets consist of the open-source dataset from NLPCC& 2014 and the dataset of manual annotation. Experimental results show that our models can achieve outstanding performance for emotion classification in Chinese Weibo texts and outperform classical baselines. Kai Gao 0006, Hua Xu 0003, Chengliang Gao, Hanyong Hao, Junhui Deng, Xiaomin Sun 0001 |
IJCNN | 1 |
| 2017 | A Convolutional Neural Network Based Sentiment Classification and the Convolutional Kernel Representation
Shen Gao, Huaping Zhang, Kai Gao 0006 |
NLDB | 3 |
| 2015 | Emotion Cause Detection for Chinese Micro-Blogs Based on ECOCC Model
Kai Gao 0006, Hua Xu 0003, Jiushuo Wang |
PAKDD (2) | 1 |
| 2015 | A rule-based approach to emotion cause detection for Chinese micro-blogs
Kai Gao 0006, Hua Xu 0003, Jiushuo Wang |
Expert Syst. Appl. | 1 |