Qintai Hu

dblp:168/2738 · DBLP profile ↗
← Back
8ranked-venue papers in the field
1as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Information Retrieval & Web Search · 3
YearPublicationVenuePosition
2025 HGAtt-ARN: A Novel Adversarial Reconstruction Network Based on Higher-order Gate Attention for Incomplete Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) is a technique for understanding and recognizing human sentiment by learning representations of different modalities. However, most existing MSA models fail to accurately analyze the real sentiment expressed by irony in slang under incomplete modality. Meanwhile, existing modality reconstruction methods suffer from modality confusion problem, which trigger catastrophic errors in sentiment analysis. To address these challenges, we propose a novel Adversarial Reconstruction Network based on Higher-order Gate Attention for incomplete Multimodal Sentiment Analysis (HGAtt-ARN). Specifically, we propose Higher-order Gate Attention (HGAtt), which can model real representations through high-dimensional mapping and gating computations of heterogeneous modalities, and accordingly construct Adversarial Reconstruction Network (ARN), which contains two core components: (a) HGAtt Reconstructor and CA-HGAtt Reconstructor, which reconstruct modality representations of the real sentiment states implied in irony under incomplete modality through HGAtt and Cross Attention; (b) modality Discriminator, which solves the modality confusion problem by maintaining reconstructed modality invariance through adversarial optimization with Reconstructor. Furthermore, we introduce multimodal Contrastive Learning to improve the performance of our model. Experiments on three public datasets show that HGAtt-ARN achieves SOTA performance, and the case study shows it can accurately recognize the real sentiment expressed by the American ironic slang under incomplete modality.
Qingpeng Wen, Qintai Hu, Bi Zeng
ICMR4
2025 ReFNet: Rehearsal-based graph lifelong learning with multi-resolution framelet graph neural networks
Ming Li 0065, Yongchun Gu, Qintai Hu
Inf. Sci.6
2024 CGI-MRE: A Comprehensive Genetic-Inspired Model For Multimodal Relation Extraction
abstract
Multimodal Relation Extraction (MRE) is an entity relationship extraction method based on multimodal information. Most existing MRE methods have two issues: 1) Weak cross-modal correlation and poor semantic consistency. 2) They do not achieve text-guided fusion of different modalities, resulting in excessive introduction of image noise. To address these issues, we propose an innovative MRE method inspired by genetics-A Comprehensive Genetic-Inspired For Multimodal Relation Extraction (CGI-MRE). It consists of two main modules: Gene Extraction And Recombination Module (GERM) and Text-Guided Fusion Module (TGFM). In the GERM module, we regard the text features and visual features as a feature body respectively, and decompose each feature body into common sub-features and unique sub-features. For these sub-features, we designed a Common Gene Extraction Mechanism (CGEM) to extract common advantageous genes in different modalities, a Unique Gene Extraction Mechanism (UGEM) to extract unique advantageous genes in each modality, and we finally use a Gene Recombination Mechanism (GRM) to obtain recombinant features that highly correlated with different modalities and have strong semantic consistency. In TGFM module, we organically fuse and extract the features in the recombined features that are beneficial to MRE. We use gate to adjust the text-guided original attention score and pooling attention score to obtain the text-guided saliency attention score. We can use this score to strictly extract information that is text-guided and beneficial to MRE from the image recombinant feature. Experimental results on the MNRE dataset show that our model outperforms the state-of-the-art performance and achieves F1-score of 84.62%.
Zhaokang Huang, Hongjun Ouyang, Qintai Hu, Bi Zeng
ICMR4
2024 VEC-MNER: Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction for Multimodal NER
abstract
Multimodal Named Entity Recognition (MNER) aims to leverage visual information to identify entity boundaries and categories in social media posts. Existing methods mainly adopt heterogeneous architecture, with ResNet (CNN-based) and BERT (Transformer-based) dedicated to modeling visual and textual features, respectively. However, current approaches still face the following issues: (1) Weak cross-modal correlations and poor semantic consistency. (2) Suboptimal fusion results when visual objects and textual entities are inconsistent. To this end, we propose a Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction (VEC-MNER) model for MNER. Specifically, compared to heterogeneous architectures, we propose a new homogeneous Hybrid Transformer Architecture, which naturally reduces the heterogeneity. Moreover, we design the Correlation-Aware Alignment (CAA-Encoder) layer and the Correlation-Aware Deep Fusion (CADF-Encoder) layer, combined with contrastive learning, to achieve more effective implicit alignment and deep semantic fusion between modalities, respectively. We also construct a Correlation-Aware (CA) module that can effectively reduce heterogeneity between modalities and alleviate visual deviation. Experimental results demonstrate that our approach achieves SOTA performance, achieving 74.89% and 87.51% F1-score on Twitter-2015 and Twitter-2017, respectively.
Hongjun Ouyang, Qintai Hu, Bi Zeng, Qingpeng Wen
ICMR3
2024 SYEnet: Simple yet effective network for palmprint recognition
Qintai Hu, Shuping Zhao
Inf. Sci.2
2019 Learning peer recommendation using attention-driven CNN with interaction tripartite graph
Qintai Hu, Zhongmei Han, Xiao-Fan Lin 0001, Qionghao Huang
Inf. Sci.1
2018 A novel approach for entity resolution in scientific documents using context graphs
Changqin Huang, Jia Zhu 0003, Xiaodi Huang 0001, Min Yang 0007, Gabriel Pui Cheong Fung, Qintai Hu
Inf. Sci.6
2018 Personalized learning full-path recommendation model based on LSTM neural networks
Yuwen Zhou, Changqin Huang, Qintai Hu, Jia Zhu 0003, Yong Tang 0001
Inf. Sci.3