Qintai Hu

dblp:168/2738 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NEXPRO: Multimodal negative expression prompting for open-set video-based facial expression recognition
Qintai Hu, Yifei Su, Fan Jiang 0017, Xiaodi Huang 0001, Qionghao Huang, Changqin Huang
Expert Syst. Appl.1
2026 Adaptive Pseudo-labeling for Causal-driven Cross-Modal hashing
Qintai Hu, Min Meng 0001, Zetao Ma, Jigang Wu
Expert Syst. Appl.2
2026 Dynamic comprehensive difficulty knowledge cells based on KAN network and stable learning for knowledge tracing
Qintai Hu, Qingpeng Wen, Bi Zeng, Liangda Fang
Neural Networks1
2026 GaitPR: A high-confidence multi-granular residual attention network for gait recognition
Xiaona Zheng, Qintai Hu, Shuping Zhao, Jigang Wu
Pattern Recognit. Lett.2
2025 DAGait: Enhancing Gait Recognition with Dynamic Adversarial Training
abstract
Gait recognition systems are increasingly deployed in real-world scenarios where adversarial robustness is critical but remains underexplored. To this end, we propose DAGait, a novel adversarial defense framework for gait recognition, which operates in the gait embedding space to address the incompatibility of conventional image-space defenses with metric-based gait models. Specifically, a feature-level progressive adversarial training strategy (FPA) is proposed, which applies bounded perturbations and gradually integrates adversarial examples into the training phase, enabling a smooth transition from clean to robust representations. We further propose TriWeight, a novel triplet loss that emphasizes adversarially hard examples while employing three dynamic weights to mitigate view variations and sample interference. Extensive experiments conducted on the CASIA-B* dataset demonstrate that DAGait significantly enhances adversarial robustness, offering a practical and effective solution for secure gait-based biometric systems.
Xiaona Zheng, Qintai Hu, Shuping Zhao, Jigang Wu
IJCB2
2025 HGAtt-ARN: A Novel Adversarial Reconstruction Network Based on Higher-order Gate Attention for Incomplete Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) is a technique for understanding and recognizing human sentiment by learning representations of different modalities. However, most existing MSA models fail to accurately analyze the real sentiment expressed by irony in slang under incomplete modality. Meanwhile, existing modality reconstruction methods suffer from modality confusion problem, which trigger catastrophic errors in sentiment analysis. To address these challenges, we propose a novel Adversarial Reconstruction Network based on Higher-order Gate Attention for incomplete Multimodal Sentiment Analysis (HGAtt-ARN). Specifically, we propose Higher-order Gate Attention (HGAtt), which can model real representations through high-dimensional mapping and gating computations of heterogeneous modalities, and accordingly construct Adversarial Reconstruction Network (ARN), which contains two core components: (a) HGAtt Reconstructor and CA-HGAtt Reconstructor, which reconstruct modality representations of the real sentiment states implied in irony under incomplete modality through HGAtt and Cross Attention; (b) modality Discriminator, which solves the modality confusion problem by maintaining reconstructed modality invariance through adversarial optimization with Reconstructor. Furthermore, we introduce multimodal Contrastive Learning to improve the performance of our model. Experiments on three public datasets show that HGAtt-ARN achieves SOTA performance, and the case study shows it can accurately recognize the real sentiment expressed by the American ironic slang under incomplete modality.
Qingpeng Wen, Qintai Hu, Bi Zeng
ICMR4
2025 ReFNet: Rehearsal-based graph lifelong learning with multi-resolution framelet graph neural networks
Ming Li 0065, Yongchun Gu, Qintai Hu
Inf. Sci.6
2024 Adaptive Neighbor Guided View Reconstruction for Incomplete Multiview Clustering
abstract
Graph-based multi-view clustering methods have gained significant attention due to their outstanding ability of clustering-structure representation. Considering the influence on the quality of pre-constructed graphs by noise, in this paper, we propose a novel method called ANGVR. Unlike existing methods that aim to directly learn a consensus graph from multi-view data, ANGVR rebuilds the graph constructed from raw data to seek a consensus graph across views for clustering. Furthermore, to guide the construction of graph, an embedding constraint based on neighboring group structures is introduced, which explores the neighborhood structure information corresponding to neighborhood sets. The experimental results show improvements in accuracy of 5.56%, 6.77% and 4.02%, respectively on 3Sources dataset with 10%, 30% and 50% missing view compared to the existing works.
Zhuojie Huang, Shuping Zhao, Qintai Hu
ISPA4
2024 CGI-MRE: A Comprehensive Genetic-Inspired Model For Multimodal Relation Extraction
abstract
Multimodal Relation Extraction (MRE) is an entity relationship extraction method based on multimodal information. Most existing MRE methods have two issues: 1) Weak cross-modal correlation and poor semantic consistency. 2) They do not achieve text-guided fusion of different modalities, resulting in excessive introduction of image noise. To address these issues, we propose an innovative MRE method inspired by genetics-A Comprehensive Genetic-Inspired For Multimodal Relation Extraction (CGI-MRE). It consists of two main modules: Gene Extraction And Recombination Module (GERM) and Text-Guided Fusion Module (TGFM). In the GERM module, we regard the text features and visual features as a feature body respectively, and decompose each feature body into common sub-features and unique sub-features. For these sub-features, we designed a Common Gene Extraction Mechanism (CGEM) to extract common advantageous genes in different modalities, a Unique Gene Extraction Mechanism (UGEM) to extract unique advantageous genes in each modality, and we finally use a Gene Recombination Mechanism (GRM) to obtain recombinant features that highly correlated with different modalities and have strong semantic consistency. In TGFM module, we organically fuse and extract the features in the recombined features that are beneficial to MRE. We use gate to adjust the text-guided original attention score and pooling attention score to obtain the text-guided saliency attention score. We can use this score to strictly extract information that is text-guided and beneficial to MRE from the image recombinant feature. Experimental results on the MNRE dataset show that our model outperforms the state-of-the-art performance and achieves F1-score of 84.62%.
Zhaokang Huang, Hongjun Ouyang, Qintai Hu, Bi Zeng
ICMR4
2024 VEC-MNER: Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction for Multimodal NER
abstract
Multimodal Named Entity Recognition (MNER) aims to leverage visual information to identify entity boundaries and categories in social media posts. Existing methods mainly adopt heterogeneous architecture, with ResNet (CNN-based) and BERT (Transformer-based) dedicated to modeling visual and textual features, respectively. However, current approaches still face the following issues: (1) Weak cross-modal correlations and poor semantic consistency. (2) Suboptimal fusion results when visual objects and textual entities are inconsistent. To this end, we propose a Hybrid Transformer with Visual-Enhanced Cross-Modal Multi-level Interaction (VEC-MNER) model for MNER. Specifically, compared to heterogeneous architectures, we propose a new homogeneous Hybrid Transformer Architecture, which naturally reduces the heterogeneity. Moreover, we design the Correlation-Aware Alignment (CAA-Encoder) layer and the Correlation-Aware Deep Fusion (CADF-Encoder) layer, combined with contrastive learning, to achieve more effective implicit alignment and deep semantic fusion between modalities, respectively. We also construct a Correlation-Aware (CA) module that can effectively reduce heterogeneity between modalities and alleviate visual deviation. Experimental results demonstrate that our approach achieves SOTA performance, achieving 74.89% and 87.51% F1-score on Twitter-2015 and Twitter-2017, respectively.
Hongjun Ouyang, Qintai Hu, Bi Zeng, Qingpeng Wen
ICMR3
2024 An Enhanced Dual-Channel-Omni-Scale 1DCNN for Fault Diagnosis
Xiaona Zheng, Qintai Hu, Shuping Zhao
PRCV (1)2
2024 Mutual dimensionless improved bearing fault diagnosis based on Bp-increment broad learning system in computer vision
Qintai Hu, Shuping Zhao, Jigang Wu, Jianbin Xiong
Eng. Appl. Artif. Intell.2
2024 SYEnet: Simple yet effective network for palmprint recognition
Qintai Hu, Shuping Zhao
Inf. Sci.2
2024 Clustering Environment Aware Learning for Active Domain Adaptation
abstract
Despite the significant progress in unsupervised domain adaptation (UDA), the performance of UDA methods is still far inferior to that of the fully supervised ones. In practical scenarios, it is usually feasible to acquire labels on a small portion of the target data through active learning (AL), which aims to train an effective model with as few queried instances as possible. However, due to the domain shift, the instances selected by existing AL algorithms can be uninformative, redundant, or outlying. To address this issue, we propose a novel approach, namely, clustering environment-aware learning (CEAL), for active domain adaptation (ADA). CEAL selects potentially the most valuable instances under domain shift by exploring the informativeness and representativeness of target samples in a clustering environment-aware manner. Specifically, for the informativeness, we not only leverage the knowledge of individual points but also their nearby neighbors, by measuring the proposed clustering environment aware informativeness score (CEAIS), thus ensuring that the selected samples are highly informative. For the representativeness, we design two schemes called point distance release (PDR) and informativeness score difference exclusion (ISDE) to guarantee the diversity and validity of the selected samples. Furthermore, we fully utilize the large amount of unlabeled data from target domain via pseudo labeling and adopt information maximization to improve the reliability of the target pseudo labels, thereby further improving the performance of the model. The effectiveness of our method is empirically verified on various benchmark datasets against recent state-of-the-art algorithms.
Jian Zhu 0001, Qintai Hu, Yutang Xiao, Boyu Wang 0004, Bin Sheng 0001, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.3
2023 MKB: Multi-Kernel Bures Metric for Nighttime Aerial Tracking
Peipei Kang, Qintai Hu, Xiaozhao Fang
PRCV (10)3
2021 Enhanced Discriminant Local Direction Pattern Learning for Robust Palmprint Identification
Qintai Hu, Shuping Zhao, Wenyan Wu 0007
PDCAT2
2019 Learning peer recommendation using attention-driven CNN with interaction tripartite graph
Qintai Hu, Zhongmei Han, Xiao-Fan Lin 0001, Qionghao Huang
Inf. Sci.1
2018 A novel approach for entity resolution in scientific documents using context graphs
Changqin Huang, Jia Zhu 0003, Xiaodi Huang 0001, Min Yang 0007, Gabriel Pui Cheong Fung, Qintai Hu
Inf. Sci.6
2018 Personalized learning full-path recommendation model based on LSTM neural networks
Yuwen Zhou, Changqin Huang, Qintai Hu, Jia Zhu 0003, Yong Tang 0001
Inf. Sci.3
2015 Trends in E-Learning Research from 2002-2013: A Co-citation Analysis
abstract
As research in e-learning areas may have progressively become an important field, it is essential to view the emerging trends and critical turns of the publications of research findings. This study utilized Cite Space, a citation analysis software for analyzing and visualizing trends in e-learning research between 2002 and 2013. A total of 4898 research papers were analyzed in terms of the timelines of co-citation, a visualized network of co-cited references and burst terms. It is found that a major ongoing research trend is concerned with the topics of self-efficacy, self-explanation and virtual reality. Moreover, recent research fronts focus on the distributed learning environments, learning communities, computer-mediated communication, and collaborative learning, etc. The research trends and research fronts are distinct and have different implications on future research, so as to track the development of new and critical emerging trends.
Xiao-Fan Lin 0001, Qintai Hu
ICALT2
2015 Exploring the effects of motivational videos for hearing-impaired children
Xiao-Fan Lin 0001, Cailing Deng, Jiayu He, Yinneng Zhang, Qintai Hu
ICCE5