Huajie Jiang

dblp:176/1538 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-1158-6321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-agent role-playing by LLMs and LMMs: An explainable open-world multi-modal crisis tweet classification method
Tong Bie, Yongli Hu, Linjia Hao, Tengfei Liu 0005, Huajie Jiang, Junbin Gao
Expert Syst. Appl.6
2026 Modality-Agnostic Hybrid Federated Learning via Knowledge Distillation and Reinforcement Learning-Based Aggregation
abstract
Due to multidimensional heterogeneity, Multimodal Federated Learning (MMFL) confronts fundamental challenges including modality incongruence, modality agnosticism, and modality incompleteness. Existing methods face a trilemma: leveraging external data with privacy risks, isolating features to restrict cross-modal interaction, or incurring high overhead from complex graph-based coordination, all culminating in suboptimal performance. In this paper, we propose a Modality-Agnostic Hybrid Federated Learning (MA-HyFL) framework that synergistically integrates unimodal and multimodal federated processes in modality-agnostic scenarios. Specifically, a bidirectional cross-modal knowledge distillation is employed to promote comprehensive collaboration at inter-client and intra-client levels, enabling robust knowledge transfer among heterogeneous modalities. A reinforcement learning-based aggregation mechanism is further introduced to orchestrate federated workflows through reward-driven policy optimization, dynamically integrating contributive client selection and adaptive aggregation weighting for closed-loop decision-making. Extensive experiments show that MA-HyFL significantly outperforms the other baseline methods in four realistic real-world applications, each exhibiting varying degrees of statistical heterogeneity and missing rates.
Yongli Hu, Huajie Jiang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.3
2026 Multimodal Knowledge Graph Completion by Cross-Modal Interaction With Similarity Enhancing and Difference Embracing
abstract
Multimodal knowledge graph completion (MMKGC) enhances the precision and breadth of application of knowledge graphs by integrating rich data from various modalities, steadily increasing its appeal in the research community. Prior studies mainly focus on the common representation of different modalities while neglecting the different and complementary features. On the contrary, some works tend to model triples of each modality separately while overlooking the similarities between modalities. It is challenging to associate the heterogeneous modalities effectively for MMKGC. In this article, we introduce a novel MMKGC framework by cross-modal interaction with similarity-enhancing and difference-embracing (CISEDE), which leverages both the similarities and differences among multimodal entities by a proposed cross-modal interaction mechanism. In the cross-modal interaction, multihead attention is employed to enhance similarity information from multimodal entities and embrace different information by linking various modal triples. Through relation-guided fusion, the modal triples are decoded and merged for MMKGC. The experimental results on three commonly used datasets, FB15k-237, WN9, and WN18RR, show that the proposed method achieves state-of-the-art performance.
Linjia Hao, Yongli Hu, Tong Bie, Huajie Jiang, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.4
2025 Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain semantic-related visual features, which may cause overfitting on seen classes with a limited number of training images. This paper proposes a novel visual and semantic prompt collaboration framework, which utilizes prompt tuning techniques for efficient feature adaptation. Specifically, we design a visual prompt to integrate the visual information for discriminative feature learning and a semantic prompt to integrate the semantic formation for visual-semantic alignment. To achieve effective prompt information integration, we further design a weak prompt fusion mechanism for the shallow layers and a strong prompt fusion mechanism for the deep layers in the network. Through the collaboration of visual and semantic prompts, we can obtain discriminative semantic-related features for generalized zero-shot image recognition. Extensive experiments demonstrate that our framework consistently achieves favorable performance in both conventional zero-shot learning and generalized zero-shot learning benchmarks compared to other state-of-the-art methods.
Huajie Jiang, Zhengxian Li, Xiaohan Yu 0001, Yongli Hu, Jian Yang 0001, Yuankai Qi
CVPR1
2025 MVDT: Multiview Distillation Transformer for View-Invariant Sign Language Translation
abstract
ABSTRACT Sign language translation based on machine learning plays a crucial role in facilitating communication between deaf and hearing individuals. However, due to the complexity and variability of sign language, coupled with limited observation angles, single‐view sign language translation models often underperform in real‐world applications. Although some studies have attempted to improve translation efficiency by incorporating multiview data, challenges, such as feature alignment, fusion, and the high cost of capturing multiview data, remain significant barriers in many practical scenarios. To address these issues, we propose a multiview distillation transformer model (MVDT) for continuous sign language translation. The MVDT introduces a novel distillation mechanism, where a teacher model is designed to learn common features from multiview data, subsequently guiding a student model to extract view‐invariant features using only single‐view input. To evaluate the proposed method, we construct a multiview sign language dataset comprising five distinct views and conduct extensive experiments comparing the MVDT with state‐of‐the‐art methods. Experimental results demonstrate that the proposed model exhibits superior view‐invariant translation capabilities across different views.
Yongli Hu, Huajie Jiang
IET Comput. Vis.3
2025 Latent attribute augmented network for few-shot class-incremental learning
Yongli Hu, Jiasen Zhang, Huajie Jiang
Neurocomputing3
2025 PRNet: A Progressive Refinement Network for referring image segmentation
Jing Liu 0059, Huajie Jiang, Yongli Hu
Neurocomputing2
2025 Multi-view Isolated sign language recognition based on cross-view and multi-level transformer
Yongli Hu, Huajie Jiang
Multim. Syst.3
2025 Dual Prototype Contrastive Network for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) requires that models are able to recognize classes they were trained on, and new classes they haven't seen before. Feature-generation approaches are popular due to their effectiveness in mitigating overfitting to the training classes. Existing generative approaches usually adopt simple discriminators for distribution or classification supervision, however, thus limiting their ability to generate visual features that are discriminative of and transferable to novel categories. To overcome this limitation and improve the quality of generated features, we propose a dual prototype contrastive augmented discriminator for the generative adversarial network. Specifically, we design a Dual Prototype Contrastive Network (DPCN), which leverages complementary information between visual space and semantic space through multi-task prototype contrastive learning. Contrastive learning of the visual prototypes enhances the ability of the generated features to distinguish between classes, while the contrastive learning of the semantic prototypes improves their transferability. Furthermore, we introduce margins into the contrastive learning process to ensure both intra-class compactness and inter-class separation. To demonstrate the effectiveness of the proposed approach, we conduct experiments on three widely-used zero-shot learning benchmark datasets, where DPCN achieves state-of-the-art performance for GZSL.
Huajie Jiang, Zhengxian Li, Yongli Hu, Jian Yang 0001, Anton van den Hengel, Ming-Hsuan Yang 0001, Yuankai Qi
IEEE Trans. Circuits Syst. Video Technol.1
2024 VQA-PDF: Purifying Debiased Features for Robust Visual Question Answering Task
Yandong Bi, Huajie Jiang, Jing Liu 0059, Yongli Hu
ICIC (12)2
2024 Referring Image Segmentation Without Text Annotations
Jing Liu 0059, Huajie Jiang, Yandong Bi, Yongli Hu
ICIC (12)2
2024 Generating Graph-Based Rules for Enhancing Logical Reasoning
Huajie Jiang, Yongli Hu
ICIC (12)2
2024 MBMF: Constructing memory banks of multi-scale features for anomaly detection
abstract
Abstract In industrial manufacturing, how to accurately classify defective products and locate the location of defects has always been a concern. Previous studies mainly measured similarity based on extracting single‐scale features of samples. However, only using the features of a single scale is hard to represent different sizes and types of anomalies. Therefore, the authors propose a set of memory banks of multi‐scale features (MBMF) to enrich feature representation and detect and locate various anomalies. To extract features of different scales, different aggregation functions are designed to produce the feature maps at different granularity. Based on the multi‐scale features of normal samples, the MBMF are constructed. Meanwhile, to better adapt to the feature distribution of the training samples, the authors proposed a new iterative updating method for the memory banks. Testing on the widely used and challenging dataset of MVTec AD, the proposed MBMF achieves competitive image‐level anomaly detection performance (Image‐level Area Under the Receiver Operator Curve (AUROC)) and pixel‐level anomaly segmentation performance (Pixel‐level AUROC). To further evaluate the generalisation of the proposed method, we also implement anomaly detection on the BeanTech AD dataset, a commonly used dataset in the field of anomaly detection, and the Fashion‐MNIST dataset, a widely used dataset in the field of image classification. The experimental results also verify the effectiveness of the proposed method.
Yongli Hu, Huajie Jiang
IET Comput. Vis.4
2024 Self-supervised knowledge distillation in counterfactual learning for VQA
Yandong Bi, Huajie Jiang, Hanfu Zhang, Yongli Hu
Pattern Recognit. Lett.2
2024 See and Learn More: Dense Caption-Aware Representation for Visual Question Answering
abstract
With the rapid development of deep learning models, great improvements have been achieved in the Visual Question Answering (VQA) field. However, modern VQA models are easily affected by language priors, which ignore image information and learn the superficial relationship between questions and answers, even in the optimal pre-training model. The main reason is that visual information is not fully extracted and utilized, which results in a domain gap between vision and language modalities to a certain extent. In order to mitigate the circumstances, we propose to extract dense captions (auxiliary semantic information) from images to enhance the visual information for reasoning and utilize them to release the gap between vision and language since the dense captions and the questions are from the same language modality (i.e., phrase or sentence). In this paper, we propose a novel dense caption-aware visual question answering model called DenseCapBert to enhance visual reasoning. Specifically, we generate dense captions for the images and propose a multimodal interaction mechanism to fuse dense captions, images, and questions in a unified framework, which makes the VQA models more robust. The experimental results on GQA, GQA-OOD, VQA v2, and VQA-CP v2 datasets show that dense captions are beneficial to improving the model generalization and our model effectively mitigates the language bias problem.
Yandong Bi, Huajie Jiang, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Fair Attention Network for Robust Visual Question Answering
abstract
As a prevailing cross-modal reasoning task, Visual Question Answering (VQA) has achieved impressive progress in the last few years, where the language bias is widely studied to learn more robust VQA models. However, the visual bias, which also influences the robustness of VQA models, is seldomly considered, resulting in weak inference ability. Therefore, how to balance the effect of language bias and visual bias has become essential in the current VQA task. In this paper, we devise a new reweighting strategy taking both the language bias and visual bias into account, and propose a Fair Attention Network for Robust Visual Question Answering (named as FAN-VQA). It first constructs a question bias branch and a visual bias branch to estimate the bias information from two modalities, which are utilized to judge the importance of samples. Then, adaptive importance weights are learned from the bias information and assigned to the candidate answers to adjust the training losses, enabling the model to shift more attention to the difficult samples that need less-salient visual clues to infer the correct answer. In order to improve the robustness of the VQA model, we design a progressive strategy to balance the influence of original training loss and adjusted training loss. Extensive experiments on the VQA-CP v2, VQA v2, and VQA-CE datasets demonstrate the effectiveness of the proposed FAN-VQA method.
Yandong Bi, Huajie Jiang, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Domain-Aware Prototype Network for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning(GZSL) aims to recognize images from seen and unseen classes with side information, such as manually annotated attribute vectors. Traditional methods focus on mapping images and semantics into a common latent space, thus achieving the visual-semantics alignment. Since the unseen classes are unavailable during training, there is a serious problem of recognition bias, which will tend to recognize unseen classes as seen classes. To solve this problem, we propose a Domain-aware Prototype Network(DPN), which splits the GZSL problem into the seen class recognition and unseen class recognition problem. For the seen classes, we design a domain-aware prototype learning branch with a dual attention feature encoder to capture the essential visual information, which aims to recognize the seen classes and discriminate the novel categories. To further recognize the fine-grained unseen classes, a visual-semantic embedding branch is designed, which aims to align the visual and semantic information for unseen-class recognition. Through the multi-task learning of the prototype learning branch and visual-semantic embedding branch, our model can achieve excellent performance on three popular GZSL datasets.
Yongli Hu, Lincong Feng, Huajie Jiang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Incorporating Multi-Level Sampling with Adaptive Aggregation for Inductive Knowledge Graph Completion
abstract
In recent years, Graph Neural Networks (GNNs) have achieved unprecedented success in handling graph-structured data, thereby driving the development of numerous GNN-oriented techniques for inductive knowledge graph completion (KGC). A key limitation of existing methods, however, is their dependence on pre-defined aggregation functions, which lack the adaptability to diverse data, resulting in suboptimal performance on established benchmarks. Another challenge arises from the exponential increase in irrelated entities as the reasoning path lengthens, introducing unwarranted noise and consequently diminishing the model’s generalization capabilities. To surmount these obstacles, we design an innovative framework that synergizes M ulti- L evel S ampling with an A daptive A ggregation mechanism (MLSAA). Distinctively, our model couples GNNs with enhanced set transformers, enabling dynamic selection of the most appropriate aggregation function tailored to specific datasets and tasks. This adaptability significantly boosts both the model’s flexibility and its expressive capacity. Additionally, we unveil a unique sampling strategy designed to selectively filter irrelevant entities, while retaining potentially beneficial targets throughout the reasoning process. We undertake an exhaustive evaluation of our novel inductive KGC method across three pivotal benchmark datasets and the experimental results corroborate the efficacy of MLSAA.
Huajie Jiang, Yongli Hu
ACM Trans. Knowl. Discov. Data2
2023 Breaking the Barrier Between Pre-training and Fine-tuning: A Hybrid Prompting Model for Knowledge-Based VQA
abstract
Considerable performance gains have been achieved for knowledge-based visual question answering due to the visual-language pre-training models with pre-training-then-fine-tuning paradigm. However, because the targets of the pre-training and fine-tuning stages are different, there is an evident barrier that prevents the cross-modal comprehension ability developed in the pre-training stage from fully endowing the fine-tuning task. To break this barrier, in this paper, we propose a novel hybrid prompting model for knowledge-based VQA, which inherits and incorporates the pre-training and fine-tuning tasks with a shared objective. Specifically, based on static declaration prompt, we construct a consistent goal with the fine-tuning via masked language modeling to inherit capabilities of pre-training task, while selecting the top-t relevant knowledge in a dense retrieval manner. Additionally, a dynamic knowledge prompt is learned from retrieved knowledge, which not only alleviates the length constraint on inputs for visual-language pre-trained models but also assists in providing answer features via fine-tuning. Combining and unifying the aims of the two stages could fully exploit the abilities of pre-training and fine-tuning to predict answer. We evaluate the proposed model on the OKVQA dataset, and the result shows that our model outperforms the state-of-the-art methods based on visual-language pre-training models with a noticeable performance gap and even exceeds the large-scale language model of GPT-3, which proves the benefits of the hybrid prompts and the advantages of unifying pre-training to fine-tuning.
Zhongfan Sun, Yongli Hu, Qingqing Gao, Huajie Jiang, Junbin Gao
ACM Multimedia4
2023 Multi-level attention for referring expression comprehension
Yunru Zhang, Huajie Jiang, Yongli Hu
Pattern Recognit. Lett.3
2023 Substructure-aware subgraph reasoning for inductive relation prediction
Huajie Jiang, Yongli Hu
J. Supercomput.2
2022 Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency
abstract
In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the style effect of generated images, video cartoonization has additional requirements on the temporal consistency. In this paper, we propose a spatially-adaptive semantic alignment framework with perceptual motion consistency for coherent video cartoonization in an unsupervised manner. The semantic alignment module is designed to restore deformation of semantic structure caused by spatial information lost in the encoder-decoder architecture. Furthermore, we introduce the spatio-temporal correlative map as a style-independent, global-aware regularization on perceptual motion consistency. Deriving from similarity measurement of high-level features in photo and cartoon frames, it captures global semantic information beyond raw pixel-value of optical flow. Besides, the similarity measurement disentangles temporal relationship from domain-specific style properties, which helps regularize the temporal consistency without hurting style effects of cartoon images. Qualitative and quantitative experiments demonstrate our method is able to generate highly stylistic and temporal consistent cartoon videos.
Zhenhuan Liu, Liang Li 0003, Huajie Jiang, Xin Jin 0004, Dandan Tu, Shuhui Wang, Zhengjun Zha
AAAI3
2022 CGNN: Caption-assisted graph neural network for image-text retrieval
Yongli Hu, Hanfu Zhang, Huajie Jiang, Yandong Bi
Pattern Recognit. Lett.3
2021 Fine-Grained Image-Text Retrieval via Complementary Feature Learning
Yantao Jia, Huajie Jiang
MMM (1)3
2019 Transferable Contrastive Network for Generalized Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is a challenging problem that aims to recognize the target categories without seen data, where semantic information is leveraged to transfer knowledge from some source classes. Although ZSL has made great progress in recent years, most existing approaches are easy to overfit the sources classes in generalized zero-shot learning (GZSL) task, which indicates that they learn little knowledge about target classes. To tackle such problem, we propose a novel Transferable Contrastive Network (TCN) that explicitly transfers knowledge from the source classes to the target classes. It automatically contrasts one image with different classes to judge whether they are consistent or not. By exploiting the class similarities to make knowledge transfer from source images to similar target classes, our approach is more robust to recognize the target images. Experiments on five benchmark datasets show the superiority of our approach for GZSL.
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
ICCV1
2019 Adaptive Metric Learning For Zero-Shot Recognition
abstract
Zero-shot learning (ZSL) has enjoyed great popularity in recent years due to its ability to recognize novel objects, where semantic information is exploited to build up relations among different categories. Traditional ZSL approaches usually focus on learning more robust visual-semantic embeddings among seen classes and directly apply them to the unseen classes without considering whether they are suitable. It is well known that domain gap exists between seen and unseen classes. In order to tackle such problem, we propose a novel adaptive metric learning approach to measure the compatibility between visual samples and class semantics, where class similarities are utilized to adapt the visual-semantic embedding to the unseen classes. Extensive experiments on four benchmark ZSL datasets show the effectiveness of the proposed approach.
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Signal Process. Lett.1
2018 Learning Class Prototypes via Structure Alignment for Zero-Shot Recognition
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001
ECCV (10)1
2018 Attribute annotation on large-scale image database by active knowledge transfer
Huajie Jiang, Ruiping Wang 0001, Yan Li 0014, Haomiao Liu, Shiguang Shan, Xilin Chen 0001
Image Vis. Comput.1
2017 Learning Discriminative Latent Attributes for Zero-Shot Classification
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Yi Yang 0001, Xilin Chen 0001
ICCV1
2015 Two Birds, One Stone: Jointly Learning Binary Code for Large-Scale Face Image Retrieval and Attributes Prediction
abstract
We address the challenging large-scale content-based face image retrieval problem, intended as searching images based on the presence of specific subject, given one face image of him/her. To this end, one natural demand is a supervised binary code learning method. While the learned codes might be discriminating, people often have a further expectation that whether some semantic message (e.g., visual attributes) can be read from the human-incomprehensible codes. For this purpose, we propose a novel binary code learning framework by jointly encoding identity discriminability and a number of facial attributes into unified binary code. In this way, the learned binary codes can be applied to not only fine-grained face image retrieval, but also facial attributes prediction, which is the very innovation of this work, just like killing two birds with one stone. To evaluate the effectiveness of the proposed method, extensive experiments are conducted on a new purified large-scale web celebrity database, named CFW 60K, with abundant manual identity and attributes annotation, and experimental results exhibit the superiority of our method over state-of-the-art.
Yan Li 0014, Ruiping Wang 0001, Haomiao Liu, Huajie Jiang, Shiguang Shan, Xilin Chen 0001
ICCV4