Qiang Jiao

dblp:166/3845 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-4725-5284ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 QANet: Query-Aware Multi-Modal Prior Refinement for Few-Shot Segmentation
abstract
Recent Few-Shot Segmentation (FSS) approaches incorporate vision-language models to improve segmentation by constructing multi-modal priors, including visual, CAM and textual priors. However, these priors are often suboptimal, suffering from background interference, incomplete activation and semantic misalignment. To address these limitations, we propose a query-aware multi-modal prior refinement framework (QANet) for FSS, in which the multi-modal priors are jointly refined for more accurate segmentation by fully exploiting target-relevant cues from query features. Specifically, QANet comprises two modules, i.e., a Query-aware Multi-modal Prior Refinement (QMPR) module for visual and CAM prior refinement, and a Query-Aware Textual Embedding and Prior Refinement (QTEPR) module for textual prior refinement. More specifically, in QMPR, a query self-attention bilateral-guided enhancement strategy and a query self-similarity prior refinement strategy are carefully designed for suppressing background interferences and recovering foreground completeness, respectively. And in QTEPR, such refined visual and CAM priors are first leveraged to extract target-relevant cues for visual-textual domain gap mitigation and text embedding enhancement, producing query-aware textual priors. These textual priors are then enhanced under the guidance of those refined multi-modal priors. Extensive experiments on PASCAL-5iand COCO-20idatasets demonstrate that QANet achieves new state-of-the-art performance. Moreover, QANet remains competitive on cross-domain FSS and weak-label FSS tasks.
Qiang Jiao, Mengrui Shi, Qiang Zhang 0020
IEEE Trans. Circuits Syst. Video Technol.1
2025 Resolving semantic conflicts in RGB-T semantic segmentation
Shenlu Zhao, Ziniu Jin, Qiang Jiao, Qiang Zhang 0020, Jungong Han
Pattern Recognit.3
2024 Deep unsupervised shadow detection with curriculum learning and self-training
Qiang Zhang 0020, Hongyuan Guo, Guanghe Li, Tianlu Zhang, Qiang Jiao
Comput. Vis. Image Underst.5
2024 AMNet: Learning to Align Multi-Modality for RGB-T Tracking
abstract
RGB-T tracking has attracted increasing attention recently due to the all-weather and all-day working capability. However, most current RGB-T trackers usually assume that RGB data and thermal infrared (TIR) data are well spatially aligned, which is difficult to be achieved in practice. Such spatial misalignment between RGB data and TIR data may lead to the ineffective cross-modal information propagation during multi-modal feature fusion, thus reducing the tracking performance. In addition, due to the discrepancy in imaging characteristics of RGB images and TIR images, there also exist great differences between the information captured by the two modality data. The differences in characteristics of RGB and TIR modalities in different local areas will cause a single fusion strategy to be unable to fully explore the complementary information within multi-modal data. For that, we propose an RGB-T tracker, referred to as AMNet, to specifically solve such two problems with two dedicated modules, i.e., a Mutual-interacted Spatial Alignment (MSA) module and an Information Matching Fusion (IMF) module. The former spatially aligns the two modality data through three essential parts, including interactions of multi-modal features, prediction of cross-modal offset map, and enhancement of the aligned features. While the latter first discriminates different types of local regions by employing several intra-modal attention modules and then uses a divide-and-conquer fusion strategy to exploit such discriminative information within RGB and TIR features of different cases for tracking. We validate the effectiveness of our AMNet with extensive experiments on three RGB-T benchmarks, which achieves new state-of-the-art performance.
Tianlu Zhang, Xiaoyi He, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.3
2024 Feature Calibrating and Fusing Network for RGB-D Salient Object Detection
abstract
Due to their imaging mechanisms and techniques, some depth images inevitably have low visual qualities or have some inconsistent foregrounds with their corresponding RGB images. Directly using such depth images will deteriorate the performance of RGB-D SOD. In view of this, a novel RGB-D salient object detection model is presented, which follows the principle of calibration-then-fusion to effectively suppress the influence of such two types of depth images on final saliency prediction. Specifically, the proposed model is composed of two stages, i.e., an image generation stage and a saliency reasoning stage. The former generates high-quality and foreground-consistent pseudo depth images via an image generation network. While the latter first calibrates the original depth information with the aid of those newly generated pseudo depth images and then performs cross-modal feature fusion for the final saliency reasoning. Especially, in the first stage, a Two-steps Sample Selection (TSS) strategy is employed to select such reliable depth images from the original RGB-D image pairs as supervision information to optimize the image generation network. Afterwards, in the second stage, a Feature Calibrating and Fusing Network (FCFNet) is proposed to achieve the calibration-then-fusion of cross-modal information for the final saliency prediction, which is achieved by a Depth Feature Calibration (DFC) module, a Shallow-level Feature Injection (SFI) module and a Multi-modal Multi-scale Fusion (MMF) module. Moreover, a loss function, i.e., Region Consistency Aware (RCA) loss, is presented as an auxiliary loss for FCFNet to facilitate the completeness of salient objects together with the reduction of background interference by considering the local regional consistency in the saliency maps. Experiments on six benchmark datasets demonstrate the superiorities of our proposed RGB-D SOD model over some state-of-the-arts.
Qiang Zhang 0020, Yang Yang 0132, Qiang Jiao, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.4
2024 Exploring Multi-Modal Spatial-Temporal Contexts for High-Performance RGB-T Tracking
abstract
In RGB-T tracking, there exist rich spatial relationships between the target and backgrounds within multi-modal data as well as sound consistencies of spatial relationships among successive frames, which are crucial for boosting the tracking performance. However, most existing RGB-T trackers overlook such multi-modal spatial relationships and temporal consistencies within RGB-T videos, hindering them from robust tracking and practical applications in complex scenarios. In this paper, we propose a novel Multi-modal Spatial-Temporal Context (MMSTC) network for RGB-T tracking, which employs a Transformer architecture for the construction of reliable multi-modal spatial context information and the effective propagation of temporal context information. Specifically, a Multi-modal Transformer Encoder (MMTE) is designed to achieve the encoding of reliable multi-modal spatial contexts as well as the fusion of multi-modal features. Furthermore, a Quality-aware Transformer Decoder (QATD) is proposed to effectively propagate the tracking cues from historical frames to the current frame, which facilitates the object searching process. Moreover, the proposed MMSTC network can be easily extended to various tracking frameworks. New state-of-the-art results on five prevalent RGB-T tracking benchmarks demonstrate the superiorities of our proposed trackers over existing ones.
Tianlu Zhang, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Image Process.2
2024 Mitigating Modality Discrepancies for RGB-T Semantic Segmentation
abstract
Semantic segmentation models gain robustness against adverse illumination conditions by taking advantage of complementary information from visible and thermal infrared (RGB-T) images. Despite its importance, most existing RGB-T semantic segmentation models directly adopt primitive fusion strategies, such as elementwise summation, to integrate multimodal features. Such strategies, unfortunately, overlook the modality discrepancies caused by inconsistent unimodal features obtained by two independent feature extractors, thus hindering the exploitation of cross-modal complementary information within the multimodal data. For that, we propose a novel network for RGB-T semantic segmentation, i.e. MDRNet+, which is an improved version of our previous work ABMDRNet. The core of MDRNet+ is a brand new idea, termed the strategy of bridging-then-fusing, which mitigates modality discrepancies before cross-modal feature fusion. Concretely, an improved Modality Discrepancy Reduction (MDR+) subnetwork is designed, which first extracts unimodal features and reduces their modality discrepancies. Afterward, discriminative multimodal features for RGB-T semantic segmentation are adaptively selected and integrated via several channel-weighted fusion (CWF) modules. Furthermore, a multiscale spatial context (MSC) module and a multiscale channel context (MCC) module are presented to effectively capture the contextual information. Finally, we elaborately assemble a challenging RGB-T semantic segmentation dataset, i.e., RTSS, for urban scene understanding to mitigate the lack of well-annotated training data. Comprehensive experiments demonstrate that our proposed model surpasses other state-of-the-art models on the MFNet, PST900, and RTSS datasets remarkably.
Shenlu Zhao, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.3
2023 Efficient RGB-T Tracking via Cross-Modality Distillation
abstract
Most current RGB-T trackers adopt a two-stream structure to extract unimodal RGB and thermal features and complex fusion strategies to achieve multi-modal feature fusion, which require a huge number of parameters, thus hindering their real-life applications. On the other hand, a compact RGB-T tracker may be computationally efficient but encounter non-negligible performance degradation, due to the weakening of feature representation ability. To remedy this situation, a cross-modality distillation framework is presented to bridge the performance gap between a compact tracker and a powerful tracker. Specifically, a specific-common feature distillation module is proposed to transform the modality-common information as well as the modality-specific information from a deeper two-stream network to a shallower single-stream network. In addition, a multi-path selection distillation module is proposed to instruct a simple fusion module to learn more accurate multi-modal information from a well-designed fusion mechanism by using multiple paths. We validate the effectiveness of our method with extensive experiments on three RGB-T benchmarks, which achieves state-of-the-art performance but consumes much less computational resources.
Tianlu Zhang, Hongyuan Guo, Qiang Jiao, Qiang Zhang 0020, Jungong Han
CVPR3
2022 Middle-Level Feature Fusion for Lightweight RGB-D Salient Object Detection
abstract
Most existing RGB-D salient object detection (SOD) models adopt a two-stream structure to extract the information from the input RGB and depth images. Since they use two subnetworks for unimodal feature extraction and multiple multi-modal feature fusion modules for extracting cross-modal complementary information, these models require a huge number of parameters, thus hindering their real-life applications. To remedy this situation, we propose a novel middle-level feature fusion structure that allows to design a lightweight RGB-D SOD model. Specifically, the proposed structure first employs two shallow subnetworks to extract low- and middle-level unimodal RGB and depth features, respectively. Afterward, instead of integrating middle-level unimodal features multiple times at different layers, we just fuse them once via a specially designed fusion module. On top of that, high-level multi-modal semantic features are further extracted for final salient object detection via an additional subnetwork. This will greatly reduce the network's parameters. Moreover, to compensate for the performance loss due to parameter deduction, a relation-aware multi-modal feature fusion module is specially designed to effectively capture the cross-modal complementary information during the fusion of middle-level multi-modal features. By enabling the feature-level and decision-level information to interact, we maximize the usage of the fused cross-modal middle-level features and the extracted cross-modal high-level features for saliency prediction. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed method over some state-of-the-art methods. Remarkably, our proposed model has only 3.9M parameters and runs at 33 FPS.
Nianchang Huang, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Image Process.2
2021 Software/Hardware Co-Design Optimization for Sparse Convolutional Neural Networks
abstract
Deep convolutional neural network (DNN) has been widely used in image recognition, target detection, and natural language processing. Unfortunately, the previous state-of-the-art convolutional neural network (CNN) increase model accuracy by increasing the number of layers of the network, which leads to the problem of large models and complex calculations. Previous work has demonstrated that network compression can be achieved by weight pruning approaches, and that specialized hardware can be used to speed up reasoning. However, the format of compressed sparse rows (CSR) or compressed sparse columns (CSC) is generally used in model pruning. The approach leads to coding operations before calculation and decoding operations during calculation by processing units on specialized hardware. And the approach degrades the performance of the model.In this paper, we propose a combination of hardware and software to address above bottlenecks. Firstly, we propose a novel software-based structured pruning, called vector-pruning. Secondly, we design a novel dataflow for our pruning method, called Vector-Sparse, our method eliminates complex coding and decoding Then, we design the corresponding hardware architecture for vector-prune. Finally, we evaluate the performance of our pruning approach for AlexNet on Xilinx xcvu19p-fsva3824-2-e. Experiments show that compared with a state-of-the-art convolution network accelerator, we proposed a software/hardware co-design approach that is 1.01x and 1.3x better in performance, effectively improving the inference speed of the model on Field Programmable Gate Array (FPGA).
Wei Hu 0001, Yong Dong, Fang Liu 0031, Qiang Jiao
SMC4
2021 RISC-VTF: RISC-V Based Extended Instruction Set for Transformer
abstract
Deep learning model Transformer has been widely used in natural language processing(NLP) filed, and its demand for computing resources is also growing. However, general-purpose processors(CPU and GPU) have invested excessive hardware resources in the design because they have to flexibly support a variety of tasks, they are not efficient for the implementation of Transformer. Consequently, various software optimizations that towards general-purpose processors have been proposed one after another. But, under the condition of ensuring sufficient accuracy, the degree of software optimization is limited. It is requisite to friendly-support Transformer at the hardware level.After analysing the computational characteristics of the Transformer model, based on RISC-V, we designed a hardware friendly instruction set architecture for the Transformer model. In addition to the basic instruction, for the intensive and general computing part of the model, according to the expansion rules of RISC-V instruction, we design the matrix load/store instruction calculation instruction, softmax instruction, activation instruction and other user-defined instructions. They support any matrix scale, and deploy it on FPGA to realize a flexible and efficient custom processor RISC-VTF for Transformer. The design is integrated on the Xilinx toolkit zynq-7000 FPGA, and the resource consumption and performance are analyzed. Compared with the traditional common ISA(Instruction Set Architecture) such as x86, arm or MIPs, RISC-VTF provides higher code density and performance efficiency.
Qiang Jiao, Wei Hu 0001, Fang Liu 0031, Yong Dong
SMC1
2020 Optimizing Accelerator on FPGA for Deep Convolutional Neural Networks
Yong Dong, Wei Hu 0001, Yonghao Wang, Qiang Jiao
ICA3PP (2)4
2020 Design of a Convolutional Neural Network Instruction Set Based on RISC-V and Its Microarchitecture Implementation
Qiang Jiao, Wei Hu 0001, Yuan Wen, Yong Dong, Zhenhao Li 0004, Yu Gan 0004
ICA3PP (2)1
2019 Pinning Controllers for Activation Output Tracking of Boolean Network Under One-Bit Perturbation
abstract
This paper studies pinning controllers for activation output tracking (AOT) of Boolean network under one-bit perturbation, based on the semitensor product of matrices. First, the definition of AOT with respect to an activation number is presented, where the activation number means the number of active outputs whose logical variables are 1 s. Then, several criteria are established for AOT issue. Further, the impact of one-bit perturbation on AOT is studied, where one-bit perturbation means that only one logical function has one-bit change of its truth table by flipping the value from 1 to 0 or 0 to 1. In addition, if a one-bit perturbation is a valid perturbation on AOT, an output feedback pinning control is designed to recover AOT. The obtained results are effectively illustrated by a D. melanogaster segmentation polarity gene network and a reduced signal transduction network.
Jie Zhong 0005, Daniel W. C. Ho, Jianquan Lu, Qiang Jiao
IEEE Trans. Cybern.4