VLDB 2026 Research / reviewers in the wild / expert
Qiang Jiao
dblp:166/3845
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-4725-5284ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QANet: Query-Aware Multi-Modal Prior Refinement for Few-Shot SegmentationabstractRecent Few-Shot Segmentation (FSS) approaches incorporate vision-language models to improve segmentation by constructing multi-modal priors, including visual, CAM and textual priors. However, these priors are often suboptimal, suffering from background interference, incomplete activation and semantic misalignment. To address these limitations, we propose a query-aware multi-modal prior refinement framework (QANet) for FSS, in which the multi-modal priors are jointly refined for more accurate segmentation by fully exploiting target-relevant cues from query features. Specifically, QANet comprises two modules, i.e., a Query-aware Multi-modal Prior Refinement (QMPR) module for visual and CAM prior refinement, and a Query-Aware Textual Embedding and Prior Refinement (QTEPR) module for textual prior refinement. More specifically, in QMPR, a query self-attention bilateral-guided enhancement strategy and a query self-similarity prior refinement strategy are carefully designed for suppressing background interferences and recovering foreground completeness, respectively. And in QTEPR, such refined visual and CAM priors are first leveraged to extract target-relevant cues for visual-textual domain gap mitigation and text embedding enhancement, producing query-aware textual priors. These textual priors are then enhanced under the guidance of those refined multi-modal priors. Extensive experiments on PASCAL-5iand COCO-20idatasets demonstrate that QANet achieves new state-of-the-art performance. Moreover, QANet remains competitive on cross-domain FSS and weak-label FSS tasks. Qiang Jiao, Mengrui Shi, Qiang Zhang 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Resolving semantic conflicts in RGB-T semantic segmentation
Shenlu Zhao, Ziniu Jin, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
Pattern Recognit. | 3 |
| 2024 | Deep unsupervised shadow detection with curriculum learning and self-training
Qiang Zhang 0020, Hongyuan Guo, Guanghe Li, Tianlu Zhang, Qiang Jiao |
Comput. Vis. Image Underst. | 5 |
| 2024 | AMNet: Learning to Align Multi-Modality for RGB-T TrackingabstractRGB-T tracking has attracted increasing attention recently due to the all-weather and all-day working capability. However, most current RGB-T trackers usually assume that RGB data and thermal infrared (TIR) data are well spatially aligned, which is difficult to be achieved in practice. Such spatial misalignment between RGB data and TIR data may lead to the ineffective cross-modal information propagation during multi-modal feature fusion, thus reducing the tracking performance. In addition, due to the discrepancy in imaging characteristics of RGB images and TIR images, there also exist great differences between the information captured by the two modality data. The differences in characteristics of RGB and TIR modalities in different local areas will cause a single fusion strategy to be unable to fully explore the complementary information within multi-modal data. For that, we propose an RGB-T tracker, referred to as AMNet, to specifically solve such two problems with two dedicated modules, i.e., a Mutual-interacted Spatial Alignment (MSA) module and an Information Matching Fusion (IMF) module. The former spatially aligns the two modality data through three essential parts, including interactions of multi-modal features, prediction of cross-modal offset map, and enhancement of the aligned features. While the latter first discriminates different types of local regions by employing several intra-modal attention modules and then uses a divide-and-conquer fusion strategy to exploit such discriminative information within RGB and TIR features of different cases for tracking. We validate the effectiveness of our AMNet with extensive experiments on three RGB-T benchmarks, which achieves new state-of-the-art performance. Tianlu Zhang, Xiaoyi He, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Feature Calibrating and Fusing Network for RGB-D Salient Object DetectionabstractDue to their imaging mechanisms and techniques, some depth images inevitably have low visual qualities or have some inconsistent foregrounds with their corresponding RGB images. Directly using such depth images will deteriorate the performance of RGB-D SOD. In view of this, a novel RGB-D salient object detection model is presented, which follows the principle of calibration-then-fusion to effectively suppress the influence of such two types of depth images on final saliency prediction. Specifically, the proposed model is composed of two stages, i.e., an image generation stage and a saliency reasoning stage. The former generates high-quality and foreground-consistent pseudo depth images via an image generation network. While the latter first calibrates the original depth information with the aid of those newly generated pseudo depth images and then performs cross-modal feature fusion for the final saliency reasoning. Especially, in the first stage, a Two-steps Sample Selection (TSS) strategy is employed to select such reliable depth images from the original RGB-D image pairs as supervision information to optimize the image generation network. Afterwards, in the second stage, a Feature Calibrating and Fusing Network (FCFNet) is proposed to achieve the calibration-then-fusion of cross-modal information for the final saliency prediction, which is achieved by a Depth Feature Calibration (DFC) module, a Shallow-level Feature Injection (SFI) module and a Multi-modal Multi-scale Fusion (MMF) module. Moreover, a loss function, i.e., Region Consistency Aware (RCA) loss, is presented as an auxiliary loss for FCFNet to facilitate the completeness of salient objects together with the reduction of background interference by considering the local regional consistency in the saliency maps. Experiments on six benchmark datasets demonstrate the superiorities of our proposed RGB-D SOD model over some state-of-the-arts. Qiang Zhang 0020, Yang Yang 0132, Qiang Jiao, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Exploring Multi-Modal Spatial-Temporal Contexts for High-Performance RGB-T TrackingabstractIn RGB-T tracking, there exist rich spatial relationships between the target and backgrounds within multi-modal data as well as sound consistencies of spatial relationships among successive frames, which are crucial for boosting the tracking performance. However, most existing RGB-T trackers overlook such multi-modal spatial relationships and temporal consistencies within RGB-T videos, hindering them from robust tracking and practical applications in complex scenarios. In this paper, we propose a novel Multi-modal Spatial-Temporal Context (MMSTC) network for RGB-T tracking, which employs a Transformer architecture for the construction of reliable multi-modal spatial context information and the effective propagation of temporal context information. Specifically, a Multi-modal Transformer Encoder (MMTE) is designed to achieve the encoding of reliable multi-modal spatial contexts as well as the fusion of multi-modal features. Furthermore, a Quality-aware Transformer Decoder (QATD) is proposed to effectively propagate the tracking cues from historical frames to the current frame, which facilitates the object searching process. Moreover, the proposed MMSTC network can be easily extended to various tracking frameworks. New state-of-the-art results on five prevalent RGB-T tracking benchmarks demonstrate the superiorities of our proposed trackers over existing ones. Tianlu Zhang, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Image Process. | 2 |
| 2024 | Mitigating Modality Discrepancies for RGB-T Semantic SegmentationabstractSemantic segmentation models gain robustness against adverse illumination conditions by taking advantage of complementary information from visible and thermal infrared (RGB-T) images. Despite its importance, most existing RGB-T semantic segmentation models directly adopt primitive fusion strategies, such as elementwise summation, to integrate multimodal features. Such strategies, unfortunately, overlook the modality discrepancies caused by inconsistent unimodal features obtained by two independent feature extractors, thus hindering the exploitation of cross-modal complementary information within the multimodal data. For that, we propose a novel network for RGB-T semantic segmentation, i.e. MDRNet+, which is an improved version of our previous work ABMDRNet. The core of MDRNet+ is a brand new idea, termed the strategy of bridging-then-fusing, which mitigates modality discrepancies before cross-modal feature fusion. Concretely, an improved Modality Discrepancy Reduction (MDR+) subnetwork is designed, which first extracts unimodal features and reduces their modality discrepancies. Afterward, discriminative multimodal features for RGB-T semantic segmentation are adaptively selected and integrated via several channel-weighted fusion (CWF) modules. Furthermore, a multiscale spatial context (MSC) module and a multiscale channel context (MCC) module are presented to effectively capture the contextual information. Finally, we elaborately assemble a challenging RGB-T semantic segmentation dataset, i.e., RTSS, for urban scene understanding to mitigate the lack of well-annotated training data. Comprehensive experiments demonstrate that our proposed model surpasses other state-of-the-art models on the MFNet, PST900, and RTSS datasets remarkably. Shenlu Zhao, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Efficient RGB-T Tracking via Cross-Modality DistillationabstractMost current RGB-T trackers adopt a two-stream structure to extract unimodal RGB and thermal features and complex fusion strategies to achieve multi-modal feature fusion, which require a huge number of parameters, thus hindering their real-life applications. On the other hand, a compact RGB-T tracker may be computationally efficient but encounter non-negligible performance degradation, due to the weakening of feature representation ability. To remedy this situation, a cross-modality distillation framework is presented to bridge the performance gap between a compact tracker and a powerful tracker. Specifically, a specific-common feature distillation module is proposed to transform the modality-common information as well as the modality-specific information from a deeper two-stream network to a shallower single-stream network. In addition, a multi-path selection distillation module is proposed to instruct a simple fusion module to learn more accurate multi-modal information from a well-designed fusion mechanism by using multiple paths. We validate the effectiveness of our method with extensive experiments on three RGB-T benchmarks, which achieves state-of-the-art performance but consumes much less computational resources. Tianlu Zhang, Hongyuan Guo, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
CVPR | 3 |
| 2022 | Middle-Level Feature Fusion for Lightweight RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) models adopt a two-stream structure to extract the information from the input RGB and depth images. Since they use two subnetworks for unimodal feature extraction and multiple multi-modal feature fusion modules for extracting cross-modal complementary information, these models require a huge number of parameters, thus hindering their real-life applications. To remedy this situation, we propose a novel middle-level feature fusion structure that allows to design a lightweight RGB-D SOD model. Specifically, the proposed structure first employs two shallow subnetworks to extract low- and middle-level unimodal RGB and depth features, respectively. Afterward, instead of integrating middle-level unimodal features multiple times at different layers, we just fuse them once via a specially designed fusion module. On top of that, high-level multi-modal semantic features are further extracted for final salient object detection via an additional subnetwork. This will greatly reduce the network's parameters. Moreover, to compensate for the performance loss due to parameter deduction, a relation-aware multi-modal feature fusion module is specially designed to effectively capture the cross-modal complementary information during the fusion of middle-level multi-modal features. By enabling the feature-level and decision-level information to interact, we maximize the usage of the fused cross-modal middle-level features and the extracted cross-modal high-level features for saliency prediction. Experimental results on several benchmark datasets verify the effectiveness and superiority of the proposed method over some state-of-the-art methods. Remarkably, our proposed model has only 3.9M parameters and runs at 33 FPS. Nianchang Huang, Qiang Jiao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Image Process. | 2 |
| 2021 | Software/Hardware Co-Design Optimization for Sparse Convolutional Neural NetworksabstractDeep convolutional neural network (DNN) has been widely used in image recognition, target detection, and natural language processing. Unfortunately, the previous state-of-the-art convolutional neural network (CNN) increase model accuracy by increasing the number of layers of the network, which leads to the problem of large models and complex calculations. Previous work has demonstrated that network compression can be achieved by weight pruning approaches, and that specialized hardware can be used to speed up reasoning. However, the format of compressed sparse rows (CSR) or compressed sparse columns (CSC) is generally used in model pruning. The approach leads to coding operations before calculation and decoding operations during calculation by processing units on specialized hardware. And the approach degrades the performance of the model.In this paper, we propose a combination of hardware and software to address above bottlenecks. Firstly, we propose a novel software-based structured pruning, called vector-pruning. Secondly, we design a novel dataflow for our pruning method, called Vector-Sparse, our method eliminates complex coding and decoding Then, we design the corresponding hardware architecture for vector-prune. Finally, we evaluate the performance of our pruning approach for AlexNet on Xilinx xcvu19p-fsva3824-2-e. Experiments show that compared with a state-of-the-art convolution network accelerator, we proposed a software/hardware co-design approach that is 1.01x and 1.3x better in performance, effectively improving the inference speed of the model on Field Programmable Gate Array (FPGA). Wei Hu 0001, Yong Dong, Fang Liu 0031, Qiang Jiao |
SMC | 4 |
| 2021 | RISC-VTF: RISC-V Based Extended Instruction Set for TransformerabstractDeep learning model Transformer has been widely used in natural language processing(NLP) filed, and its demand for computing resources is also growing. However, general-purpose processors(CPU and GPU) have invested excessive hardware resources in the design because they have to flexibly support a variety of tasks, they are not efficient for the implementation of Transformer. Consequently, various software optimizations that towards general-purpose processors have been proposed one after another. But, under the condition of ensuring sufficient accuracy, the degree of software optimization is limited. It is requisite to friendly-support Transformer at the hardware level.After analysing the computational characteristics of the Transformer model, based on RISC-V, we designed a hardware friendly instruction set architecture for the Transformer model. In addition to the basic instruction, for the intensive and general computing part of the model, according to the expansion rules of RISC-V instruction, we design the matrix load/store instruction calculation instruction, softmax instruction, activation instruction and other user-defined instructions. They support any matrix scale, and deploy it on FPGA to realize a flexible and efficient custom processor RISC-VTF for Transformer. The design is integrated on the Xilinx toolkit zynq-7000 FPGA, and the resource consumption and performance are analyzed. Compared with the traditional common ISA(Instruction Set Architecture) such as x86, arm or MIPs, RISC-VTF provides higher code density and performance efficiency. Qiang Jiao, Wei Hu 0001, Fang Liu 0031, Yong Dong |
SMC | 1 |
| 2020 | Optimizing Accelerator on FPGA for Deep Convolutional Neural Networks
Yong Dong, Wei Hu 0001, Yonghao Wang, Qiang Jiao |
ICA3PP (2) | 4 |
| 2020 | Design of a Convolutional Neural Network Instruction Set Based on RISC-V and Its Microarchitecture Implementation
Qiang Jiao, Wei Hu 0001, Yuan Wen, Yong Dong, Zhenhao Li 0004, Yu Gan 0004 |
ICA3PP (2) | 1 |
| 2019 | Pinning Controllers for Activation Output Tracking of Boolean Network Under One-Bit PerturbationabstractThis paper studies pinning controllers for activation output tracking (AOT) of Boolean network under one-bit perturbation, based on the semitensor product of matrices. First, the definition of AOT with respect to an activation number is presented, where the activation number means the number of active outputs whose logical variables are 1 s. Then, several criteria are established for AOT issue. Further, the impact of one-bit perturbation on AOT is studied, where one-bit perturbation means that only one logical function has one-bit change of its truth table by flipping the value from 1 to 0 or 0 to 1. In addition, if a one-bit perturbation is a valid perturbation on AOT, an output feedback pinning control is designed to recover AOT. The obtained results are effectively illustrated by a D. melanogaster segmentation polarity gene network and a reduced signal transduction network. Jie Zhong 0005, Daniel W. C. Ho, Jianquan Lu, Qiang Jiao |
IEEE Trans. Cybern. | 4 |