Qiang Qi

dblp:38/10180 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel Context
abstract
Recent research has focused on using convolutional neural networks (CNNs) as the backbones in two-view correspondence learning, demonstrating significant superiority over methods based on multilayer perceptrons. However, CNN backbones that are not tailored to specific tasks may fail to effectively aggregate global context and oversmooth dense motion fields in scenes with large disparity. To address these problems, we propose a novel network named SC-Net, which effectively integrates bilateral context from both spatial and channel perspectives. Specifically, we design an adaptive focused regularization module (AFR) to enhance the model's position-awareness and robustness against spurious motion samples, thereby facilitating the generation of a more accurate motion field. We then propose a bilateral field adjustment module (BFA) to refine the motion field by simultaneously modeling long-range relationships and facilitating interaction across spatial and channel dimensions. Finally, we recover the motion vectors from the refined field using a position-aware recovery module (PAR) that ensures consistency and precision. Extensive experiments demonstrate that SC-Net outperforms state-of-the-art methods in relative pose estimation and outlier removal tasks on YFCC100M and SUN3D datasets.
Shuyuan Lin, Hailiang Liao, Qiang Qi, Taotao Lai, Jian Weng 0001
AAAI3
2026 MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud Registration
abstract
Robust and discriminative feature learning is critical for high-quality point cloud registration. However, existing deep learning–based methods typically rely on Euclidean neighborhood-based strategies for feature extraction, which struggle to effectively capture the implicit semantics and structural consistency in point clouds. To address these issues, we propose a multi-domain context integration network (MCI-Net) that improves feature representation and registration performance by aggregating contextual cues from diverse domains. Specifically, we propose a graph neighborhood aggregation module, which constructs a global graph to capture the overall structural relationships within point clouds. We then propose a progressive context interaction module to enhance feature discriminability by performing intra-domain feature decoupling and inter-domain context interaction. Finally, we design a dynamic inlier selection method that optimizes inlier weights using residual information from multiple iterations of pose estimation, thereby improving the accuracy and robustness of registration. Extensive experiments on indoor RGB-D and outdoor LiDAR datasets show that the proposed MCI-Net significantly outperforms existing state-of-the-art methods, achieving the highest registration recall of 96.4% on 3DMatch.
Shuyuan Lin, Wenwu Peng, Qiang Qi, Miaohui Wang, Jian Weng 0001
AAAI4
2026 MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object Detection
abstract
Video object detection is a fundamental yet challenging task in computer vision. Recently, DETR-based methods have gained prominence in this domain owing to their powerful global modeling capabilities. However, these methods are still confronted with two key limitations: frame-agnostic initialization of object queries and scale-agnostic attention mechanisms, which hinder their capability to capture the appearance variations of dynamic objects and model the temporal consistency across frames. To alleviate these limitations, we propose a multiscale-aware transformer diffusion network (MSTDiff), a novel framework designed for the video object detection task, including two technical improvements over existing methods. First, we design a diffusion-driven adaptive query module, which models the object query distribution through a diffusion process conditioned on input frames, enabling an adaptive and content-aware initialization of object queries. Second, we develop a multiscale-aware transformer encoder module, which combines multi-head convolutional units with attention mechanisms to enhance multi-scale feature representations while preserving global dependence modeling. We conduct extensive experiments on the public ImageNet VID dataset, and the results demonstrate that our MSTDiff achieves 87.7% mAP with ResNet-101, outperforming most previous state-of-the-art video object detection methods.
Qiang Qi, Wenqi Shang, Shuyuan Lin
AAAI1
2026 Multi-agent reinforcement learning method for joint optimization of block assignment and yard crane redeployment at river-sea intermodal container terminal
Huakun Liu, Shuzheng Yang, Hongbin Tian, Qiang Qi
Adv. Eng. Informatics6
2026 HAVEN: Hierarchical diffusion and value-based trajectory selection for offline safe reinforcement learning
Erlie Wang, He Diao, Xianglin Chen, Jingkui Zhang, Xiaofeng Chai, Qiang Qi, Ping Zhang 0023
Neurocomputing6
2025 TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
abstract
Video object detection has made significant progress in recent years thanks to convolutional neural networks (CNNs) and vision transformers (ViTs). Typically, CNNs excel at capturing local features but struggle to model global representations. Conversely, ViTs are adept at capturing long-range global features but face challenges in representing local feature details. Off-the-shelf video object detection methods solely rely on CNNs or ViTs to conduct feature aggregation, which hampers their capability to simultaneously leverage global and local information, thereby resulting in limited detection performance. In this paper, we propose a Transformer-GraphFormer Blender Network (TGBFormer) for video object detection, with three key technical improvements to fully exploit the advantages of transformers and graph convolutional networks while compensating for their limitations. First, we develop a spatial-temporal transformer module to aggregate global contextual information, constituting global representations with long-range feature dependencies. Second, we introduce a spatial-temporal GraphFormer module that utilizes local spatial and temporal relationships to aggregate features, generating new local representations that are complementary to the transformer outputs. Third, we design a global-local feature blender module to adaptively couple transformer-based global representations and GraphFormer-based local representations. Extensive experiments demonstrate that our TGBFormer establishes new state-of-the-art results on the ImageNet VID dataset. Particularly, our TGBFormer achieves 86.5% mAP while running at around 41.0 FPS on a single Tesla A100 GPU.
Qiang Qi, Xiao Wang 0001
AAAI1
2025 IMC-Det: Intra-Inter Modality Contrastive Learning for Video Object Detection
Qiang Qi, Zhenyu Qiu, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
Int. J. Comput. Vis.1
2025 DGC-Net: Dynamic Graph Contrastive Network for Video Object Detection
abstract
Video object detection is a challenging task in computer vision since it needs to handle the object appearance degradation problem that seldom occurs in the image domain. Off-the-shelf video object detection methods typically aggregate multi-frame features at one stroke to alleviate appearance degradation. However, these existing methods do not take supervision knowledge into consideration and thus still suffer from insufficient feature aggregation, resulting in the false detection problem. In this paper, we take a different perspective on feature aggregation, and propose a dynamic graph contrastive network (DGC-Net) for video object detection, including three improvements against existing methods. First, we design a frame-level graph contrastive module to aggregate frame features, enabling our DGC-Net to fully exploit discriminative contextual feature representations to facilitate video object detection. Second, we develop a proposal-level graph contrastive module to aggregate proposal features, making our DGC-Net sufficiently learn discriminative semantic feature representations. Third, we present a graph transformer to dynamically adjust the graph structure by pruning the useless nodes and edges, which contributes to improving accuracy and efficiency as it can eliminate the geometric-semantic ambiguity and reduce the graph scale. Furthermore, inherited from the framework of DGC-Net, we develop DGC-Net Lite to perform real-time video object detection with a much faster inference speed. Extensive experiments conducted on the ImageNet VID dataset demonstrate that our DGC-Net outperforms the performance of current state-of-the-art methods. Notably, our DGC-Net obtains 86.3%/87.3% mAP when using ResNet-101/ResNeXt-101.
Qiang Qi, Hanzi Wang, Yan Yan 0001, Xuelong Li 0001
IEEE Trans. Image Process.1
2024 Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object Detection
abstract
Video object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for lacking of motion robustness. We propose a novel proposal distillation of multi-modal feature aggregation network (PDMAN). Specially, it initially aligns the feature domain and flow domain via a lightweight flow module (LFM) and then facilities frame-level feature aggregation. Subsequently, a global-based semantic embedding module (GSEM) is designed to incorporate global semantic features into instance features and introduce a global multi-label classification loss to guide encoding with high class-wise responsiveness. Finally, to alleviate the presence of insufficient and redundant information in multi-modal instance-level feature aggregation, a proposal distilled aggregation module (PDAM) is employed. By distilling the instance set, this approach realizes a fine-grained feature aggregation, ultimately boosting the detection performance. Experimental results demonstrate that the proposed PDMAN achieves a favorable result on the most representative large-scale ImageNet VID dataset.
Zhenyu Qiu, Qiang Qi, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
ICASSP2
2024 Class-Aware Dual-Supervised Aggregation Network for Video Object Detection
abstract
Video object detection has attracted increasing attention in recent years. Although great success has been achieved by off-the-shelf video object detection methods through delicately designing various types of feature aggregation, they overlook the class-aware supervision and thus still suffer from the problem of classification incapability, which means the classification between objects with deteriorated or similar appearances is error-prone. In this article, we propose a novel class-aware dual-supervised aggregation network (CDANet) for video object detection, including three substantial improvements to effectively alleviate the classification incapability problem of previous methods. First, we develop a class-aware cross-modality distillation supervision that transfers the semantic knowledge of label data to the features of video data, effectively enhancing the semantic representations of features. Second, we design a graph-guided feature aggregation module that effectively models the structural relations between features by leveraging the dynamic residual graph convolutional network, enabling our CDANet to perform more effective feature aggregation in the temporal domain. Third, we present a class-aware proposal contrastive supervision to maximize the intra-class agreement and inter-class disagreement, which is conducive to improving the semantic discriminability of features. The class-aware dual supervision and feature aggregation are tightly tied into a unified end-to-end framework to make our CDANet fully exploit class-specific semantic knowledge and inter-frame temporal dependencies to enhance object appearance representations, which facilitates the classification of detected objects. We conduct experiments on the challenging ImageNet VID dataset, and the results demonstrate the superiority of our CDANet against state-of-the-art methods. More remarkably, our CDANet achieves 85.4% mAP with ResNet-101 or 86.5% mAP with ResNeXt-101.
Qiang Qi, Yan Yan 0001, Hanzi Wang
IEEE Trans. Multim.1
2023 DF-Net: Diversity-Focused Network for Video Object Detection
abstract
Video object detection is a challenging task due to deteriorated object appearances. To enhance per-frame features, one way is to aggregate features from several support frames. However, proposals generated by the region proposal network may not be precise and diverse due to the fixed anchors, limiting the detection performance. We propose a novel architecture called Diversity-Focused Network (DF-Net), which consists of three modules: 1) An affine transform module (ATM), which is proposed to model the deblurring process and fuse the feature maps of different receptive fields by a multi-level attention block; 2) A label assignment module (LAM), which is proposed to assign the labels to the proposals used in a fine-grained aggregation manner; 3) A regression-guided diffusion module (RGDM), which is proposed to obtain the features of diversity and higher quality. Experiments show that DF-Net achieves favorable results on the most representative large-scale ImageNet VID dataset. Remarkably, the DF-Net achieves 84.8% mAP with ResNet-101 without post-processing steps.
Zhenyu Qiu, Qiang Qi, Yan Yan 0001, Hanzi Wang
ICIP2
2023 TCNet: A Novel Triple-Cooperative Network for Video Object Detection
abstract
Video object detection aims at accurately localizing the objects in videos and correctly recognizing their categories. Off-the-shelf video object detection methods have made some progress in recent years but they still suffer from the problems of inaccurate object localization, incorrect object recognition or insufficient relation learning, resulting in limited detection performance. In this paper, we propose a novel triple-cooperative network (TCNet) for high-performance video object detection, with three substantial improvements to ameliorate the problems of existing methods. First, we develop a context-aware proposal refinement module to generate high-quality proposals, enabling our TCNet to achieve more accurate object localization. Second, we present a similarity-aware semantic distillation module that innovatively leverages the semantic knowledge of class labels as additional supervisory signals to enhance the object recognition ability of our TCNet. Third, we design a structure-aware relation learning module to effectively model the structural relations between features with an adaptive-pruning residual graph convolutional network, making our TCNet perform more effective feature aggregation. We conduct extensive experiments on the challenging ImageNet VID dataset and the experimental results demonstrate that our TCNet outperforms current state-of-the-art methods. More remarkably, our TCNet achieves 85.2% mAP and 86.3% mAP with ResNet-101 and ResNeXt-101, respectively.
Qiang Qi, Tianxiang Hou, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.1
2023 DGRNet: A Dual-Level Graph Relation Network for Video Object Detection
abstract
Video object detection is a fundamental and important task in computer vision. One mainstay solution for this task is to aggregate features from different frames to enhance the detection on the current frame. Off-the-shelf feature aggregation paradigms for video object detection typically rely on inferring feature-to-feature (Fea2Fea) relations. However, most existing methods are unable to stably estimate Fea2Fea relations due to the appearance deterioration caused by object occlusion, motion blur or rare poses, resulting in limited detection performance. In this paper, we study Fea2Fea relations from a new perspective, and propose a novel dual-level graph relation network (DGRNet) for high-performance video object detection. Different from previous methods, our DGRNet innovatively leverages the residual graph convolutional network to simultaneously model Fea2Fea relations at two different levels including frame level and proposal level, which facilitates performing better feature aggregation in the temporal domain. To prune unreliable edge connections in the graph, we introduce a node topology affinity measure to adaptively evolve the graph structure by mining the local topological information of pairwise nodes. To the best of our knowledge, our DGRNet is the first video object detection method that leverages dual-level graph relations to guide feature aggregation. We conduct experiments on the ImageNet VID dataset and the results demonstrate the superiority of our DGRNet against state-of-the-art methods. Especially, our DGRNet achieves 85.0% mAP and 86.2% mAP with ResNet-101 and ResNeXt-101, respectively.
Qiang Qi, Tianxiang Hou, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
IEEE Trans. Image Process.1
2022 Dual Selection Network for Video Object Detection
abstract
Some off-the-shelf video object detection methods usually enhance the degraded proposal features of target frames by aggregating the proposal features from support frames. However, the proposals generated by region proposal network may not be accurate, resulting in inaccurate proposal features and limited performance. To mitigate this, we propose a novel dual selection network (DSNet) for video object detection, which contains two successive stages: selecting proposals that fit objects more closely, and selecting proposal features that are more conducive to feature aggregation. Correspondingly, the proposal selection module (PSM) aims to select better proposals by exploiting their boundary information, and the selective aggregation module (SAM) aims to select better proposal features for aggregation. Consequently, DSNet can generate more robust proposal features through the novel dual selection mechanism implemented by PSM and SAM. Extensive experiments show that our DSNet obtains 83.7% mAP and achieves superior performance over several state-of-the-art methods.
Tianxiang Hou, Qiang Qi, Yang Lu 0009, Kaiwen Du, Hanzi Wang
ICME2
2022 Fuzzy Optimal Tracking Control of Hypersonic Flight Vehicles via Single-Network Adaptive Critic Design
abstract
Optimal performance is extremely important for hypersonic flight control. Different from most existing methodologies, which only consider basic control performance including stability, robustness, and transient performance, this article deals with the design of nearly optimal tracking controllers for hypersonic flight vehicles (HFVs). First, main controllers are developed for the velocity subsystem and the altitude subsystem of HFVs via concise fuzzy approximations. Then, optimal controllers are nearly implemented utilizing single-network adaptive critic design. Moreover, the stability of closed-loop systems and the convergence of optimal controllers are theoretically proved. Finally, compared simulation results are given to verify the superiority. The special contribution is the application of a low-complex control structure owing to the critic-only network and advanced learning laws developed for fuzzy approximations, which is expected to guarantee satisfied real-time performance.
Xiangwei Bu, Qiang Qi
IEEE Trans. Fuzzy Syst.2
2022 A Simplified Finite-Time Fuzzy Neural Controller With Prescribed Performance Applied to Waverider Aircraft
abstract
This article addresses a finite-time prescribed performance controller within the concise fuzzy-neural framework with application to a waverider aircraft. First, new finite-time performance functions are developed to construct a constraint funnel, which accomplishes that tracking errors converge to their steady-state values in a given time (i.e., finite-time convergence), being expected to guarantee tracking errors with small overshoots. Then, the equivalent transformation approach is introduced to unify unknown dynamics such that the control complexity is reduced. Moreover, to further reduce computational costs, a single-learning-parameter-based regulation scheme is developed for fuzzy-neural approximation. Finally, the proposed method is applied to a waverider aircraft to test its effectiveness and superiority.
Xiangwei Bu, Qiang Qi, Baoxu Jiang
IEEE Trans. Fuzzy Syst.2
2022 FastVOD-Net: A Real-Time and High-Accuracy Video Object Detector
abstract
Video object detection is a tough task due to the severe appearance degradation caused by rapid motion, sudden occlusion or rare poses. The great challenge facing video object detection is the simultaneous requirements on both accuracy and speed because the pursuit of one aspect usually causes significant expense to the other. Most existing methods mainly focus on improving detection accuracy with little attention to computationally efficient solutions, and thus they are impractical for many real-world applications. This motivates us to develop a real-time and high-accuracy video object detection method. In this paper, we propose a novel video object detector, called FastVOD-Net, which can yield highly accurate detection results at real-time speed. Specifically, we first develop a temporally-cascaded deformable alignment (TCDA) module to model the object displacements induced by video motion. Then, we introduce another two modules, namely spatially-refined temporal aggregation (SRTA) and attention-guided semantic distillation (AGSD), to improve the appearance feature of the currently processed frame and enhance the semantic representation of non-keyframes, respectively. For keyframe scheduling, we design an adaptive keyframe selection scheduler (AKSS) to adjust the keyframe interval online, making the keyframe usage more rational. On one hand, the characteristics of our FastVOD-Net enable it to sparsely perform expensive feature extraction, which significantly reduces the computational cost and thus guarantees real-time speed. On the other hand, the collaboration of the above tightly-coupled modules and adaptive keyframe scheduler makes FastVOD-Net fully exploit inter-frame temporal dependencies and thus guarantees high accuracy. Experiments on the ImageNet VID dataset show that our FastVOD-Net achieves 79.3% mAP at 29.6 fps or 81.2% mAP at 23.0 fps on an Nvidia RTX 2080 Ti GPU, which is the state-of-the-art performance in real time.
Qiang Qi, Xiao Wang 0072, Tianxiang Hou, Yan Yan 0001, Hanzi Wang
IEEE Trans. Intell. Transp. Syst.1
2021 Prophet: Speeding up Distributed DNN Training with Predictable Communication Scheduling
abstract
Optimizing performance for Distributed Deep Neural Network (DDNN) training has recently become increasingly compelling, as the DNN model gets complex and the training dataset grows large. While existing works on communication scheduling mostly focus on overlapping the computation and communication to improve DDNN training performance, the GPU and network resources are still under-utilized in DDNN training clusters. To tackle this issue, in this paper, we design and implement a predictable communication scheduling strategy named Prophet to schedule the gradient transfer in an adequate order, with the aim of maximizing the GPU and network resource utilization. Leveraging our observed stepwise pattern of gradient transfer start time, Prophet first uses the monitored network bandwidth and the profiled time interval among gradients to predict the appropriate number of gradients that can be grouped into blocks. Then, these gradient blocks can be transferred one by one to guarantee high utilization of GPU and network resources while ensuring the priority of gradient transfer (i.e., low-priority gradients cannot preempt high-priority gradients in the network transfer). Prophet can make the forward propagation start as early as possible so as to greedily reduce the waiting (idle) time of GPU resources during the DDNN training process. Prototype experiments with representative DNN models trained on Amazon EC2 demonstrate that Prophet can improve the DDNN training performance by up to 40% compared with the state-of-the-art priority-based communication scheduling strategies, yet with negligible runtime performance overhead.
Qiang Qi, Ruitao Shang, Li Chen 0019, Fei Xu 0009
ICPP2
2021 Rationing bandwidth resources for mitigating network resource contention in distributed DNN training clusters
Qiang Qi, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006
CCF Trans. High Perform. Comput.1
2020 Global and local feature alignment for video object detection
abstract
Extending image-based object detectors into video domain suffers from immense inadaptability due to the deteriorated frames caused by motion blur, partial occlusion or strange poses. Therefore, the generated features of deteriorated frames encounter the poor quality of misalignment, which degrades the overall performance of video object detectors. How to capture valuable information locally or globally is of importance to feature alignment but remains quite challenging. In this paper, we propose a Global and Local Feature Alignment (abbreviated as GLFA) module for video object detection, which can distill both global and local information to excavate the deep relationship between features for feature alignment. Specifically, GLFA can model the spatial-temporal dependencies over frames based on propagating global information and capture the interactive correspondences within the same frame based on aggregating valuable local information. Moreover, we further introduce a Self-Adaptive Calibration (SAC) module to strengthen the semantic representation of features and distill valuable local information in a dual local-alignment manner. Experimental results on the ImageNet VID dataset show that the proposed method achieves high performance as well as a good trade-off between real-time speed and competitive accuracy.
Haihui Ye, Qiang Qi, Yang Lu 0009, Hanzi Wang
MMAsia2
2018 Integrating QDWD with pattern distinctness and local contrast for underwater saliency detection
Muwei Jian, Qiang Qi, Junyu Dong, Yilong Yin, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.2
2018 Saliency detection using quaternionic distance based weber local descriptor and level priors
Muwei Jian, Qiang Qi, Junyu Dong, Xin Sun 0003, Yujuan Sun, Kin-Man Lam 0001
Multim. Tools Appl.2
2017 The OUC-vision large-scale underwater image database
abstract
In this paper, a large-scale underwater image database for underwater salient object detection or saliency detection is presented in detail. This database is called the OUC-VISION underwater image database, which contains 4400 underwater images of 220 individual objects. Each object is captured with four pose variations (the frontal-, the opposite-, the left-, and the right-views of each underwater object) and five spatial locations (the underwater object is located at the top-left corner, the top-right corner, the center, the bottom-left corner, and the bottom-right corner) to obtain 20 images. Meanwhile, this publicly available OUC-VISION database also provides relevant industrial fields, and academic researchers with underwater images under different sources of variations, especially pose, spatial location, illumination, turbidity of water, etc. Ground-truth information is also manually labelled for this database. The OUC-VISION database can not only be widely used to assess and evaluate the performance of the state-of-the-art salient-object detection and saliency-detection algorithms for general images, but also will particularly benefit the development of underwater vision technology in the future.
Muwei Jian, Qiang Qi, Junyu Dong, Yinlong Yin, Wenyin Zhang, Kin-Man Lam 0001
ICME2