Zhiqiang Zhou 0001

dblp:60/8343-1 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0001-6871-8236ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding Space
abstract
Infrared-visible image fusion (IVIF) has attracted much attention owing to the highly-complementary properties of the two image modalities. Due to the lack of ground-truth fused images, the fusion output of current deep-learning based methods heavily depends on the loss functions defined mathematically. As it is hard to well mathematically define the fused image without ground truth, the performance of existing fusion methods is limited. In this paper, we propose to use natural language to express the objective of IVIF, which can avoid the explicit mathematical modeling of fusion output in current losses, and make full use of the advantage of language expression to improve the fusion performance. For this purpose, we present a comprehensive language-expressed fusion objective, and encode relevant texts into the multi-modal embedding space using CLIP. A language-driven fusion model is then constructed in the embedding space, by establishing the relationship among the embedded vectors representing the fusion objective and input image modalities. Finally, a language-driven loss is derived to make the actual IVIF aligned with the embedded language-driven fusion model via supervised training. Experiments show that our method can obtain much better fusion results than existing techniques.
Lingjuan Miao, Zhiqiang Zhou 0001, Lei Zhang 0243, Yajun Qiao
ACM Multimedia3
2024 Gradient Calibration Loss for Fast and Accurate Oriented Bounding Box Regression
abstract
Oriented object detection has a very wide range of application scenarios. In recent years, a lot of rotation detectors have been designed to achieve high-performance oriented object detection. Intersection-over-Union (IoU) is the commonly used indicator to evaluate the accuracy of detection performance. Many methods introduce IoU into the bounding box regression loss to achieve the aligned training and evaluation process for better performance. However, in this paper, we demonstrate several drawbacks of rotated IoU loss through both experiments and theoretical derivation: 1) There is a negative correlation between the loss gradient and the angular error. 2) The optimization process of rotated IoU loss suffers from scale sensitivity, which is not conducive to the model convergence. To solve the problems, we propose a Gradient Calibration Loss (GCL) that optimizes the rotated IoU loss via gradient analysis and correction. We construct the optimized gradient in GCL to avoid IoU loss oscillation and scale sensitivity, thereby accelerating model convergence. Models supervised by GCL have a more stable training process, faster convergence, and better performance. Moreover, GCL can be easily introduced into the existing rotation detectors to achieve performance gains without extra inference overhead. Extensive experiments on multiple oriented object detection datasets and models demonstrate the superiority of our method. Our method achieves state-of-the-art performance on the mainstream benchmark datasets. The source code and models are available at https://github.com/ming71/GCL.
Qi Ming, Lingjuan Miao, Zhiqiang Zhou 0001, Junjie Song, Aleksandra Pizurica
IEEE Trans. Geosci. Remote. Sens.3
2024 Not All Boxes Are Equal: Learning to Optimize Bounding Boxes With Discriminative Distributions in Optical Remote Sensing Images
abstract
Detecting oriented objects in optical remote sensing images has been consistently challenging due to difficulties in bounding boxes localization. The cascaded regression framework, widely employed for high-quality bounding box refinement, has demonstrated effectiveness in this domain. However, our experiments reveal a discontinuity issue in bounding box optimization in cascaded regression framework. As a result, performance gain is not guaranteed across all stages in this framework. In this paper, we propose a Distribution Discriminative Detector(DDDet) to address the above issues and enhance the optimization of bounding boxes in oriented object detection. Specifically, a novel Conditional Anchor Refinement Framework(CARF) is designed to improve cascaded regression structure. CARF distinguishes bounding boxes with different distributions, adaptively optimizing them within the well-assigned regressors. Subsequently, the Aligned Convolution Module(ACM) is integrated into each regressor, facilitating the continuous alignment between features and refined anchors. Furthermore, the Geometry-guided Training Sample Selection(GTSS) method is incorporated into CARF to assign labels based on object shape priors. Experimental results show that DDDet obtains state-of-the-art performance on mainstream datasets for oriented object detection in remote sensing image, which demonstrates the effectiveness of the proposed method. Our method surpasses many current single-stage detectors, two-stage detectors, and refine-stage detectors, achieving the mAP of 79.41% on DOTA dataset, and 44.15% on FAIR1M dataset.
Qi Ming, Lingjuan Miao, Zhiqiang Zhou 0001, Nicolas Vercheval, Aleksandra Pizurica
IEEE Trans. Geosci. Remote. Sens.3
2023 Deep Dive into Gradients: Better Optimization for 3D Object Detection with Gradient-Corrected IoU Supervision
abstract
Intersection-over-Union (IoU) is the most popular metric to evaluate regression performance in 3D object detection. Recently, there are also some methods applying IoU to the optimization of 3D bounding box regression. However, we demonstrate through experiments and mathematical proof that the 3D IoU loss suffers from abnormal gradient w.r.t. angular error and object scale, which further leads to slow convergence and suboptimal regression process, respectively. In this paper, we propose a Gradient-Corrected IoU (GCIoU) loss to achieve fast and accurate 3D bounding box regression. Specifically, a gradient correction strategy is designed to endow 3D IoU loss with a reasonable gradient. It ensures that the model converges quickly in the early stage of training, and helps to achieve fine-grained refinement of bounding boxes in the later stage. To solve suboptimal regression of 3D IoU loss for objects at different scales, we introduce a gradient rescaling strategy to adaptively optimize the step size. Finally, we integrate GCIoU Loss into multiple models to achieve stable performance gains and faster model convergence. Experiments on KITTI dataset demonstrate superiority of the proposed method. The code is available at https://github.com/ming71/GCIoU-loss.
Qi Ming, Lingjuan Miao, Zhe Ma 0001, Zhiqiang Zhou 0001, Xuhui Huang, Yuanpei Chen, Yufei Guo 0001
CVPR5
2023 Optimized Point Set Representation for Oriented Object Detection in Remote-Sensing Images
abstract
How to represent the object more appropriately in oriented object detection is an essential problem to be solved, there are many solutions for the object represented. It is a relatively novel approach to represent objects as a number of sample points useful for both localization and recognition. However, the current point-set-based representation methods do not effectively supervise all points for learning, and the internal information of the convex hull in the point set cannot be effectively learned. Therefore, this letter proposes point set distance (PSD) loss, which learns set-to-set supervision of objects to effectively represent objects. Besides, most of the current sample selection strategies are based on the Intersection over Union (IoU), but these methods cannot comprehensively measure candidate samples quality. To select high-quality point sets, we propose to use the probability distribution of point sets to select the positive samples. Our probabilistic point set sample selection (PPSS) scheme effectively exploits the classification information, regression information, and the distribution characteristics of the point set. Experimental results on remote sensing image datasets including DOTA, DIOR-R, and HRSC2016, demonstrate the proposed method for arbitrary-oriented object detection achieves consistent and substantial improvements.
Junjie Song, Lingjuan Miao, Zhiqiang Zhou 0001, Qi Ming, Yunpeng Dong
IEEE Geosci. Remote. Sens. Lett.3
2023 A Novel Object Detector Based on High-Quality Rotation Proposal Generation and Adaptive Angle Optimization
abstract
Currently, reliable and accurate oriented detection in remote sensing images still needs to be improved. The wide variation of object shapes and orientations in the remote sensing images usually leads to two issues in two-stage oriented object detectors. One issue is how to generate high-quality rotation proposals. The other is the angular error sensitivity to the aspect ratios in the angle optimization process. In this paper, we propose a novel rotation proposal generation and optimization detector, which is based on high-quality rotation proposal generation and adaptive angle optimization to solve these two issues. The proposed method mainly establishes the geometric relationship guided region proposal networks (GRG-RPN) and the adaptive angle optimization head (AAO-Head) to achieve more accurate oriented object detection. The GRG-RPN only uses a simple network and a small number of horizontal anchors to predict high-quality rotation proposals. This approach was derived via the calculation based on the theoretical analysis of the geometric relationship between the oriented bounding boxes (OBBs) and their external horizontal bounding boxes (EHBBs). The AAO-Head solves the angular error sensitivity to the aspect ratios and achieves adaptive angle optimization using a new regression parameter, which is defined based on the theoretical analysis of the relationship between the Intersection over Union (IoU), the angular errors and the aspect ratios. Experiments show that our method can achieve 2.5% mAP improvement averagely versus the compared SOTA methods, and achieve 0.5% mAP improvement versus the next best method Oriented R-CNN with fewer regression parameters and a simpler regression approach.
Yajun Qiao, Lingjuan Miao, Zhiqiang Zhou 0001, Qi Ming
IEEE Trans. Geosci. Remote. Sens.3
2022 Optimization for Arbitrary-Oriented Object Detection via Representation Invariance Loss
abstract
Arbitrary-oriented objects exist widely in remote sensing images. The mainstream rotation detectors use oriented bounding boxes (OBBs) or quadrilateral bounding boxes (QBBs) to represent the rotating objects. However, these methods suffer from the representation ambiguity for oriented object definition, which leads to suboptimal regression optimization and the inconsistency between the loss metric and the localization accuracy of the predictions. In this letter, we propose a representation invariance loss (RIL) to optimize the bounding box regression for the rotating objects in the remote sensing images. RIL treats multiple representations of an oriented object as multiple equivalent local minima and hence transforms bounding box regression into an adaptive matching process with these local minima. Next, the Hungarian matching algorithm is adopted to obtain the optimal regression strategy. Besides, we propose a normalized rotation loss to alleviate the weak correlation between different variables and their unbalanced loss contribution in OBB representation. Extensive experiments on remote sensing datasets show that our method achieves consistent and substantial improvement. The code and models are available athttps://github.com/ming71/RIDetto facilitate future research.
Qi Ming, Lingjuan Miao, Zhiqiang Zhou 0001, Xue Yang 0005, Yunpeng Dong
IEEE Geosci. Remote. Sens. Lett.3
2022 CFC-Net: A Critical Feature Capturing Network for Arbitrary-Oriented Object Detection in Remote-Sensing Images
abstract
Object detection in optical remote-sensing images is an important and challenging task. In recent years, the methods based on convolutional neural networks (CNNs) have made good progress. However, due to the large variation in object scale, aspect ratio, as well as the arbitrary orientation, the detection performance is difficult to be further improved. In this article, we discuss the role of discriminative features in object detection, and then propose a critical feature capturing network (CFC-Net) to improve detection accuracy from three aspects: building powerful feature representation, refining preset anchors, and optimizing label assignment. Specifically, we first decouple the classification and regression features, and then construct robust critical features adapted to the respective tasks of classification and regression through the polarization attention module (PAM). With the extracted discriminative regression features, the rotation anchor refinement module (R-ARM) performs localization refinement on preset horizontal anchors to obtain superior rotation anchors. Next, the dynamic anchor learning (DAL) strategy is given to adaptively select high-quality anchors based on their ability to capture critical features. The proposed framework creates more powerful semantic representations for objects in remote-sensing images and achieves high-performance real-time object detection. Experimental results on three remote-sensing datasets including HRSC2016, DOTA, and UCAS-AOD show that our method achieves superior detection performance compared with many state-of-the-art approaches. Code and models are available athttps://github.com/ming71/CFC-Net.
Qi Ming, Lingjuan Miao, Zhiqiang Zhou 0001, Yunpeng Dong
IEEE Trans. Geosci. Remote. Sens.3
2021 Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
abstract
Arbitrary-oriented objects widely appear in natural scenes, aerial photographs, remote sensing images, etc., and thus arbitrary-oriented object detection has received considerable attention. Many current rotation detectors use plenty of anchors with different orientations to achieve spatial alignment with ground truth boxes. Intersection-over-Union (IoU) is then applied to sample the positive and negative candidates for training. However, we observe that the selected positive anchors cannot always ensure accurate detections after regression, while some negative samples can achieve accurate localization. It indicates that the quality assessment of anchors through IoU is not appropriate, and this further leads to inconsistency between classification confidence and localization accuracy. In this paper, we propose a dynamic anchor learning (DAL) method, which utilizes the newly defined matching degree to comprehensively evaluate the localization potential of the anchors and carries out a more efficient label assignment process. In this way, the detector can dynamically select high-quality anchors to achieve accurate object detection, and the divergence between classification and regression will be alleviated. With the newly introduced DAL, we can achieve superior detection performance for arbitrary-oriented objects with only a few horizontal preset anchors. Experimental results on three remote sensing datasets HRSC2016, DOTA, UCAS-AOD as well as a scene text dataset ICDAR 2015 show that our method achieves substantial improvement compared with the baseline model. Besides, our approach is also universal for object detection using horizontal bound box. The code and models are available at https://github.com/ming71/DAL.
Qi Ming, Zhiqiang Zhou 0001, Lingjuan Miao, Linhao Li
AAAI2
2021 Wide-Context Attention Network for Remote Sensing Image Retrieval
abstract
Remote sensing image retrieval (RSIR) has broad application prospects, but related challenges still exist. One of the most important challenges is how to obtain discriminative features. In recent years, although the powerful feature learning ability of convolutional neural networks (CNNs) has significantly improved RSIR, their performance can be restricted by the complexity of remote sensing (RS) images, such as small objects, varying scales, and wide scope. To address these problems, we propose a novel wide-context attention network (W-CAN). It leverages two attention modules to adaptively learn local features correlated in the spatial and channel dimensions, respectively, which can obtain discriminative features with extensive context information. During training, a hybrid loss is introduced to enhance the intraclass compactness and interclass separability of the features. Moreover, we add a branch to learn binary descriptors and realize the end-to-end descriptor aggregation. Experiments on four RS benchmark data sets demonstrate that the proposed method can outperform some state-of-the-art RSIR methods.
Honghu Wang, Zhiqiang Zhou 0001, Hua Zong, Lingjuan Miao
IEEE Geosci. Remote. Sens. Lett.2
2021 A Novel CNN-Based Method for Accurate Ship Detection in HR Optical Remote Sensing Images via Rotated Bounding Box
abstract
Currently, reliable and accurate ship detection in optical remote sensing images is still challenging. Even the state-of-the-art convolutional neural network (CNN)-based methods cannot obtain very satisfactory results. To more accurately locate the ships in diverse orientations, some recent methods conduct the detection via the rotated bounding box. However, it further increases the difficulty of detection because an additional variable of ship orientation must be accurately predicted in the algorithm. In this article, a novel CNN-based ship-detection method is proposed by overcoming some common deficiencies of current CNN-based methods in ship detection. Specifically, to generate rotated region proposals, current methods have to predefine multioriented anchors and predict all unknown variables together in one regression process, limiting the quality of overall prediction. By contrast, we are able to predict the orientation and other variables independently, and yet more effectively, with a novel dual-branch regression network, based on the observation that the ship targets are nearly rotation-invariant in remote sensing images. Next, a shape-adaptive pooling method is proposed to overcome the limitation of a typical regular region of interest (ROI) pooling in extracting the features of the ships with various aspect ratios. Furthermore, we propose to incorporate multilevel features via the spatially variant adaptive pooling. This novel approach, called multilevel adaptive pooling, leads to a compact feature representation more qualified for the simultaneous ship classification and localization. Finally, a detailed ablation study performed on the proposed approaches is provided, along with some useful insights. Experimental results demonstrate the great superiority of the proposed method in ship detection.
Linhao Li, Zhiqiang Zhou 0001, Bo Wang 0013, Lingjuan Miao, Hua Zong
IEEE Trans. Geosci. Remote. Sens.2
2019 Multi-focus image fusion using boosted random walks-based algorithm with two-scale focus maps
Jinlei Ma, Zhiqiang Zhou 0001, Bo Wang 0013, Lingjuan Miao, Hua Zong
Neurocomputing2
2018 Scale-Aware Edge-Preserving Image Filtering via Iterative Global Optimization
abstract
Presently, few filters are able to smooth images in a scale-aware manner like Gaussian filtering while not blurring the edges of large-scale features, whereas this kind of filter can be important in many visual applications requiring scale-aware manipulation while avoiding halos. In this paper, we propose a filtering technique through iterative global optimization (IGO), enabling to achieve both good scale-aware and edge-preserving performance. Our method is based on a filtering idea of selective gradient suppression and guidance gradient correction in the framework of IGO, which has the advantages of avoiding halos and preventing oversharpening of edges, and a scale-aware measure can be introduced to further control the way of gradient suppression. The proposed measure is spatially varying and oriented by coarse-scale local extrema at each pixel to better preserve the natural boundaries of large-scale structures. Besides, we show that our method can be fast implemented with a sequence of 1-D filtering. In the experiments, we demonstrate the effectiveness of our method by comparing it with current state-of-the-art filtering methods and using it in a variety of applications.
Zhiqiang Zhou 0001, Bo Wang 0013, Jinlei Ma
IEEE Trans. Multim.1
2016 A Novel Inshore Ship Detection via Ship Head Classification and Body Boundary Determination
abstract
In this letter, we propose a novel method for inshore ship detection via ship head classification and body boundary determination. Compared with some traditional ship head detection methods depending on accurate ship head segmentation, we generate novel ship head features in the transformed domain of polar coordinate, where the ship heads have an approximate trapezoid shape and can be more easily detected. Then, these features are used in the classification based on support vector machine to detect the ship head candidates, and give the important information of initial ship head direction. Next, the surrounding consistent line segments are utilized to refine the ship direction, and the ship boundary is determined based on the saliency of directional gradient information symmetrical about the ship body. Finally, the context information of sea areas is introduced to remove false alarms. Experimental results show that the proposed method can accurately and robustly detect the inshore ships in high-resolution optical remote sensing images.
Sun Li, Zhiqiang Zhou 0001, Bo Wang 0013, Fei Wu 0021
IEEE Geosci. Remote. Sens. Lett.2
2012 An Optimized Approach for Pansharpening Very High Resolution Multispectral Images
abstract
State-of-the-art pansharpening methods generally inject the spatial details extracted from the panchromatic (Pan) image into the multispectral (MS) images by considering different injection models. The fusion performances severely rely on the accuracy of the modeling and the estimation of model parameters. In this letter, we propose an optimized approach to avoid explicitly modeling the detail injection process. The solution employs the gradient field of the Pan image for spatial enhancement. The low-pass (LP) version of the fused bands are constrained to be the most similar to the original MS bands to preserve the spectral characteristics. We use the local correlation coefficients between the MS band and the LP version of the Pan image to adjust the two sources of information based on a simple observation, and it is further optimized by considering the overall quality index Q4. Experimental results demonstrate that the proposed method outperforms the state-of-the-art multiresolution analysis-based methods.
Zhiqiang Zhou 0001, Silong Peng, Bo Wang 0013, Zhihui Hao, Shaolin Chen
IEEE Geosci. Remote. Sens. Lett.1