Guanqun Wang

dblp:225/1415 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A defect severity assessment method based on empirical feature attribute scoring from semantic segmentation data for reliability analysis of steel structure connections
Zhigang Lü, Chuchao He 0001, Leiguang Duan, Guanqun Wang
Adv. Eng. Informatics7
2025 Dynamic volumetric cloud modeling with edge details refinement
Guanqun Wang, Shaoze Su, Xinjie Wang 0003
Comput. Graph.1
2025 When Automation Fails: Examining the Effect of a Verbal Recovery Strategy on User Experience in Automated Driving
abstract
Automated agents’ errors will cause various negative influences on humans and their relationships with humans (e.g., reducing user experience). They are increasingly required to have social recovery strategies (e.g., human-like apology and explanation) to mitigate the negative impacts of their errors and maintain resilient human–automation relationships. However, the efficacy of these strategies in human–automation interaction (HAI) largely remains unknown, especially in less controlled environments. Here we conducted a test track experiment and designed a verbal recovery design (consisting of an apology, explanation, and promise) by an automated driving system (ADS) installed in a real automated vehicle after an ADS failure. We utilized a Wizard of Oz design to simulate the ADS’ failure and its verbal recovery attempt (through a voice by a human or Apple Siri). Participants (N = 389) were assigned to four groups: normal (without experiencing the ADS failure), fault (experiencing the ADS failure), Siri-voice-recovery, and human-voice-recovery. The major measures were positive experience and negative experience while riding in the automated vehicle and perceived ADS usability. Overall, we found that the human-voice-recovery can to some degree mitigate the negative impacts of the ADS failure on user experience. The Siri-voice-recovery worked on positive experience but cannot restore it to that in the normal group. It implies that more empirical efforts are needed to examine social recovery strategies in HAI in natural environments, develop strategies specific to HAI, and offer effective guidelines for social recovery design.
Zhigang Xu 0001, Guanqun Wang, Siming Zhai, Peng Liu 0030
Int. J. Hum. Comput. Interact.2
2025 LLaMA-Unidetector: An LLaMA-Based Universal Framework for Open-Vocabulary Object Detection in Remote Sensing Imagery
abstract
Object detection is a crucial task in computer vision for remote sensing applications. However, the reliance of traditional methods on predefined and trained object categories limits their applicability in open-world scenarios. A key challenge in open-vocabulary object detection lies in accurately identifying unseen objects. Existing approaches often focus solely on detecting object locations, struggling to recognize the categories of previously unseen targets. To address this issue, we propose a novel benchmark where models are trained on known base classes and evaluated on their performance in detecting and recognizing unseen or novel classes. To this end, we introduce llama-Unidetector, a universal framework that incorporates textual information into a closed-set detector, enabling the generalization to open-set scenarios. Our llama-Unidetector leverages a decoupled learning strategy that separates localization and recognition. In the first stage, a class-agnostic detector identifies objects, distinguishing only between foreground and background. In the second stage, the detected foreground objects are passed through TerraOV-LLM, a multimodal large language model, for recognition, utilizing the strong generalization capabilities of large language models to infer the correct categories. We propose a self-built Vision Question Answering (VQA) remote sensing dataset, TerraVQA, and conduct extensive experiments on the NWPU-VHR10, DOTA1.0, and DIOR datasets. The llama-Unidetector achieves impressive results, with a performance of 75.46% AP, 50.22% AP and 51.38% AP on the zero-shot detection benchmarks for the NWPU-VHR10, DOTA1.0 and DIOR datasets, respectively. Our source code is available at: https://github.com/ChloeeGrace/LLaMA-Unidetector.
Jianlin Xie, Guanqun Wang, Tong Zhang 0028, Yikang Sun, He Chen 0004, Yin Zhuang, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.2
2025 A Unified Remote Sensing Object Detector Based on Fourier Contour Parametric Learning
abstract
A unified object detector needs to integrate various abilities for adapting to different remote sensing object detection tasks. However, there is a lack of a feasible way to integrate multigrained object detection requirements i.e., horizontal bounding box (HBB), oriented bounding box (OBB), and instance segmentation (InSeg) into a unified detection way. Then, it often has to design specific parametric learning ways and their corresponding architectures, which cannot be finely adaptive to various kinds of object detection tasks. Therefore, in this article, a new benchmark is set up to integrate multigrained object detection requirements of HBB, OBB, and InSeg into one challenging task of arbitrary-shaped object contour detection. At the same time, a unified object contour detector (UniconDet) is proposed for achieving multigrained object detection from complicated remote sensing scenes. First, a Fourier contour parametric modeling (FCPM) is defined to project arbitrary-shaped object contours from the spatial domain into the frequency domain. Then, it can unify spatial parametric representations of HBB, OBB, and InSeg as frequency coefficient representations, which can be used for realizing a more generic and robust parametric regression. Second, a multiview cross-attention (MVCA) feature extraction way is designed at each scale of the regression layer, which can assist UniconDet in perceiving Fourier contour parameters by exploring the coupled relations between different discrete contour sampling periods of each object. Third, a center-contour enhancing regression layer (C2-ERL) is designed to generate regional guidance and cascade contour propagation, which can ensure a more accurate center point prediction and Fourier contour parameter regression. Finally, extensive experiments are carried out on benchmarks of HBB, OBB, InSeg, and new multigrained object detection, and the results indicate that our proposed UniconDet can obtain superior performance. The source code is available athttps://github.com/ZhAnGToNG1/UniconDet.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.3
2025 Controllable Generative Knowledge-Driven Few-Shot Object Detection From Optical Remote Sensing Imagery
abstract
Few-shot object detection (FSOD) has to learn classification and localization information for unseen object detection under very low-data resource regimes. However, when deficient samples are adopted for model training, it is hard to build powerful location-aware and identification abilities for well coping with agnostic bias from diverse testing scenarios; at the same time, the overfitting phenomenon is easily occurring. Therefore, in this article, a controllable generative knowledge-driven FSOD called CGK-FSOD is proposed for unseen object detection from optical remote sensing imagery. Specifically, to enrich the learnable data space of scarce samples for preventing incomplete agnostic-bias learning, while avoiding the overfitting phenomenon, a visual-textual prompt-based controllable data generation is designed to generate high-quality object detection data based on pretrained foundational models [i.e., the stable diffusion (SD) and contrastive language-image pre-training (CLIP)], which not only can introduce the generalized domain-level knowledge into the remote sensing domain but also sets up an all-round data space to support complete learning of potential agnostic bias. Furthermore, with respect to the denoising generative process of SD, a series of cross-modality generative features in latent representation space are reused for few-shot fine-tuning by the designed cross-modality feature embedding (CMFE), which not only can bring diverse generative abilities into the feature fusion step of the detector but also gracefully sets up feature representation scalability to make the detector better adapt to agnostic bias from diverse testing scenarios of FSOD. Finally, extensive experiments are executed on two public remote sensing datasets (e.g., DIOR and NWPUVHR-10), and the results indicate that the proposed CGK-FSOD is very effective and flexible for FSOD.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.3
2024 Cloud-Device Collaborative Learning for Multimodal Large Language Models
abstract
The burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene understanding. However, the deployment of these large-scale MLLMs on client devices is hindered by their extensive model parameters, leading to a notable de-cline in generalization capabilities when these models are compressed for device deployment. Addressing this chal-lenge, we introduce a Cloud-Device Collaborative Contin-ual Adaptation framework, designed to enhance the performance of compressed, device-deployed MLLMs by lever-aging the robust capabilities of cloud-based, larger-scale MLLMs. Our framework is structured into three key components: a device-to-cloud uplink for efficient data transmission, cloud-based knowledge adaptation, and an optimized cloud-to-device downlink for model deployment. In the up-link phase, we employ an Uncertainty-guided Token Sam-pling (UTS) strategy to effectively filter out-of-distribution tokens, thereby reducing transmission costs and improving training efficiency. On the cloud side, we propose Adapter-based Knowledge Distillation (AKD) method to transfer refined knowledge from large-scale to compressed, pocket-size MLLMs. Furthermore, we propose a Dynamic Weight update Compression (DWC) strategy for the down-link, which adaptively selects and quantizes updated weight parameters, enhancing transmission efficiency and reducing the representational disparity between cloud and de-vice models. Extensive experiments on several multimodal benchmarks demonstrate the superiority of our proposed framework over prior Knowledge Distillation and device-cloud collaboration methods. Notably, we also validate the feasibility of our approach to real-world experiments.
Guanqun Wang, Jiaming Liu 0003, Chenxuan Li 0003, Yuan Zhang 0020, Junpeng Ma, Maurice Chong, Renrui Zhang, Yijiang Liu, Shanghang Zhang
CVPR1
2024 Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer
abstract
Neuromorphic spike data, an upcoming modality with high temporal resolution, has shown promising potential in autonomous driving by mitigating the challenges posed by high-velocity motion blur. However, training the spike depth estimation network holds significant challenges in two aspects: sparse spatial information for pixel-wise tasks and difficulties in achieving paired depth labels for temporally intensive spike streams. Therefore, we introduce open-source RGB data to support spike depth estimation, leveraging its annotations and spatial information. The inherent differences in modalities and data distribution make it challenging to directly apply transfer learning from open-source RGB to target spike data. To this end, we propose a cross-modality cross-domain (BiCross) framework to realize unsupervised spike depth estimation by introducing simulated mediate source spike data. Specifically, we design a Coarse-to-Fine Knowledge Distillation (CFKD) approach to facilitate comprehensive cross-modality knowledge transfer while preserving the unique strengths of both modalities, utilizing a spike-oriented uncertainty scheme. Then, we propose a Self-Correcting Teacher-Student (SCTS) mechanism to screen out reliable pixel-wise pseudo labels and ease the domain shift of the student model, which avoids error accumulation in target spike data. To verify the effectiveness of BiCross, we conduct extensive experiments on four scenarios, including Synthetic to Real, Extreme Weather, Scene Changing, and Real Spike. Our method achieves state-of-the-art (SOTA) performances, compared with RGB-oriented unsupervised depth estimation methods. Code and dataset: https://github.com/Theia-4869/BiCross.
Jiaming Liu 0003, Qizhe Zhang, Xiaoqi Li 0020, Jianing Li 0001, Guanqun Wang, Ming Lu 0002, Tiejun Huang 0001, Shanghang Zhang
ICRA5
2024 Regression-Guided Positive Sample Refocusing Paradigm for Tiny Object Detection in Aerial Images
abstract
Tiny object detection represents a pivotal challenge in remote sensing intelligent interpretation, necessitating detectors to exhibit heightened precision in object localization. However, typical model optimization strategies cannot release the detector’s potential for precisely localizing objects. And the lack of interpretability in detection box filtering based on object classification scores serves as a constraint on further performance improvement. Therefore, this paper proposed a novel model optimization strategy to thoroughly unleash the potential of the detector for precise localization. Then, the utilization of object comprehensive confidence score enhances the interpretability of the post-processing step for detection boxes. Rigorous experiments on the AI-TOD dataset have demonstrated the effectiveness of our method, achieving state-of-the-art performance.
Lihui Ge, He Chen 0004, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, Fukun Bi, Liang Chen 0004
IGARSS3
2024 Advancing Controllable Diffusion Model for Few-Shot Object Detection in Optical Remote Sensing Imagery
abstract
Few-shot object detection (FSOD) from optical remote sensing imagery has to detect rare objects given only a few annotated bounding boxes. The limited training data is hard to represent the data distribution of realistic remote sensing scenes, restricting the performance of FSOD. Recently, learning conditional controls for text-to-image diffusion model has achieved great progress, which is capable of precisely generating the controllable yet imaginational images by text prompt and spatially localized input conditions. Accordingly, in this work, we aim to explore the potential of diffusion model and propose a solution for few-shot object detection by controllable data generation. Firstly, draw upon a few annotated objects, their bounding boxes and categories are respectively used as the spatial conditions and text prompts, then employ them into large text-to-image diffusion models for controlled image generation. Secondly, based the generated images, in order to adapt to the scale and orientation variances of remote sensing objects, a data transformation is devised for boosting the robustness of model training. Finally, some experiments were conducted on public remote sensing dataset DIOR, and the results proved its effectiveness.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, Fukun Bi
IGARSS4
2024 Regression-Guided Refocusing Learning With Feature Alignment for Remote Sensing Tiny Object Detection
abstract
Tiny object detection is a formidable challenge in remote sensing intelligent interpretation. Tiny objects are usually fuzzy, densely distributed and highly sensitive to positioning errors, which leads to the mainstream detector usually achieving suboptimal detection performance when facing tiny objects. To address the mismatch of mainstream detector architectures and model optimization strategies in the context of tiny object detection, this paper presents an efficient and interpretable algorithm for tiny object detection, termed the Cross-Attention based Feature Fusion Enhanced tiny object detection Network (CAF2ENet). First, the cross-attention mechanism is introduced to refine the upsampling results of deep features. This refinement improves the precision of multi-scale feature fusion. Second, a training strategy named regression-based refocusing learning is introduced. Deviating from the conventional optimization strategy, our method guides the optimizer to prioritize higher-quality detection boxes by adjusting sample weights. This adjustment significantly amplifies the detector’s potential to achieve superior detection results. Finally, the object composite confidence score is employed for the interpretable filtering of detection boxes. Extensive experiments on Tiny Object Detection in Aerial Images (AI-TOD) and object Detection in Optical Remote sensing images (DIOR) datasets are carried out, and comparison indicate that the proposed CAF2ENet can perform the remarkable performance compared to other state-of-the-art (SOTA) tiny object detection detectors, as it can reach 63.7% Average Precision (AP50) on AI-TOD and 75.4%AP50on DIOR, achieve SOTA performance.
Lihui Ge, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Hao Dong 0003, Liang Chen 0004
IEEE Trans. Geosci. Remote. Sens.2
2024 DECOR: Dynamic Decoupling and Multiobjective Optimization for Long-Tailed Remote Sensing Image Classification
abstract
In the realm of remote sensing, targets of interest span a range of categories. However, their distribution is not always uniform. Certain categories substantially outnumber others, resulting in what’s termed a ‘long-tailed distribution’ in remote sensing imagery. This imbalanced distribution often biases a classifier’s focus toward the more abundant (head) classes, at the detriment of the less-represented (tail) classes. Such biases undermine the classifier’s generalization performance, particularly in the context of remote sensing image classification (RSIC). While existing mitigation approaches such as resampling, reweighting, and transfer learning offer some respite, they often miss out on in-depth knowledge refinement, rendering them less effective for severe long-tailed RSIC scenarios. To counter these challenges, we introduce DECOR, a dynamic decoupling and multi-objective optimization framework. Within DECOR, the feature extractor and classifier are dynamically decoupled, promoting superior feature representation and classifier training. Then, a multi-objective optimization approach is proposed to delve deeper, refining feature representation at the knowledge level using learnable feature centroids coupled with masked world knowledge learning. Moreover, to combat the pronounced effects of sample imbalance on classifier training, we employ a class-balanced re-sampling technique paired with a parameter-efficient adapter, which sharpens the classifier’s decision boundary and bridges the gap between representation and classification. DECOR’s efficacy is validated through comprehensive experiments on several datasets, including the NWPU-RESISC45-LT (NWPU-LT), AID-LT, and our self-built BIT-AFGR50-LT. Experimental results demonstrate DECOR’s marked enhancement in performance on long-tailed datasets. Our source code is available at: https://github.com/ChloeeGrace/DECOR.
Jianlin Xie, Guanqun Wang, Yin Zhuang, Can Li 0005, Tong Zhang 0028, He Chen 0004, Liang Chen 0004, Shanghang Zhang
IEEE Trans. Geosci. Remote. Sens.2
2024 To Err is Automation: Can Trust be Repaired by the Automated Driving System After its Failure?
abstract
Failures of the automated driving system (ADS) in automated vehicles (AVs) can damage driver–ADS cooperation (e.g., causing trust damage) and traffic safety. Researchers suggest infusing a human-like ability, active trust repair, into automated systems, to mitigate broken trust and other negative impacts resulting from their failures. Trust repair is regarded as a key ergonomic design in automated systems. Trust repair strategies (e.g., apology) are examined and supported by some evidence in controlled environments, however, rarely subjected to empirical evaluations in more naturalistic environments. To fill this gap, we conducted a test track study, invited participants (N= 257) to experience an ADS failure, and tested the influence of the ADS’ trust repair on trust and other psychological responses. Half of participants (n= 128) received the ADS’ verbal message (consisting of apology, explanation, and promise) by a human voice (n= 63) or by Apple's Siri (n= 65) after its failure. We measured seven psychological responses to AVs and ADS [e.g., trust and behavioral intention (BI)]. We found that both strategies cannot repair damaged trust. The human-voice-repair strategy can to some degree mitigate other detrimental influences (e.g., reductions in BI) resulting from the ADS failure, but this effect is only notable among participants without substantial driving experience. It points to the importance of conducting ecologically valid and field studies for validating human-like trust repair strategies in human–automation interaction and of developing trust repair strategies specific to safety-critical situations.
Peng Liu 0030, Yueying Chu, Guanqun Wang, Zhigang Xu 0001
IEEE Trans. Hum. Mach. Syst.3
2023 Multi-Grained Global-Local Semantic Feature Fusion for Few Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification aims to classify unseen scenes by using only a few labeled samples. Hence, how to set up a more effective feature description according to a few labeled samples, becomes an important issue. In this paper, in view of more complicated remote sensing scenes containing several hierarchical and coupled spatial relations (e.g., internal and external spatial contexts), which severely hinder the feature extraction under few-shot learning scenarios, a multi-grained global-local semantic feature fusion (MGGL-SFF) method is proposed for few-shot remote sensing scene classification, which can better combine the global discriminative spatial semantic features with local transferable fragment features to set a powerful prototype representation up for few shot learning. Finally, experiments are carried out on defined few-shot remote sensing scene classification benchmark, and results proved the proposed MGGL-SFF can achieve a new state-of-the-art performance.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004
IGARSS4
2023 Posterior Instance Injection Detector for Arbitrary-Oriented Object Detection From Optical Remote-Sensing Imagery
abstract
Arbitrary-oriented object detection (AOOD) from optical remote sensing imagery has to correctly generate delicate oriented boundary boxes (OBBs) and meanwhile identify their specific categories. However, how to make detectors learn delicate parameters of OBBs, especially for the crucial orientation information, and identify object category from complex background becomes a challenge task. Therefore, in this article, for exploring a better way to guide the detector to learn specific category and parametric information of OBBs, a novel one-stage anchor-free detector called Posterior Instance Injection Detector (PIIDet) is proposed for AOOD. First, as the anchor-free manner lacks prior information, an object-aware posterior guidance (OAPG) structure is proposed to generate specific-category instances used for conditioning on OBB prediction. This structure can assist the proposed PIIDet in better learning the relative parametric information of OBBs corresponding to their specific categories. Besides, to guarantee a high quality injection of specific-category instances, a new hierarchical feature fusion module is developed to establish a suitable multi-scale feature mapping space. Second, considering the negative optimization of angle regression, which is caused by the boundary discontinuity of angular periods and sudden shifts of the relation between width and height in training phase, a novel binary classification embedded angle regression space (BCE-RegSpace) is devised for providing continuous angle regression space and stable relation between width and height. Finally, extensive experiments are executed on three AOOD benchmarks (e.g., DOTA, DIOR-R and HRSC2016), and results proved that the proposed concise one-stage anchor-free PIIDet can reach the state-of-the-art (SOTA) performance and meanwhile have an impressive inference speed.
Tong Zhang 0028, Yin Zhuang, He Chen 0004, Guanqun Wang, Lihui Ge, Liang Chen 0004, Hao Dong 0003, LianLin Li
IEEE Trans. Geosci. Remote. Sens.4
2022 Adaptive Local Context Embedding for Small Vehicle Detection from Aerial Optical Remote Sensing Images
abstract
Small vehicle detection is one of the remaining challenging task because the ambiguous appearance is against complex background interference. Consequently, in order to improve the performance of small vehicle detection from aerial optical remote sensing images, a novel adaptive local context (ALC) embedding way is designed and further introduced into an anchor free detection manner which is called ALC-Net, and in ALC-Net, it can adaptively set up the effective local context feature to improve keypoint description of small vehicles and boost the detection performance without adding extra prior information. Finally, several experiments are carried out on two widely used datasets (e.g., UCAS-AOD [1] and VEDAI [2]) and the results indicate that the proposed ALC-Net can exhibit the competitive small vehicle detection performance than other detectors.
Shanjunyu Liu, Yin Zhuang, Hao Dong 0003, Peng Gao 0007, Guanqun Wang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li
IGARSS5
2022 FSoD-Net: Full-Scale Object Detection From Optical Remote Sensing Imagery
abstract
Object detection is an essential task in computer vision. Recently, several convolution neural network (CNN)-based detectors have achieved a great success in natural scenes. However, for optical remote sensing images with a large scale of view, lower proportion of foreground target pixels and drastic differences in object scale present considerable challenges. To address these problems, we propose a novel one-stage detector called the full-scale object detection network (FSoD-Net) which consists of proposed multiscale enhancement network (MSE-Net) backbone cascaded with scale-invariant regression layers (SIRLs). First, MSE-Net provides the multiscale description enhancement by integrated the Laplace kernel with fewer parallel multiscale convolution layers. Second, SIRLs contain three different isolated regression branch layers (i.e., corresponding to small, medium, and large scales), which make default discrete scale bounding boxes (bboxes) cover full-scale object information in regression procedure. A novel specific scale joint loss is also designed that uses the softmax function combined with a strong$L_{1}$-norm constraint in each regression branch layer. It can further speed up the convergence and improve the classification scores of predicted bboxes. Finally, extensive experiments are carried on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) which contain multiple instances from different imaging platforms, and these results demonstrate that FSoD-Net can achieve better performance than other state-of-the-art one-stage detectors, and it can reach a mean average precision (mAP) of 75.33% on DOTA and 71.80% mAP on DIOR, respectively. Especially, the average precision (AP) of tiny object detection can improve 10%–20% approximately.
Guanqun Wang, Yin Zhuang, He Chen 0004, Tong Zhang 0028, LianLin Li, Shan Dong, Qianbo Sang
IEEE Trans. Geosci. Remote. Sens.1
2022 Multiscale Semantic Fusion-Guided Fractal Convolutional Object Detection Network for Optical Remote Sensing Imagery
abstract
Optical remote sensing object detection is a challenging task, because of the complex background interference, ambiguous appearances of tiny objects, densely arranged circumstances, and multiclass object with vaster scale variances and irregular aspect ratios. The performance of object detection is seriously restricted. Thus, in this article, inspired by the anchor-free object detection framework, and aiming to solve these difficulties to improve the optical remote sensing object detection performance, a powerful one-stage detector of multiscale semantic fusion-guided fractal convolution network (MSFC-Net) is proposed. First, facing these strong-coupled semantic relations in each complex scene, a compound semantic feature fusion (CSFF) way is designed for generating an effective semantic description, which is a benefit to pixel-wise object center point interpretation. In addition, it can be easily extended into a semantic segmentation task. Second, in view of accurate multiclass pixel-wise center point predictions based on an effective compound semantic description, a novel fractal convolution (FC) regression layer is designed, which adaptively achieves the regression of multiscale bounding boxes (bboxes) with irregular aspect ratio under no priori information. Third, related to the set up FC regression layer, a specific hybrid loss is designed to make the proposed MSFC-Net converge better. Finally, the extensive experiments on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) datasets are carried out, and comparisons indicate that the proposed MSFC-Net can perform the remarkable performance than other state-of-the-art one-stage detectors, as it can reach 80.26% mean average precision (mAP) and 79.33% mF1 on DOTA and 70.08% mAP and 73.45% mF1 on DIOR. Then, our work is available athttps://github.com/ZhAnGToNG1/MSFC-Net.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, Shan Dong, He Chen 0004, LianLin Li
IEEE Trans. Geosci. Remote. Sens.3
2021 Trajectory Optimization for a Connected Automated Traffic Stream: Comparison Between an Exact Model and Fast Heuristics
abstract
Numerous fast heuristic algorithms, including shooting heuristics (SH), have been developed for real-time trajectory optimization, although their optimality has not yet been quantified. This paper compares the performance between fast heuristics and exact optimization models. We investigate a core trajectory optimization problem as a building block for numerous trajectory optimization problems, i.e., guiding movements of connected automated vehicles on a one-lane highway when the arrival and departure times and velocity are given. To apply the SH algorithm to this problem, we adapt it to a fast-simplified shooting heuristic (FSSH) model to solve the trajectory smoothing problems with different arrival and departure velocities. An exact trajectory optimization (ETO) model is formulated that takes the vehicle position and velocity as the decision variables, and the fuel consumption and driving comfort as the objective function. The constraints of the model are based on the limits and safety of the vehicle dynamics between consecutive vehicles. We demonstrate the convexity of the ETO objective function, ensuring the solvability of the ETO model at the true optimum using gradient descent algorithms supplied by the MATLAB optimization toolbox. Six groups of numerical experiments using different input parameters and one experiment using real Next Generation Simulation (NGSIM) data are conducted. ETO can improve the objective values by a few to tens of percentage points. However, FSSH achieves a greater solution efficiency with an average solution time of less than 0.1 s compared to ~450 s for ETO.
Zhigang Xu 0001, Yu Wang 0084, Guanqun Wang, Xiaopeng Shaw Li, Robert L. Bertini, Xiaobo Qu 0002, Xiangmo Zhao
IEEE Trans. Intell. Transp. Syst.3
2020 Feature Enhanced Centernet for Object Detection in Remote Sensing Images
abstract
Multi-scale object detection in optical remote sensing imagery is a challenging task due to the varied object scales. Existed state-of-art object detection methods have achieved significant growth. However, most of the methods are based on default anchors, which need to be predefined. The multi-scale object detection accuracy still needs to be improved, especially for small and dense objects. To improve the robustness of the detection algorithm and the performance of multi-scale object detection, a novel anchor-free multi-scale object detection method Feature Enhanced CenterNet is proposed in this paper. First, we use the “encoder-decoder” structure and introduce horizontal connections to enhance feature representation capabilities. Second, an context-aware up-sampling method is proposed to obtain feature maps with suitable scale. To demonstrate the performance of the proposed method, we perform abundant experiments on the public remote sensing datasets. The experimental results demonstrate the robustness and effectiveness of the proposed method.
Tong Zhang 0028, Guanqun Wang, Yin Zhuang, He Chen 0004, Hao Shi 0006, Liang Chen 0004
IGARSS2
2020 FRF-Net: Land Cover Classification From Large-Scale VHR Optical Remote Sensing Images
abstract
Deep learning (DL) technique is widely applied in remote sensing (RS) applications because of its outstanding nonlinear feature extraction ability. However, with regard to the issues of large-scale and very high-resolution (VHR) land cover classification, multi-object distributions and clear appearance with large intraclass difference become challenges for refined pixelwise land cover mapping. Focusing on these problems, the letter proposed a novel encoding-to-decoding method called the full receptive field (RF) network (FRF-Net) based on two types of attention mechanism. In the FRF-Net, ResNet-101 is used as the basic backbone. Then, the ensemble feature is generated by encoding the high-level features based on the self-attention mechanism which could achieve full RF to capture long-range semantic. Next, the encoding result is decoded by the fusion attention mechanism combined with the low-level feature to produce a fusion feature which contains a refined semantic description for accurate land cover mapping. Extensive experiments based on the GID and ISPRS data sets proved that the proposed network outperforms the state-of-the-art methods. The FRF-Net achieved 66.71% and 64.17% of the mean of classwise Intersection over Union (mIOU) with smaller computation cost on ISPRS and GID, respectively.
Qianbo Sang, Yin Zhuang, Shan Dong, Guanqun Wang, He Chen 0004
IEEE Geosci. Remote. Sens. Lett.4
2019 Spatial Enhanced-SSD For Multiclass Object Detection in Remote Sensing Images
abstract
Accurate multiclass object detection in remote sensing images is a challenging task, especially for small objects. Since the scales of objects in remote sensing images have a great variance, almost all of the advanced detection methods have shortcomings. Consequently, improving the accuracy of multiclass objects detection has always been the direction of researchers' efforts. In this paper, a spatial enhanced-Single Shot MultiBox Detector (SE-SSD) is proposed. First, to enhance the spatial information, we enlarge the input image channels with embedding oriented-gradients feature maps. Second, the multiple output layers in the backbone network are changed to reduce one pooling operation. Finally, we design a context module to enhance the receptive field for feature layer description in SE-SSD framework. Experimental results on DOTA dataset demonstrate that Spatial Enhanced-SSD method reaches a much higher mean average precision (mAP) than Faster R-CNN, SSD and other classic detection network.
Guanqun Wang, Yin Zhuang, Zhiru Wang, He Chen 0004, Hao Shi 0006, Liang Chen 0004
IGARSS1
2018 A Novel Harbor Detection Method Based on Pattern Coding Algorithm
abstract
Harbor automatic detection is a scene interpretation in remote sensing image processing. Fast and accurate harbor detection can significantly improve the performance of inshore ship detection. In order to achieve harbor detection in complex remote sensing images, in this paper, a novel pattern coding algorithm is proposed. The proposed harbor detection method has three steps: First, the harbor water area is extracted by using the definition circle (DC) model. Secondly, pre-generated eight KEY patterns are multiplied with the local scenes, and the probability density function (PDF) of the multiplied local scene is recorded. Finally, the Euclidean distance between the eight patterns' PDFs and the original local scene's PDF is calculated, then compared with the threshold and coded, so as to realize the harbor area detection. Experimental results demonstrate that the novel method has outstanding performance on harbor area detection in complex broad width remote sensing images.
Guanqun Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004
IGARSS1