EDBT 2026 Demo / reviewers in the wild / expert
Qiang Li 0042
dblp:72/872-42
· DBLP profile ↗
52ranked-venue papers
10as first author
52since 2021 · last 2026
0000-0002-6736-3389ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 7 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Generalizable Remote Sensing Change Detection via Low-Rank Exchange Adaptation of Vision Foundation ModelabstractRemote sensing change detection (CD) has achieved remarkable progress in recent years. However, little attention has been paid to generalizable change detection (GCD) methods that can effectively generalize to unseen scenarios or domains beyond the training distribution. The major challenges in GCD arise from domain diversity and bitemporal domain shifts in remote sensing images, caused by variations in imaging platforms, acquisition times, geographic regions, and observed events. To tackle these challenges, we propose GenCD, a GCD framework built upon vision foundation models (VFMs). Specifically, GenCD introduces two key components: (1) a Low-Rank Exchange Adaptation (LREA) strategy of VFMs that aligns bitemporal representations while preserving the generalization capacity of VFMs on single-temporal inputs; and (2) a Token-Guided Feature Refinement (TGFR) mechanism that leverages an input-independent token as a guide to refine difference features, improving the discrimination between changed and unchanged regions. We conduct extensive cross-dataset evaluations on eight diverse datasets across three binary CD tasks: land cover, land use, and building-only CD. The results consistently demonstrate the superior generalization of GenCD over SoTA methods, highlighting its effectiveness in GCD. Jingtao Hu, Qiang Li 0042, Qi Wang 0009 |
AAAI | 3 |
| 2026 | Remote sensing imagery shadow detection via physical constraint
Kaichen Chi, Qiang Li 0042, Qi Wang 0009 |
Pattern Recognit. | 4 |
| 2026 | Heterogeneous change detection via frequency domain interaction and statistical style embedding
Qiang Li 0042, Qi Wang 0009 |
Pattern Recognit. | 2 |
| 2026 | MSDP-Net: Multi-scale distribution perception network for rotating object detection in remote sensing
Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009 |
Pattern Recognit. | 4 |
| 2026 | RAPTOR: Rotational Adaptive Parallel Topology for Object Detection in remote sensing
Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009 |
Pattern Recognit. | 4 |
| 2026 | Cross-Modal Spherical Aggregation for Weakly Supervised Remote Sensing Shadow RemovalabstractShadows are dark areas, typically rendering low illumination intensity. Admittedly, the infrared image can provide robust illumination cues that the visible image lacks, but existing methods ignore the collaboration between heterogeneous modalities. To fill this gap, we propose a weakly supervised shadow removal network with a spherical feature space, dubbed S2-ShadowNet, to explore the best of both worlds for visible and infrared modalities. Specifically, we employ a modal translation (visible-to-infrared) model to learn the cross-domain mapping, thus generating realistic infrared samples. Then, Swin Transformer is utilized to extract strong representational visible/infrared features. Simultaneously, the extracted features are mapped to the smooth spherical manifold, which alleviates the domain shift through regularization. Well-designed similarity loss and orthogonality loss are embedded into the spherical space, prompting the separation of private visible/infrared features and the alignment of shared visible/infrared features through constraints on both representation content and orientation. Such a manner encourages implicit reciprocity between modalities, thus providing a novel insight into shadow removal. Notably, ground truth is not available in practice, thus S2-ShadowNet is trained by cropping shadow and shadow-free patches from the shadow image itself, avoiding stereotypical and strict pair data acquisition. More importantly, we contribute a largescale weakly supervised shadow removal benchmark that makes shadow removal independent of specific scenario constraints possible. Extensive experiments demonstrate that S2-ShadowNet outperforms state-of-the-art methods in both qualitative and quantitative comparisons. The code and benchmark are available at https://github.com/chi-kaichen/S2-ShadowNet. Kaichen Chi, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Multim. | 4 |
| 2026 | Deep Reinforcement Learning for Lunar Polar Low-Light EnhancementabstractAs a bridge between the moon and human perception, the lunar optical image reflects lunar topography, geology, and evolution. Unfortunately, the permanent shadow regions (PSRs) near the lunar poles suffer from information contamination due to insufficient illumination. Low-light enhancement is a subjective process whose target is tied to human visual perception. However, existing low-light enhancement methods often operate as opaque “black box”, lacking transparency and failing to accommodate diverse perceptual preferences. To this end, we explore a PSRs Low-Light Enhancer (PSRs-LLE) that treats low-light enhancement as a Markov decision process, thereby dynamically fitting perceptual preferences. Specifically, a deep Q network as an agent integrates multiple user-friendly attributes (e.g., brightness, contrast, chroma, and detail) through actions recursion (i.e., a candidate set of image enhancement operations). Such transparent and specific action sequences satisfy customization preferences of users while providing convincing interpretability, compared with the “black box” paradigm of deep learning. More importantly, a well-designed non-reference loss function liberates PSRs-LLE from the dilemma of virtual assumptions and paired data, which further enhances usability. Extensive experiments demonstrate that PSRs-LLE outperforms state-of-the-art methods in both qualitative and quantitative comparisons. Kaichen Chi, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Multim. | 2 |
| 2025 | Quantized Memory-Efficient Full-Parameter Tuning with Sign Descent OptimizationabstractFull Parameter Fine-Tuning (FPFT) has become the preferred method for adapting LLMs to downstream tasks due to its exceptional performance. Current methods primarily utilize zeroth-order optimizers or integrate gradient computation and updates to conserve GPU memory. However, they fail to consider the optimizer states information (e.g., momentum, variance), leading to suboptimal convergence and instability during training. To address this, we propose a Quantized Memory-Efficient Full-Parameter Tuning with Sign descent optimization training framework (SQ-MEFT). Firstly, we construct a novel optimizer that uses the sign of momentum as the update amount to maximize the potential of momentum. In addition, to better maintain memory efficiency, we apply 4-bit quantization to the momentum while synchronously computing and updating gradients. When trained with mixed precision, our optimizer can reduce the total memory footprint by up to 7× compared to AdamW. Xuezhi Zhao, Haichen Bai, Qiang Li 0042, Qi Wang 0009 |
ICME | 3 |
| 2025 | Underwater image enhancement via color constraints and transmission-guided modeling
Kaichen Chi, Qiang Li 0042 |
Pattern Recognit. | 2 |
| 2025 | RSMamba: Biologically Plausible Retinex-Based Mamba for Remote Sensing Shadow RemovalabstractShadow removal is an essential task for remote sensing imagery analysis, which is tricky due to spatial irregular and inhomogeneous degradation distribution. Unfortunately, current shadow removal pipelines face challenges with suboptimal performance and insufficient interpretability. To this end, we unleash the long-sequence modeling potential of State Space Models (SSMs) in the context of shadow removal. Coupled with the accurate perception of traditional Retinex decomposition towards illumination, the well-designed RSMamba enjoys the best of both worlds between superior competitiveness and theoretical intuitiveness. Specifically, RSMamba mimics the retina and cerebral cortex to explore illumination and reflectance. The former drives the selective scan mechanism to enhance the response towards contamination, while the latter serves as a tool to preserve illumination fidelity. In addition, contour and gradient regularizations of illumination and reflectance components reflect the spatial opponency of shadows, which are consistent with the center-surround opponent receptive field of the human visual system. Such a manner incorporates the domain knowledge of neurophysiological mechanisms into neural networks, providing new insights into shadow removal. Extensive experiments demonstrate that RSMamba outperforms state-of-the-art methods. Kaichen Chi, Sai Guo, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | DualStrip-Net: A Strip-Based Unified Framework for Weakly- and Semi-Supervised Road Segmentation From Satellite ImagesabstractAutomated road segmentation from remote sensing imagery remains a fundamental challenge in Earth observation systems. The primary bottleneck lies in acquiring dense pixel-wise annotations, which is both labor-intensive and time-prohibitive. This article presents DualStrip-Net, a novel deep learning framework for weakly supervised and semi-supervised road segmentation that effectively handles both sparse annotations and limited labeled data. Unlike conventional convolutional neural network (CNN)-based segmentation methods that lack explicit road topology modeling, DualStrip-Net exploits the inherent linear topology of road networks through a dual-stream architecture that combines patch-level annotation strategy and strip-based feature learning. The framework captures road characteristics through orthogonal strip processing in horizontal and vertical orientations. The proposed DualStrip Learning mechanism enables robust feature representation of road structures through complementary views. Extensive evaluations on the DeepGlobe, Massachusetts, and CHN6-CUG benchmark datasets demonstrate that DualStrip-Net achieves superior performance in both weakly supervised and semi-supervised settings. Notably, with only 20% of labeled training data, our method outperforms the supervised-only baselines on both Massachusetts and CHN6-CUG datasets. The code is available athttps://github.com/jasonnhu/DualStrip-Net/. Jingtao Hu, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Adaptive Frequency Separation Enhancement Network for Infrared Small Target Detection
Wuwei Wang, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Global-Local Feature Collaborative Fusion and Edge Refinement Network for Remote Sensing Change DetectionabstractRemote sensing image change detection (RSI-CD) has witnessed significant progress. The core focus lies in accurately identifying spatio-temporal changes in remote sensing images and precisely extracting edge details. However, current remote sensing image change detection primarily faces two challenges: the diversity and complexity of change regions, and the blurred edges of change regions. These issues lead to pseudo-change interference and incomplete detection of change targets. To alleviate these limitations, this study first develops a Local-Global Differential Fusion Module (LGDFM) that integrates local feature extraction with global self-attention mechanism to enhance change feature representation through effective bi-temporal image fusion. This module simultaneously captures both global and local information from source images, thereby overcoming the insufficient feature fusion in existing bi-temporal change detection methods. Subsequently, we propose an Edge-Aware Enhancement Module (EAEM) that employs morphological operations to extract edge features. The module progressively refines prediction results through multi-stage edge feature integration that ensures precise contour delineation. Experiments on two datasets demonstrate that the proposed method achieves competitive performance compared to existing approaches. Zhenjian Qu, Yuemei Qin, San Zhang, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Multibranch Mutual-Guiding Learning for Infrared Small Target DetectionabstractAt present, many infrared target detection approaches focus on designing modules that address the two key characteristics of targets: their weak signals and small size. However, these approaches often fail to fully leverage guided learning for weak and small target content, resulting in sub-optimal detection performance, particularly in terms of shape preservation and target positioning. To tackle this challenge, this paper proposes a multi-branch mutual-guiding learning network (MMLNet) that enhances the accuracy of infrared target detection, even in the absence of clear morphological and textural features in images. The method consists of three branches: edge, positioning, and detection, each of which is designed with a specialized module from a unique perspective. In the detection branch, we introduce a multi-dimensional lossless encoder optimized through a downsampling strategy and multi-level feature fusion to mitigate feature loss in small targets. In the positioning branch, a target positioning strategy is proposed to explicitly identify candidate targets from the image by means of a learnable multi-kernel pattern. In the edge branch, a simple architecture is adopted to enhance the ability of the model to preserve the target shape. To effectively utilize the knowledge of different branches, a mutual-guiding fusion module is developed to adjust information within and between branches. The manner adaptively utilizes the specific knowledge from each input branch. Experiment results demonstrate that the proposed method achieves comparable performance, and the visualization results show the advantages of our method in shape preservation and positioning of the targets. Our code is publicly available at https://github.com/qianngli/MMLNet. Qiang Li 0042, Wei Zhang 0250, Wanxuan Lu, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | InterMamba: A Visual-Prompted Interactive Framework for Dense Object Detection and AnnotationabstractExisting object detection methods is constrained by the high annotation costs, particularly in remote sensing due to the diversity of targets and the large scale of data. Visual-Prompted Interactive Object Detection can enhance the efficiency of data annotation by leveraging user-provided visual prompts to iteratively refine detection results. However, current interactive annotation frameworks are hindered by their reliance on simple feature fusion strategies, which limit their ability to capture fine-grained semantic relationships. Moreover, more advanced fusion methods face computational complexity challenges, making them unsuitable for high-resolution feature spaces commonly encountered in remote sensing imagery. To address these limitations, we propose InterMamba, an efficient framework for interactive object detection in remote sensing images. InterMamba integrates the VMamba backbone and a novel Cross Vision Selective Scan Module (Cross-VSSM) to achieve linear-complexity multi-scale feature fusion, reducing memory consumption while capturing fine-grained details in high-resolution feature spaces. To further enhance interaction flexibility and detection precision, a hybrid Gaussian heatmap generation method is proposed to encodes user-provided point and bounding box annotations. Meanwhile, a User Interaction Loss function further optimizes detection accuracy in dense scenarios by aligning localization and classification with user guidance. Our experiments demonstrate that InterMamba consistently outperforms existing methods in mean Average Precision (mAP). In terms of enhancing precision and reducing annotation costs, InterMamba establishes a robust solution for interactive remote sensing object detection. Code will be available at https://github.com/lsjhaha/InterMamba. Shanji Liu, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Semantic-Guided Framework for Few-Shot Remote Sensing Object DetectionabstractFew-Shot Object Detection (FSOD) aims to recognize novel class targets using limited annotated data. Conventional approaches rely on extensive base class training, followed by fine-tuning where few instances from both base and novel classes are sampled for each category. Although they demonstrate remarkable performance in natural image domains, the specificity of remote sensing scenarios poses two critical challenges for FSOD: 1) The morphological differences between remote sensing images and natural images are significant, leading to a loss of structural priors in the Region Proposal Network (RPN). This makes it difficult for structural priors pretrained on natural images to generalize to remote sensing images, especially for novel class with scarce data; 2) Differences in imaging conditions lead to appearance variations among similar objects, leading to sparse visual features are insufficient to represent the common semantic structure of the entire class. To solve problems above, we introduce an innovative framework named ST-FSOD. Primarily, we introduce the SA-RPN module, which leverages efficient pixel association capability to generate high-quality foreground object proposals. Subsequently, through a text guiding learner module (TGL), we use textual labels of each category to generate image-agnostic text-guided prototypes. The enhanced text prototypes are fused with visual features to complement the sparse visual features. Extensive experiments conducted on the DIOR, NWPU VHR-10 and RSOD benchmarks demonstrate that the proposed method consistently surpasses strong baselines and achieves superior performance compared to previous state-of-the-art (SOTA) approaches. Our project will be open-sourced soon on https://github.com/wdcjhyy/ST-FSOD. Chenchen Sun, Yuyu Jia, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Embedding Generalized Semantic Knowledge Into Few-Shot Remote Sensing SegmentationabstractFew-shot segmentation (FSS) for remote sensing (RS) imagery leverages supporting information from limited annotated samples to achieve query segmentation of novel classes. Previous efforts are dedicated to mining segmentation-guiding visual cues from a constrained set of support samples. However, they still struggle to address the pronounced intra-class differences in RS images, as sparse visual cues make it challenging to establish robust class-specific representations. In this article, we propose a holistic semantic embedding (HSE) approach that effectively harnesses general semantic knowledge, i.e., class description (CD) embeddings. Instead of the naive combination of CD embeddings and visual features for segmentation decoding, we investigate embedding the general semantic knowledge during the feature extraction stage. Specifically, in HSE, a spatial dense interaction (SDI) module allows the interaction of visual support features with CD embeddings along the spatial dimension via self-attention. Furthermore, a global content modulation (GCM) module efficiently augments the global information of the target category in both support and query features, thanks to the transformative fusion of visual features and CD embeddings. These two components holistically synergize CD embeddings and visual cues, constructing a robust class-specific representation. Through extensive experiments on the standard FSS benchmark, the proposed HSE approach demonstrates superior performance compared to peer work, setting a new state-of-the-art. Qi Wang 0009, Yuyu Jia, Wei Huang 0068, Junyu Gao 0001, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Exploring Context Alignment and Structure Perception for Building Change DetectionabstractAutomatically monitoring building changes can assist human experts in disaster rescue, urban planning, and resource protection. Consequently, much research focuses on building change detection. Recently, the methods based on deep learning have achieved impressive performance. However, most of them ignore the effect of bitemporal image misalignment, which is prone to lead to false detection. To this end, a building change detection model with context alignment and structure perception (CASP) is proposed. First, imitating the brain logic of humans to identify changes, a bitemporal interactive alignment module (BIAM) is designed, which suppresses the spatial dislocation noise via a bidirectional reference-guided feature aggregation strategy. Building on this, a difference-induced alignment module (DIAM) is introduced to mitigate the adverse impact of misalignment errors further and improve the accuracy of building change detection. Second, a structure-aware feature fusion module is developed and integrated into the feature encoder, to enhance the discrimination of building representations and highlight the specificity of the proposed method. Extensive experiments on three representative building change detection datasets are implemented to verify the superiority of the above improvements. The quantitative and qualitative results demonstrate that the proposed method achieves competitive performance. The code is available athttps://github.com/ptdoge/CASP. Qi Wang 0009, Jiawei Ren 0004, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Robust Image Registration via Consistent Topology Sort and Vision InspectionabstractMachine vision plays a crucial role in Earth observation. As a fundamental and challenging task in vision systems, image registration faces new challenges due to increasing collaborative and customization applications. The prevalence of more false matches and low-precision matches is particularly evident in complex and changeable scenarios. In this article, we propose a robust image registration method via topology sort and vision consistence. Initial candidate matches are established via the nearest neighbor ratio of image intensity descriptors. A topological sort across the proximity structure around the point pairs is defined to assess the reliability of candidate matched pairs, effectively eliminating more false matches while retaining highly reliable point pairs. To preserve more point pairs, we develop a spatial visual inspection mechanism to further determine the potential matches from the remaining pairs that do not satisfy the previous topological constraint. During vision inspection, the spatial transformation model is simultaneously estimated. Experimental results on public datasets show that the proposed method outperforms state-of-the-art approaches in both matching accuracy and visual effect. Jian Yang 0019, Ju Huang, Qiang Li 0042, Cong Wang 0033, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Refined Cascade Cost Volume for Multiview Remote Sensing Image ReconstructionabstractResearch on remote sensing multi-view stereo has significantly advanced the development of large-scale 3D urban reconstruction. However, existing frameworks encounter challenges with blurred edge details when processing aerial image, which impedes the accuracy of depth estimation. To address these limitations, we propose RC-MVS, the deep estimation network specifically tailored for remote sensing multi-view stereo tasks. This network aims to enhance the geometric details within the view space while effectively reducing match noise, achieving high-precision depth estimation. Specifically, we introduce a refined cascade framework that integrates geometric details with semantic information, ensuring both global structural consistency and local feature expressiveness. During the feature extraction phase, we redesign the feature space construction process and introduce a denoising feature pyramid module. This module reduces feature inconsistency and employs multiple denoising strategies to purify feature representations, thereby enhancing the accuracy of the matching process. Furthermore, to achieve progressive optimization of the depth range, we propose a progressive cross-layer fusion module. This module progressively fuses low-resolution cost volumes, reducing domain shifts between different data dimensions, thereby enhancing the understanding of fine structures within the depth map and the broader context. Experimental results show that the RC-MVS model performs exceptionally well on the LuoJia-MVS and WHU datasets, achieving superior quantitative and qualitative performance. Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Semantic-Guided Multiview Stereo Reconstruction for Aerial ImageabstractThe application of learning-based Multi-view Stereo (MVS) depth estimation methods has achieved significant results in large-scale 3D reconstruction benchmarks. However, adjacent terrains in aerial image interfere with depth estimation along building edges during matching process, leading to inaccurate results. To address these challenges, we propose a new end-to-end MVS network, named FuS-MVSNet, which fuses monocular depth probability as a semantic guidance into the multi-view geometry-based MVS framework. By combining the strengths of geometric consistency and local semantics, FuS-MVSNet achieves notable enhancements in both accuracy and robustness. Specifically, we first construct a monocular branch based on the pre-trained Depth Anything model to perform monocular metric depth estimation. The non-shared parameters ensure that the depth estimation process is independent of multi-view branch, focusing exclusively on semantic depth inference. Subsequently, to incorporate monocular features into the multi-view network, we introduce a volume adaptive fusion module, which adaptively integrates monocular feature volumes into the standard cost volume via an attention mechanism and guides the cost volume regularization. Finally, confidence-based dynamic selection between the two prediction branches ensures the selection of the more robust branch result under challenging conditions. Qualitative and quantitative results indicate that we achieve competitive performance on multiple benchmarks, including the WHU and LuoJia-MVS datasets. Wei Zhang 0250, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Parameter-Efficient Transfer Learning for Remote Sensing Image CaptioningabstractRemote sensing image captioning (RSIC) aims to generate accurate and concise textual descriptions for remote sensing (RS) images. It plays a significant role in the analysis of earth observation data. The success of Vision-and-Language Pre-training (VLP) models provides the foundation for their transfer to the RSIC task. To reduce the cost of transferring VLP models to downstream tasks, numerous Parameter-Efficient Transfer Learning (PETL) techniques have been proposed. However, most of them focus on fine-tuning general-purpose foundation models without fully considering the unique characteristics of remote sensing data. In this paper, we introduce PE-RSIC, a novel PETL framework tailored for RSIC. Specifically, the framework builds on a pre-trained BLIP-2 model while further designing a lightweight Cross-modal RS adapter (CRS-Adapter) and a Class Prompt. During training, all parameters of the pre-trained model remain frozen, and the newly added CRS-Adapter modules are updated to efficiently transfer vision-and-language knowledge from the natural domain to the RS domain. The Class Prompt is obtained by projecting the vision-encoded [CLS] token into the decoder, guiding the model to generate more accurate captions. This approach enables the model to capture critical RS class features that might be lost during the query decoding process, with only a minimal increase in parameters. Extensive experiments show that our PE-RSIC framework outperforms full fine-tuning while utilizing only 5% of the trainable parameters. Xuezhi Zhao, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Like Humans to Few-Shot Learning Through Knowledge Permeation of Visual and LanguageabstractFew-shot learning aims to generalize the recognizer from seen categories to an entirely novel scenario. With only a few support samples, several advanced methods initially introduce class names as prior knowledge for identifying novel classes. However, obstacles still impede achieving a comprehensive understanding of how to harness the mutual advantages of visual and textual knowledge. In this paper, we set out to fill this gap via a coherent Bidirectional Knowledge Permeation strategy called BiKop, which is grounded in human intuition: a class name description offers a moregeneralrepresentation, whereas an image captures thespecificityof individuals. BiKop primarily establishes a hierarchical joint general-specific representation through bidirectional knowledge permeation. On the other hand, considering the bias of joint representation towards the base set, we disentangle base-class-relevant semantics during training, thereby alleviating the suppression of potential novel-class-relevant information. Experiments on four challenging benchmarks demonstrate the remarkable superiority of BiKop, particularly outperforming previous methods by a substantial margin in the 1-shot setting (improving the accuracy by 7.58% onminiImageNet). Yuyu Jia, Junyu Gao 0001, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Multim. | 4 |
| 2025 | Confident Multi-View StereoabstractSolving the Multi-View Stereo (MVS) problem is a cornerstone in computer vision, with depth map estimation and fusion being one of the most critical approaches. The depth confidence map is pivotal in ensuring the precision and completeness of the reconstruction outcomes. These algorithms frequently encounter a trade-off between completeness and accuracy in the confidence map, which can significantly impair the final reconstruction results. This paper analyzes the causes and phenomena of these issues, namely Confidence Jitter, Confidence Gap, and Confidence Disappearance. From these insights, a multi-view stereo network named CF-MVSNet is introduced, comprising three essential components. Firstly, the method mitigates the Confidence Jitter problem through two confidence fusion strategies. Secondly, it narrows the depth sampling space to near sub-pixel levels, addressing the Confidence Gap through neighborhood-average pooling. Lastly, the algorithm tackles the Confidence Disappearance problem resulting from multi-scale classification and regression with a loss function named CL. Our proposed method demonstrates superior performance across two critical metrics: the completeness of the depth map and the accuracy of the reconstructed point cloud, outperforming current state-of-the-art MVS methods. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 2 |
| 2024 | An Efficient Subspace Partition Method Using Curve Fitting for Hyperspectral Band SelectionabstractBand selection is a valid method to reduce redundant information in hyperspectral images (HSIs). Typically, two methods are used to select representative bands: ranking bands based on predefined criteria and selecting cluster centers by grouping bands. It is advantageous to combine these two methods for hyperspectral band selection tasks since their benefits are complimentary. To take full of these advantages, we propose a hyperspectral band selection method using curve-fitting subspace partition (CFSP), including the subspace partition method and local context representative band selection method. The contributions of this letter are summarized below: 1) through fitting the spectral curves and then utilizing the point of maximum curvature to partition the band set, the similar and adjacent bands can be divided into the same group, which is very consistent with the way of subspace partition and 2) a representative band selection method in the local context is proposed. The locally optimal bands are selected sequentially to constitute the candidate band set. Then, through ranking and iteratively updating the candidate band set, we can effectively find the desired bands. The experiments on three public HSI datasets show that the proposed method has significant advantages compared with some advanced competitors. In particular, on the Salinas dataset, the selected bands achieved an excellent average overall accuracy (OA) of 91.32% using the support vector machine (SVM) classifier. Long Fang, Qiang Li 0042 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Neural Implicit Fourier Transform for Remote Sensing Shadow RemovalabstractRemote sensing shadow removal is an open issue. Previous studies focus on working in the spatial dimension, ignoring the potential of the Fourier dimension, while illumination degradation typically exists in the amplitude component. To address this limitation, our insight is a fresh dual-stage Fourier-based network (NeFour), which explores the best of both worlds between frequency and spatial information. In the frequency stage, we investigate the positive correlation between amplitude and brightness from channel and spatial statistics. Coupled with implicitly defined normalization, a controllable fitting amplitude transform map recreates the illumination. In the spatial stage, the inverted dark channel prior with 3-D coordinates serves as modulation matrices that naturally reveal the spatial distribution of shadows, thus elegantly eliminating shadow remnants. With ingenious design, NeFour achieves nontrivial performance against state-of-the-art shadow removal methods in terms of both visual perception and quantitative evaluation. The code is publicly available athttps://github.com/chi-kaichen/NeFour. Kaichen Chi, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | 3-D Neighborhood Cross-Differencing: A New Paradigm Serves Remote Sensing Change DetectionabstractChange detection is a prevalent technique in remote sensing image analysis for investigating geomorphological evolution. The modeling and analysis of difference features are crucial for the precise detection of land cover changes. In order to extract difference features, previous work has either directly computed them through differential operations or implicitly modeled them via feature fusion. However, these rudimentary strategies rely heavily on a high degree of congruence within the bitemporal feature space, which results in the model’s diminished capacity to capture subtle variations induced by factors such as differences in illumination. In response to this challenge, the concept of 3-D neighborhood difference convolution (3D-NDC) is proposed for robustly aggregating the intensity and gradient information of features. Furthermore, to delve into the deep disparities within bitemporal instance features, we propose a novel paradigm for differential feature extraction based on 3D-NDC, termed 3-D neighborhood cross-differencing. This strategy is dedicated to exploring the interplay of cross-temporal features, thereby unveiling the inherent disparities among various land cover characteristics. In addition, a detail-focused refinement (DfR) decode based on the Laplace operator has been designed to synergize with the 3-D neighborhood cross-differencing, aiming to improve the detail performance of change instances. This integration forms the basis of a new change detection framework, named ChangeLN. Extensive experiments demonstrate that ChangeLN significantly outperforms other state-of-the-art change detection methods. Moreover, the 3-D neighborhood cross-difference strategy exhibits the potential for integration into other change detection frameworks to improve detection performance. Open code is available fromhttps://github.com/weiAI1996/3DNCD_ChangeLN. Kaichen Chi, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Semantic-Explicit Filtering Network for Remote Sensing Image Change DetectionabstractRemote sensing image change detection (RSI-CD) aims to explore surface change information from aligned dual-phase images. However, RSI-CD currently encounters two major challenges. The first issue is the inadequate object-level semantic representation during the feature extraction in CD networks. The other issue is the spectral resolution of the RS image is limited, which leads to a mixture of pseudochange and real change. In order to explore the above-mentioned two challenges, we propose a semantic-explicit filtering network (SFNet) based on a neighborhood feature attention module (NFAM) and multiple-receptive-field semantic filtering mechanism (MSFM). First, the NFAM exploits the correlation of multiscale features and fuses features from the proximity layer to enhance the semantic-explicit representation of the object level. Then, the MSFM takes the weight map after the enhanced semantic representation as input and progressively refines the weight map through a multiple-receptive-field parallel convolution (MPC). This process filters out pseudochange from the predicted result while retaining the real-change information. The experiments on two benchmark datasets demonstrate that the proposed approach presents satisfactory performance over the existing methods. Yuemei Qin, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Edge-Guided Perceptual Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) plays a critical role in applications such as night navigation and fire rescue. Its primary purpose is to extract small targets from cluttered backgrounds. While deep learning-based methods have made great advancements in this field, there are still some limitations. One common issue is that the detected target shape tends to be smooth, and extremely small targets may not be effectively detected due to background interference. This article proposes an edge-guided perception network (EGPNet) for IRSTD to alleviate this trouble. To maintain the information of small targets, EGPNet utilizes a multiscale feature progressive fusion (MFPF) encoder to extract features. This progressive fusion manner enhances semantic information and contextual correlation. Considering that the detected target shapes may result in smoothing effect, an edge-guided image refinement module (EIRM) is incorporated to improve the integrity of the target shape. Moreover, we introduce a local target amplifier (LTA) to boost the visibility and representation of targets, while suppressing the clutter background interference. The experimental results illustrate that the proposed model can detect the targets with small and weak in different scenes well. Our code is publicly available athttps://github.com/qianngli/EGPNet. Qiang Li 0042, Zhigang Yang 0002, Yuan Yuan 0026, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Learning Remote Sensing Aleatoric Uncertainty for Semi-Supervised Change DetectionabstractSignificant progress has been recently achieved in the field of remote sensing image (RSI) change detection based on data-driven deep learning. Fully supervised models have limitations on the availability of massive annotated training data, while semi-supervised change detection (SSCD) has garnered increasingly widespread attention. Nevertheless, existing SSCD methods do not categorize the types of remote sensing aleatoric uncertainty (RSAU), let alone investigate the impact of uncertainty on performance. To this end, we define RSAU for SSCD and introduce the progressive uncertainty-aware and uncertainty-guided framework (PUF). It consists of two crucial components to perceive and guide the RSAU in the training stage. The first component, i.e., progressive uncertainty-aware learning (PUAL), decodes and quantifies the uncertainty inherent in the samples from the weak branch. The second one, i.e., uncertainty-guided multiview learning (UML), generates multiple image pairs designed for distortion and mixing for the strong branch. UML utilizes the uncertainty values derived from PUAL to offer guidance throughout the training process, which discerns and learns discriminative features from high-quality samples. Extensive experiments are conducted on three multiclass and building change detection (CD) benchmarks, i.e., CDD, SYSU, and LEVIR-CD. Furthermore, we propose a small dataset to enhance the understanding of aleatoric uncertainty, namely, LEVIR-AU. The proposed PUF consistently achieves state-of-the-art (SOTA) performance. The dataset and codes are available athttps://github.com/shenjh0/PUF. Jinhao Shen, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Semantic-Spatial Collaborative Perception Network for Remote Sensing Image CaptioningabstractImage captioning is a fundamental vision-language task with wide-ranging applications in daily life. The existing methods often struggle to accurately interpret the semantic information in remote sensing images due to the complexity of backgrounds. Target region masks can effectively reflect the shape characteristics of targets and their potential interrelationships. Therefore, incorporating and fully integrating these features can significantly improve the quality of generated captions. However, researchers are hindered by the lack of relevant datasets that contain corresponding object masks. It is natural to ask the following: how to efficiently introduce and utilize object masks? In this article, we provide potential target masks for the publicly available remote sensing image caption (RSIC) datasets, enabling models to utilize the regional features of targets for RSIC. Meanwhile, a novel RSIC algorithm is proposed that combines regional positional features with fine-grained semantic information, abbreviated as$\text {S}^{2}$CPNet. To effectively capture the semantic information from image and position relationship from mask, respectively, the semantic and spatial feature enhancement submodules are introduced at the ends of encoder branches, respectively. Furthermore, the cross-view feature fusion module is designed to integrate regional features and semantic information efficiently. Then, a target recognition decoder is developed to enhance the ability of model to identify and describe critical targets in images. Finally, we improve the caption generation decoder by adaptively merging textual information with visual features to generate more accurate descriptions. Our model achieves satisfactory results on three RSIC datasets compared with the existing method. The related datasets and code will be open-sourced inhttps://github.com/CVer-Yang/SSCPNet. Qi Wang 0009, Zhigang Yang 0002, Weiping Ni, Junzheng Wu, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | HCNet: Hierarchical Feature Aggregation and Cross-Modal Feature Alignment for Remote Sensing Image CaptioningabstractRemote sensing image captioning aims to describe the crucial objects from remote sensing images in the form of natural language. The inefficient utilization of object texture and semantic features in images, along with the ineffective cross-modal alignment between image and text features, are the primary factors that impact the model to generate high-quality captions. To alleviate this trouble, this paper presents a network for remote sensing image captioning, namely HCNet, including hierarchical feature aggregation and cross-modal feature alignment. Specifically, a hierarchical feature aggregation module is proposed to obtain a comprehensive representation of vision features, which is beneficial for producing accurate descriptions. Considering the disparities between different modal features, we design a cross-modal feature interaction module in the decoder to facilitate feature alignment. It can fully utilize cross-modal features to localize critical objects. Besides, a cross-modal feature align loss is introduced to realize the alignment between image and text features. Extensive experiments show our HCNet can achieve satisfactory performance. Especially, we demonstrate significant performance improvements of +14.15% CIDEr score on NWPU datasets compared to existing approaches. The source code is publicly available at https://github.com/CVer-Yang/HCNet. Zhigang Yang 0002, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | C²Net: Road Extraction via Context Perception and Cross Spatial-Scale Feature InteractionabstractRoad extraction from remote sensing images (RSIs) holds significant application value in various aspects of daily scenarios. However, it is still challenging to extract high-quality road results from RSIs due to the interference of objects sharing similar structures with roads in the background and the occlusion caused by surroundings. To alleviate these problems, a road extraction network based on the global-local Context perception and Cross spatial-scale feature interaction is proposed ($\text {C}^{2}$Net). First, a global-local context perception module (GLCPM) is incorporated to capture the overall topology features of the road, which aims to improve the ability of the model to discriminate between roads and similar objects. Then, the cross spatial-scale feature interaction module is designed in the skip connection to effectively aggregate full-scale features without loss of feature information, which can provide rich and accurate road structural features for the decoder. Experiments conducted on public road datasets demonstrate that$\text {C}^{2}$Net outperforms existing methods in terms of comprehensive metrics such as intersection over union (IoU) and the$F1$-score. The results indicate that$\text {C}^{2}$Net can produce road results with superior connectivity and quality. The source code will be publicly available athttps://github.com/CVer-Yang/CCNet. Zhigang Yang 0002, Wei Zhang 0250, Qiang Li 0042, Weiping Ni, Junzheng Wu, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Boosting Binary Object Change Detection via Unpaired Image Prototypes ContrastabstractBinary object change detection aims to monitor the evolution of the object of interest in a fixed region. Constructing a relevant dataset for deep learning models is strenuous. In the existing datasets, there is usually an imbalance between changed and unchanged samples, as well as a restricted diversity within the changed samples. Aiming at that, some methods utilize unpaired images used for object segmentation to generate pseudo-bitemporal images for change detection. However, due to the existence of the domain gap between different data sources, the model obtained by these methods can not well generalize to the real bitemporal images. Inspired by them but to avoid the domain difference, we explore how to directly use the unpaired images within a real change detection dataset to complement changed samples. In detail, a concise metric-based framework is designed, which consists of two branches, a projector and a predictor. The framework obtains the change map by computing the distance between the bitemporal embedding outputted by the projector. Meanwhile, instructed by an indirect semantic supervision module (ISSM) specially designed, the predictor can generate the semantic confidence map distinguishing the pixels in an image into two categories. Based on the output of the framework, an unpaired image prototype contrast module (UIPCM) is proposed. It enriches the diversity of the change samples for training by combining the prototypes in unpaired images at the feature level, leading to alleviating the imbalance between changed and unchanged samples. Besides, a dual margin contrastive loss (DMCL) is adopted during training. It can reduce the constraint on the consistency of bitemporal embedding in unchanged regions. The benefits and the superiority of the proposed method are demonstrated on two well-recognized datasets. The code is available at https://github.com/ptdoge/UIPC. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote SensingabstractLearnable multiview stereo (MVS) aerial image depth estimation has obtained great success in 3-D digital urban reconstruction. Currently, most depth estimation methods in the large-scale sense heavily involve adapting the general MVS framework. However, these methods often overlook the cross-view interval and limited viewpoint inherent in aerial images data. In this article, we introduce an learning-based MVS method for aerial image depth estimation, which enhances visual consistency to address the insufficient accuracy caused by the characteristics of aerial image data, namely, AggrMVS. First, an optical flow-guided feature extraction module is introduced to map the dynamic relationship between reference and source images. It explicitly captures edge information of different depth components to guide the cost volume regularization. Second, a cross-view volume fusion module is proposed to enhance the interaction among reference volumes, further improving the aggregation ability of the source volume. Furthermore, AggrMVS achieves refined aerial image depth estimation results with a lightweight cascade architecture. Since low-altitude oblique aerial datasets currently lack, we reconstruct a multicategory synthetic aerial scene benchmark from general MVS datasets. The benchmark dataset is available athttps://github.com/ToscW/BlendedUAV. Experiments on public and proposed datasets confirm that AggrMVS outperforms other MVS depth estimation methods in terms of qualitative and quantitative aspects. Wei Zhang 0250, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Multi-Domain Adaptation for Motion DeblurringabstractMotion deblurring is an important topic in the field of image enhancement, which has widespread applications including video surveillance, object detection, etc. Many algorithms are designed for motion deblurring and achieve remarkable performance. However, mainstream motion blur datasets are collected under normal weather and illuminance conditions, i.e., normal domain, ignoring their variations. As a result, current methods perform poorly in dynamic real-world scenes. To address these issues, we study the work in two aspects. First, we collect the real-world motion blur dataset with a well-designed collection device from various angles, focal lengths, and street scenes. Considering its domain is single, it is augmented via a Domain Transfer Strategy (DTS) to construct a Multi-Domain dataset (MD dataset), expanding the domains of the collected dataset. Second, we propose a Multi-Domain Adaptive Deblur Network (MDADNet) with two modules. The one is the Domain Adaptation (DA) module that exploits domain invariant features to stabilize the performance of the MDADNet in multiple domains. The other is the Meta Deblurring (MDB) module that employs the auxiliary branch to enhance the deblurring ability. It also enables the MDADNet to update parameters during the testing stage, improving the generalizations of the MDADNet. Extensive experimental results demonstrate that the MD-trained methods significantly strengthen the motion deblurring ability in multiple domains. Particularly, the proposed MDADNet achieves state-of-the-art performance on the MD dataset and public motion blur datasets. Kai Zhuang, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Multim. | 2 |
| 2023 | Edge Neighborhood Contrastive Learning for Building Change DetectionabstractBuilding change detection aims to identify the change in buildings in the same geographic area. Recently, many methods based on deep learning (DL) have achieved encouraging performance. However, some challenges remain in effectively exploiting the temporal–spatial correlation and achieving good discrimination in the neighborhood of the edge. To relieve these issues, we develop a selective attention module (SAM) to model the relationship between the semantic and the state (i.e., unchanged or changed) of the pixel, which is integrated into an existing metric learning-based architecture. Moreover, inspired by recent advances in contrastive learning, we present a novel edge neighborhood contrastive learning method to force the network to learn discriminative and compact features, leading to improving the accuracy of building change detection. Experimental results demonstrate that our method achieves competitive performance in terms of objective metrics and visual comparisons. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | RGB-Induced Feature Modulation Network for Hyperspectral Image Super-ResolutionabstractSuper-resolution (SR) is one of the powerful techniques to improve image quality for low-resolution (LR) hyperspectral image (HSI) with insufficient detail and noise. Traditional methods typically perform simple cascade or addition during the fusion of the auxiliary high-resolution RGB and LR HSI. As a result, the abundant HR RGB details are not utilized as a priori information to enhance the HSI feature representation, leaving room for further improvements. To address this issue, we propose an RGB-induced feature modulation network for HSI SR (IFMSR). Considering that similar patterns are common in images, a multi-corresponding patch aggregation is designed to globally assemble this contextual information, which is beneficial for feature learning. Besides, to adequately exploit plentiful HR RGB details, an RGB-induced detail enhancement (RDE) module and a deep cross-modality feature modulation (CFM) module are proposed to transfer the supplementary materials from RGB to HSI. These modules can provide a more direct and instructive representation, leading to further edge recovery. Experiments on several datasets demonstrate that our approach achieves comparable performance under more realistic degradation condition. Our code is publicly available at https://github.com/qianngli/IFMSR. Qiang Li 0042, Maoguo Gong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Block Diagonal Representation Learning for Hyperspectral Band SelectionabstractHyperspectral band selection is viewed as an effective dimension reduction method. Recently, researchers present graph-based clustering for hyperspectral image (HSI) processing. However, most of them conduct clustering on a fixed data matrix so that it is sensitive to the quality of initial matrix. Moreover, these algorithms apply spectral clustering to obtain the final clustering result in increasing the time consumption. Based on these facts, we propose a block diagonal representation learning algorithm (BDRLA) in this paper. BDRLA generates a high-quality similarity matrix by approximating the initial affinity matrix. Meanwhile, motivated to the spectral bands distance similarity matrix has a clear diagonal structure, a block diagonal similarity matrix with ordered partition points based upon the ℓ2-norm is constructed. By doing so, the obtained similarity matrix is directly applied to subsequent processing without extracting the clustering indicators. Additionally, in order to estimate the importance of bands, dictionary learning is adopted to select the representative band in each cluster. Extensive experiment results on three public datasets indicate that the bands selected by the proposed method achieve satisfactory performance. Long Fang, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hyperspectral Band Selection via Difference Between IntergroupsabstractVarious methods are proposed to reduce the dimensions of hyperspectral image by band selection in recent years. Most methods select one band from each group to construct a band subset. However, the redundancy in the selected bands from different groups is neglected. Furthermore, the researchers do not pay enough attention to how many bands are appropriated for selection. To solve these issues, we propose a hyperspectral band selection method via difference between inter-groups (DIG), which includes grouping strategy and ranking strategy. Specifically, the grouping strategy adopts intra-group similarity to reasonably distribute all partitioning point positions. The similarity of bands within the same group is significantly improved. For the ranking strategy, it not only takes into account the knowledge and intra-group similarity of bands, but also evaluates the differences between each band and other inter-group bands. The redundancy in band subset is reduced sufficiently. In order to accurately obtain the optimal number of bands, an evaluation function is designed to measure the information content and redundancy in various band subsets. Experimental results from different aspects show that the proposed model has a large performance advantage on three public datasets. Baidong Peng, Long Fang, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Multiscale Factor Joint Learning for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (SR) using auxiliary RGB image has obtained great success. Currently, most methods respectively train single model to handle different scale factors, which may lead to the inconsistency of spatial and spectral contents when converted to the same size. In fact, the manner ignores the exploration of potential interdependence among different scale factors in single model. To this end, we propose a multi-scale factor joint learning for hyperspectral image super-resolution (MulSR). Specifically, to take advantage of the inherent priors of spatial and spectral information, a deep architecture using single scale factor is designed by terms of symmetrical guided encoder (SGE) to explore the hyperspectral image and RGB image. Considering that there are obvious differences in texture details at various scale factors, another architecture is proposed which is basically the same as above, except that its scale factor is larger. On this basis, a multi-scale information interaction (MII) unit is modeled between two architectures by a direction-aware spatial context aggregation (DSCA) module. Besides, the contents generated by the model with multi-scale factor are combined to build a learnable feedback compensation correction (LFCC). The difference is fed back to the architecture with large scale factor, forming an interactive feedback joint optimization pattern. This calibrates the representation of spatial and spectral contents in the reconstruction process. Experiments on synthetic and real datasets demonstrate that our MulSR shows superior performance in terms of qualitative and quantitative aspects. Our code is publicly available at https://github.com/qianngli/MulSR. Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Difference-Guided Aggregation Network With Multiimage Pixel Contrast for Change DetectionabstractChange detection is a critical task in remote sensing to monitor the state of the surface on Earth. This field has been dominated by deep learning-based methods recently. Many models that model the temporal-spatial correlation in bitemporal images through the non-local interaction between bitemporal features achieve impressive performance. However, under complex scenes including multiple change types or weakly discriminate objects, they suffer from achieving discriminative fusion of information due to the weak semantic discrimination of the bitemporal representations. Aiming at this problem, a difference-guided aggregation network (DGANet) is proposed, where two key modules are injected, i.e., a difference-guided aggregation module (DGAM) and a weighted metric module (WMM). The bitemporal features in DGAM are aggregated with the guidance of their differences, which focuses on their change relevance and relaxes their semantic distinction. Therefore, the fused features are change-relevant and discriminative. WMM aims to achieve adaptive distance computation between the bitemporal features by dynamic feature attention in different dimensions. It is helpful to suppress the pseudo-changes. Besides, a change magnitude contrastive loss (CMCL) is introduced to employ the dependency of bitemporal pixels in different bitemporal images, which further enhances the representation quality of the model. Meanwhile, it is further extended in this work. The effectiveness of the three improvements is demonstrated by extensive ablation studies. The results on three datasets widely used illustrate that our method achieves satisfactory performance. Qiang Li 0042, Yanling Miao, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Single-Shot Balanced Detector for Geospatial Object DetectionabstractGeospatial object detection is an essential task in remote sensing community. One-stage methods based on deep learning have faster running speed but cannot reach higher detection accuracy than two-stage methods. In this paper, to achieve excellent speed/accuracy trade-off for geospatial object detection, a single-shot balanced detector is presented. First, a balanced feature pyramid network (BFPN) is designed, which can balance semantic information and spatial information between high-level and shallow-level features adaptively. Second, we propose a task-interactive head (TIH). It can reduce the task misalignment between classification and regression. Extensive experiments show that the improved detector obtains significant detection accuracy with considerable speed on two benchmark datasets. Yanfeng Liu, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2022 | Hyperspectral image super-resolution via multi-domain feature learning
Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 1 |
| 2022 | CDD-Net: A Context-Driven Detection Network for Multiclass Object DetectionabstractUnlike object detection in natural images that usually achieved great success, remote sensing imagery has its own challenges to detect and localize multiclass objects, such as large-scale change, uncertain direction, and high density. The context information of the objects is very worthwhile for solving these challenges in remote sensing images. In this letter, we propose a context-driven detection network (CDD-Net) to improve the accuracy of multiclass object detection in remote sensing images. For capturing the local neighboring objects and features, a local context feature network (LCFN) is proposed to learn the local context of the region of interest. Meanwhile, a hybrid attention pyramid network (HAPN) is designed, which can steer the focus to more valuable features. The HAPN inserts a squeeze and excitation block (SEB) and three asymmetric convolution blocks (ACBs) in the feature pyramid network (FPN). The experimental results over the DOTA-v1.5 data set demonstrate that the proposed CDD-Net yields promising results. Ke Zhang 0014, Jingyu Wang 0002, Yezi Wang, Qi Wang 0009, Qiang Li 0042 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Symmetrical Feature Propagation Network for Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image (HSI) super-resolution (SR) methods using a auxiliary high-resolution (HR) RGB image have achieved great progress recently. However, most existing methods aggregate the information of RGB image and HSI early during input or shallow feature extraction, whose difference between two images has not been treated and discussed. Although a few methods combine both the image features in the middle layer of the network, they fail to make full use of the two inherent properties, i.e., rich spectra of HSI and HR content of RGB image, to guide model representation learning. To address these issues, in this article, we propose a dual-stage learning approach for HSI SR to learn a general spatial–spectral prior and image-specific details, respectively. In the coarse stage, we fully take advantage of two adjacent bands and RGB image to build the model. During coarse SR, a symmetrical feature propagation approach is developed to learn the inherent content of each image over a relatively long range. The symmetrical structure encourages the two streams to better retain their particularity. Meanwhile, it can realize the information interaction by the adaptive local block aggregation (ALBA) module. To learn image-specific details, a back-projection refinement network is embedded in the structure, which further improves the performance in fine stage. The experiments on four benchmark datasets demonstrate that the proposed approach presents excellent performance over the existing methods. Our code is publicly available athttps://github.com/qianngli/SFPN. Qiang Li 0042, Maoguo Gong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | ABNet: Adaptive Balanced Network for Multiscale Object Detection in Remote Sensing ImageryabstractBenefiting from the development of convolutional neural networks (CNNs), many excellent algorithms for object detection have been presented. Remote sensing object detection (RSOD) is a challenging task mainly due to: 1) complicated background of remote sensing images (RSIs) and 2) extremely imbalanced scale and sparsity distribution of remote sensing objects. Existing methods cannot effectively solve these problems with excellent detection accuracy and rapid speed. To address these issues, we propose an adaptive balanced network (ABNet) in this article. First, we design an enhanced effective channel attention (EECA) mechanism to improve the feature representation ability of the backbone, which can alleviate the obstacles of complex background on foreground objects. Then, to combine multiscale features adaptively in different channels and spatial positions, an adaptive feature pyramid network (AFPN) is designed to capture more discriminative features. Furthermore, considering that the original FPN ignores rich deep-level features, a context enhancement module (CEM) is proposed to exploit abundant semantic information for multiscale object detection. Experimental results on three public datasets demonstrate that our approach exhibits superior performance over baseline by only introducing less than 1.5M extra parameters. Yanfeng Liu, Qiang Li 0042, Yuan Yuan 0001, Qian Du 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Dual-Stage Approach Toward Hyperspectral Image Super-ResolutionabstractHyperspectral image produces high spectral resolution at the sacrifice of spatial resolution. Without reducing the spectral resolution, improving the resolution in the spatial domain is a very challenging problem. Motivated by the discovery that hyperspectral image exhibits high similarity between adjacent bands in a large spectral range, in this paper, we explore a new structure for hyperspectral image super-resolution (DualSR), leading to a dual-stage design, i.e., coarse stage and fine stage. In coarse stage, five bands with high similarity in a certain spectral range are divided into three groups, and the current band is guided to study the potential knowledge. Under the action of alternative spectral fusion mechanism, the coarse SR image is super-resolved in band-by-band. In order to build model from a global perspective, an enhanced back-projection method via spectral angle constraint is developed in fine stage to learn the content of spatial-spectral consistency, dramatically improving the performance gain. Extensive experiments demonstrate the effectiveness of the proposed coarse stage and fine stage. Besides, our network produces state-of-the-art results against existing works in terms of spatial reconstruction and spectral fidelity. Our code is publicly available at https://github.com/qianngli/DualSR. Qiang Li 0042, Yuan Yuan 0001, Xiuping Jia, Qi Wang 0009 |
IEEE Trans. Image Process. | 1 |
| 2021 | Hyperspectral Image Super-Resolution Via Adjacent Spectral Fusion StrategyabstractHyperspectral image exhibits low spatial resolution due to the limitation of imaging system. Improving it without an auxiliary high resolution (HR) image still remains a challenging problem. Recently, although many deep learning-based hyperspectral image super-resolution (SR) methods have been proposed, they make the insufficient utilization of adjacent bands to improve the reconstruction performance. To address this issue, we explore a new structure for hyperspectral image SR via adjacent spectral fusion strategy. Inspired by the high similarity among adjacent bands, neighboring band partition is proposed to divide the adjacent bands into several groups. Through the current band, the adjacent bands is guided to enhance the exploration ability. To explore more complementary information, an alternative fusion mechanism, i.e., intra-group fusion and inter-group fusion, is designed, which helps to recover the missing details in the current band. Experiments demonstrate that our approach produces the state-of-the-art results over the existing approaches. Qiang Li 0042, Qi Wang 0009, Xuelong Li 0001 |
ICASSP | 1 |
| 2021 | Hyperspectral Image Super-Resolution via Multi-Domain Feature LearningabstractHyperspectral image super-resolution (SR) methods are continually being refreshed due to deep neural networks. Despite this, the existing works barely explore more spatial information using mixed 2D/3D convolution. Moreover, they do not make full use of multi-domain features to realize information complementation. To tackle these challenges, we propose a hyperspectral image SR approach via multi-domain feature learning. To be specific, a multi-domain feature learning strategy using 2D/3D unit is presented to explore spatial and spectral information by alternate manner. To recover the more details, the edge body generation mechanism (EBGM) is introduced to learn the high frequency information, which generates the edge prior. Besides, the multi-domain feature fusion (MDFF) is designed to fully integrated hierarchical know ledge from different 2D/3D units, leading to further achieve information complementation. Experiments demonstrate that our approach attains the better performance over the state-of-the-art methods. Qiang Li 0042, Qi Wang 0009, Xuelong Li 0001 |
IGARSS | 1 |
| 2021 | Exploring the Relationship Between 2D/3D Convolution for Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (SR) methods based on deep learning have achieved significant progress recently. However, previous methods lack the joint analysis between spectrum and horizontal or vertical direction. Besides, when both 2D and 3D convolution are in the network, the existing models cannot effectively combine the two. To address these issues, in this article, we propose a novel hyperspectral image SR method by exploring the relationship between 2D/3D convolution (ERCSR). Our method alternately employs 2D and 3D units to solve the problem of structural redundancy by sharing spatial information during reconstruction for existing model, which can enhance the learning ability of 2D spatial domain. Importantly, compared with the network using 3D units, i.e., 2D units are replaced by 3D units, it can not only reduce the size of the model but also improve the performance of the model. Furthermore, to exploit the spectrum fully, the split adjacent spatial and spectral convolution (SAEC) is designed to parallelly explore information between spectrum and horizontal or vertical direction in space. Experiments on widely used benchmark datasets demonstrate that the proposed approach outperforms state-of-the-art SR algorithms across different scales in terms of quantitative and qualitative analysis. Qiang Li 0042, Qi Wang 0009, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | A Fast Neighborhood Grouping Method for Hyperspectral Band SelectionabstractHyperspectral images can provide dozens to hundreds of continuous spectral bands, so the richness of information has been greatly improved. However, these bands lead to increasing complexity of data processing, and the redundancy of adjacent bands is large. Recently, although many band selection methods have been proposed, this task is rarely handled through the context information of the whole spectral bands. Moreover, the scholars mainly focus on the different numbers of selected bands to explain the influence by accuracy measures, neglecting how many bands to choose is appropriate. To tackle these issues, we propose a fast neighborhood grouping method for hyperspectral band selection (FNGBS). The hyperspectral image cube in space is partitioned into several groups using coarse-fine strategy. By doing so, it effectively mines the context information in a large spectrum range. Compared with most algorithms, the proposed method can obtain the most relevant and informative bands simultaneously as subset in accordance with two factors, such as local density and information entropy. In addition, our method can also automatically determine the minimum number of recommended bands by determinantal point process. Extensive experimental results on benchmark data sets demonstrate the proposed FNGBS achieves satisfactory performance against state-of-the-art algorithms. Qi Wang 0009, Qiang Li 0042, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |