EDBT 2026 Demo / reviewers in the wild / expert
Dou Quan
dblp:189/3644
· DBLP profile ↗
47ranked-venue papers
12as first author
40since 2021 · last 2025
0000-0001-6943-4657ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 9 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Feature Spectrum Learning for Remote Sensing Change DetectionabstractChange detection (CD) holds significant implications for Earth observation, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing methods mainly regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). Nevertheless, their efforts are limited by the complexity of optimizing GANs and the absence of guidance from physical properties. This paper finds that the spectrum transformation (ST) has the potential to mitigate pseudo-changes by aligning in the frequency domain carrying the style. However, the benefit of ST is largely constrained by two drawbacks: 1) limited transformation space and 2) inefficient parameter search. To address these limitations, we propose a Feature Spectrum learning (FeaSpect) that adaptively eliminate pseudo-changes in the latent space. For the drawback 1), FeaSpect directs the transformation towards stylealigned discriminative features via feature spectrum transformation (FST). For the drawback 2), FeaSpect allows FST to be trainable, efficiently discovering optimal parameters via extraction box with adaptive attention and extraction box with learnable strides. Extensive experiments on challenging datasets demonstrate that our method remarkably outperforms existing methods and achieves a commendable trade-off between accuracy and efficiency. Importantly, our method can be easily injected into other frameworks, achieving consistent improvements. Qi Zang, Dong Zhao 0007, Shuang Wang 0001, Dou Quan, Zhun Zhong |
CVPR | 4 |
| 2025 | Predicting Spectral Information for Self-Supervised Signal ClassificationabstractDeep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as labeling necessitates expert knowledge. This paper proposes a novel self-supervised signal classification method called Spectral-Guided Self-Supervised Signal Classification (SGSSC). Specifically, to leverage frequency-domain information with modulation semantics as prior knowledge for the model, we design a previously unexplored pretext task tailored to the format of signal data. This task involves predicting spectral information from masked time-domain signals, enabling the model to learn implicit signal features through cross-domain pattern transformation. Furthermore, the pretext task in the SGSSC method is relevant to the downstream classification task, and using traditional fine-tuning strategies on the downstream task may lead to the loss of certain features associated with the pretext task. Therefore, we propose an attention mechanism-based fine-tuning strategy that adaptively integrates pre-trained features from different levels. Extensive experimental results validate the superiority of the SGSSC method. For instance, when the proportion of labeled samples is only 0.5%, our method achieves an average improvement of 2.3% in downstream classification tasks compared to the best-performing self-supervised training strategies. Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Dou Quan, Rui Yang 0038, Dong Zhao 0007, Luyang Mei |
IJCAI | 5 |
| 2025 | A two-stage strategy for brain-inspired unsupervised learning in spiking neural networks
Chuanfeng Ma, Biao Hou, Leida Li, Hao Zhu 0009, Dou Quan, Licheng Jiao |
Neurocomputing | 7 |
| 2025 | FASTCC: A lightweight human pose detection method leveraging SimCCabstractAs a focal area within computer vision algorithms, human pose estimation algorithms find applications in diverse fields such as security and virtual reality. Achieving a balance between speed and accuracy is imperative for practical applications. Existing methods often present a trade-off between high accuracy and real-time performance. In response, this paper introduces the Fast Coordinate Classification (FastCC) detection head. It employs a shared fully connected Transformer for global self-attention operations on feature layers from the backbone network. The spatial attention coordinate encoder then outputs the coordinates of the horizontal and vertical axes of the keypoints, which are subsequently combined to derive the actual keypoint positions. Experimental evaluations conducted on the COCO and MPII datasets demonstrate that our detection head enhances the accuracy of human pose estimation algorithms while maintaining a lightweight design, outperforming the traditional heatmap method. Yi Li 0054, Yongtao Wang, Dou Quan, Yabo Yan, Qinghai Yang |
Neurocomputing | 4 |
| 2025 | AFLNet: Auxiliary Feature Learning-Guided Cross-Channel Automatic Modulation ClassificationabstractThis paper conducted a thorough investigation into the primary difficulty of the cross-channel automatic modulation classification (AMC) task by examining data distribution and feature space of different channel conditions. We concluded that the disruption of the target channel feature space structure breakdown the mapping relationship across channels, serving as the main contributor to model performance degradation. Based on the above conclusion, in order to improve the performance of cross-channel AMC, we introduce the Auxiliary Feature Learning-Guided Network (AFLNet). This network improves the structure of the target feature space through two uniquely designed tasks and facilitates efficient cross-domain alignment via a collaborative alignment mechanism. Specifically, AFLNet integrates similarity-based and confidence-based auxiliary feature learning tasks to enhance the discriminability of the target feature space and maintain the correspondence of category structures across different channels, thereby reducing the difficulty of feature alignment. The collaborative alignment mechanism combines adversarial training-based and self-training-based feature alignment methods, leveraging their mutually reinforcing effect and complementary strengths in global alignment and class-level alignment to enhance overall alignment performance. We carried out extensive experiments across four scenarios characterized by substantial channel variations, verifying that AFLNet achieves state-of-the-art with accuracy improvement of up to 9.71%. Hantong Xing, Shuang Wang 0001, Chenxu Wang 0004, Dou Quan, Hanlin Mo, Luyang Mei, Huaji Zhou, Licheng Jiao |
IEEE Trans. Commun. | 4 |
| 2025 | Burden-Free Distillation From Foundation Model for Efficient Remote Sensing Change DetectionabstractApplying vision foundation models to remote sensing change detection (CD) has attracted extensive research attention. These studies employ inherent general knowledge from vision foundation models to enhance CD performance. Existing methods explicitly employ the foundation model as a feature extractor while designing additional learnable modules to bridge the task gap. However, these methods substantially increase the computational burden and memory demand in the inference. This paper therefore focuses on addressing the core challenge of effectively leveraging the knowledge from vision foundation models to enhance CD performance while maintaining computational efficiency. Instead of explicitly utilizing the foundation model, we propose Burden-Free Distillation (BFD), an architecture-agnostic foundation model-based distillation framework for efficient CD. BFD transfers the general knowledge from foundation models to task-specific models, thereby eliminating the dependency on foundation models during inference. Specifically, BFD transfers the foundation model knowledge through Dual-temporal Feature Matching module (DFM). This module enables multi-level feature alignment by computing pixel-wise spatial similarity between the foundation models’ general features and the CD models’ task-specific features. Additionally, we leverage patch contrastive distillation, which transfers localized structural patterns to the CD model to further mitigate task discrepancies between foundation models and CD models. We conduct extensive experiments across multiple foundation models and CD architectures, experimental results demonstrate that BFD effectively adapts the knowledge of foundation models to CD tasks without additional computational burden. Compared to other foundation model-based CD methods, BFD reduces the model parameters by 80.3% and improves IoU by 1.78% on the S2Looking dataset. The code is available at https://github.com/Younger-hua/Burden-Free-Distillation. Shuang Wang 0001, Chonghua Lv, Dou Quan, Ning Huyan, Xianwei Cao, Jingxi Sun, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Generalization-Aware Remote Sensing Change Detection via Domain-Agnostic LearningabstractChange detection has essential significance for the region's development, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing transformation-based methods regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). However, their efforts are limited by two drawbacks: 1) Transformed images suffer from distortion that reduces feature discrimination. 2) Alignment hampers the model from learning domain-agnostic representations that degrades performance on scenes with domain shifts from the training data. Therefore, oriented from pseudo-changes caused by style differences, we present a generalizable domain-agnostic difference learning network (DonaNet). For the drawback 1), we argue for local-level statistics as style proxies to assist against domain shifts. For the drawback 2), DonaNet learns domain-agnostic representations by removing domain-specific style of encoded features and highlighting the class characteristics of objects. In the removal, we propose a domain difference removal module to reduce feature variance while preserving discriminative properties and propose its enhanced version to provide possibilities for eliminating more style by decorrelating the correlation between features. In the highlighting, we propose a cross-temporal generalization learning strategy to imitate latent domain shifts, thus enabling the model to extract feature representations more robust to shifts actively. Extensive experiments conducted on three public datasets demonstrate that DonaNet outperforms existing state-of-the-art methods with a smaller model size and is more robust to domain shift. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Licheng Jiao |
IEEE Trans. Multim. | 4 |
| 2025 | Seeking a Hierarchical Prototype for Multimodal Gesture RecognitionabstractGesture recognition has drawn considerable attention from many researchers owing to its wide range of applications. Although significant progress has been made in this field, previous works always focus on how to distinguish between different gesture classes, ignoring the influence of inner-class divergence caused by gesture-irrelevant factors. Meanwhile, for multimodal gesture recognition, feature or score fusion in the final stage is a general choice to combine the information of different modalities. Consequently, the gesture-relevant features in different modalities may be redundant, whereas the complementarity of modalities is not exploited sufficiently. To handle these problems, we propose a hierarchical gesture prototype framework to highlight gesture-relevant features such as poses and motions in this article. This framework consists of a sample-level prototype and a modal-level prototype. The sample-level gesture prototype is established with the structure of a memory bank, which avoids the distraction of gesture-irrelevant factors in each sample, such as the illumination, background, and the performers' appearances. Then the modal-level prototype is obtained via a generative adversarial network (GAN)-based subnetwork, in which the modal-invariant features are extracted and pulled together. Meanwhile, the modal-specific attribute features are used to synthesize the feature of other modalities, and the circulation of modality information helps to leverage their complementarity. Extensive experiments on three widely used gesture datasets demonstrate that our method is effective to highlight gesture-relevant features and can outperform the state-of-the-art methods. Yunan Li 0001, Tianyu Qi, Zhuoqi Ma, Dou Quan, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Boosting Generalization of Semantic Segmentation With Unseen Style Seeking-Based Meta-LearningabstractThis article considers a worst and most challenging scene in domain generalization (DG), where a model aims to generalize well on unseen domains while only one single domain is available for training. Existing randomization-based methods achieve this goal by enriching the style of the training data. However, they fail to guarantee the diversity of newly generated data required for generalization and thus lead to insufficient expansion of the training distribution. Thus, we propose a novel single DG (SDG) framework, unseen style seeking-based meta-learning (USSML). In USSML, multiple plausible domains with various styles are first constructed from a single source domain and the combination is performed across generated domains to emulate unseen images, extending the distribution boundaries of the source domain. The domain combination is performed at two levels, i.e., global and instance, to meet the generalization challenge in semantic segmentation. Then, the generated diverse domains are further exploited to force the model to optimize in an unbiased manner across all domains by relearning regions lacking domain-invariant representation capability, driving the model toward domain invariance. A point worth mentioning is that the proposed method is easily integrated into existing segmentation methods with little computational cost to improve their generalization. Extensive experiments are conducted on five popular segmentation datasets and the results have verified the effectiveness of USSML in improving the model's generalization and the superiority of USSML over existing works. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wanqing Li 0001, Dou Quan, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Watching it in Dark: A Target-Aware Representation Learning Framework for High-Level Vision Tasks in Low Illumination
Yunan Li 0001, Shoude Li, Dou Quan, Chaoneng Li, Qiguang Miao |
ECCV (75) | 5 |
| 2024 | A Dual-Branch Random Mask Alignment Framework for Semi-Supervised PolSAR Terrain ClassificationabstractDespite the recent success of deep learning based polarimetric synthetic aperture radar(PolSAR) classification algorithms, it remains challenging in the scenario of limited labeled samples. Existing semi-supervised PolSAR terrain classification methods focus on the exploitation of pseudolabels, which are unreliable with limited labeled samples. To solve this problem, we propose a dual-branch random mask alignment framework for PolSAR terrain classification task. First, we propose a balanced regional expansion algorithm for labeled sample expansion. Then, to fully exploit the massive unlabeled samples, we designed a dual-branch network using two different polarization decomposition features as inputs, and a random mask alignment loss is employed to achieve consistency constraints on the unlabeled samples. Experimental results on two PolSAR datasets demonstrate that the proposed method achieve excellent performance with limited labeled samples. Tianquan Bian, Zhuangzhuang Sun, Shuang Wang 0001, Dou Quan, Yanhe Guo |
IGARSS | 5 |
| 2024 | Enhancing Change Detection Robustness: A Whitening Transformation ApproachabstractThe vast amount of remote sensing data has been instrumental in supporting research on change detection algorithms based on deep learning.However, factors such as geographical changes and variations in sensor parameters can result in significant style differences between remote sensing images at different time points, leading to a decline in model performance.To address this issue, this paper proposes a change detection algorithm based on whitening feature extraction, aiming to alleviate distribution differences by decoupling domain-invariant discriminative features from domain-specific style features.The effectiveness of the proposed method is demonstrated through transfer experiments from the SVCD dataset to the SZADA dataset and from the SYSU dataset to the SZADA dataset. Qi Zang, Zhengyao Wang, Dou Quan, Shuang Wang 0001 |
IGARSS | 5 |
| 2024 | Fourier Domain Adaptive Multi-Modal Remote Sensing Image Template Matching Based on Siamese NetworkabstractMulti-modal remote sensing image template matching is a meaningful and crucial topic in remote sensing image processing. However, due to different imaging mechanisms, there are significant nonlinear radiometric variations among multi-modal remote sensing images, increasing the matching challenge and leading to poor matching performances. To tackle this issue, this paper proposes a Fourier Domain Adaptive Network (FDANet) for multi-modal remote sensing image matching. Firstly, FDANet randomly swaps the low-frequency spectrum information between multi-modal images through the Fourier transform to reduce differences among multi-modal images, enhancing network adaptability to different image modalities and improving the multi-modal image matching performance. Secondly, FDANet extracts domain-invariant features from the transformed images through a deep Siamese network. After that, FDANet performs template matching and achieves high-precision multi-modal remote sensing image matching. In addition, we adopt the contrastive learning loss to optimize the FDANet. Extensive experiments on multi-modal remote sensing image matching demonstrate the effectiveness and advantages of the proposed FDANet. Chonghua Lv, Dou Quan, Shuang Wang 0001, Xiangming Jiang, Yu Gu 0015, Licheng Jiao |
IGARSS | 3 |
| 2024 | Generalized Source-Free Domain-adaptive Segmentation via Reliable Knowledge Propagation
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Jinlong Li 0003, Nicu Sebe, Zhun Zhong |
ACM Multimedia | 5 |
| 2024 | LM-Net: A Lightweight Matching Network for Remote Sensing Image Matching and RegistrationabstractDeep feature learning methods have shown significant advantages over handcrafted feature-based methods in remote sensing image matching and registration. Existing deep learning methods usually introduce complex modules into the deep convolutional network for more robust feature learning. However, they usually require high computation and memory resources for the computing device and have expensive time costs for image registration. As a basic image-processing task, it is crucial to build a lightweight matching network (LM-Net) for fast and accurate image matching and registration. Unfortunately, the image-matching performance will decrease significantly when we directly compress the deep model to a lightweight one. This article proposes an LM-Net based on the knowledge distillation (KD) learning framework for remote sensing image matching and registration. We first build an LM-Net with three convolutional layers. Then, this article proposes an effective KD approach for network optimization, which transfers the effective knowledge from the deep matching network to LM-Net to improve image-matching performances. Specifically, this article considers the useful information in the instance samples and the relation information between samples. It designs the feature and feature relation distillation learning for LM-Net training. Extensive experimental results and analysis have shown the effectiveness and advantages of the proposed LM-Net. LM-Net can reduce the number of parameters and computational complexity of the matching network. Meanwhile, LM-Net can significantly decrease the time cost and achieve results comparable to those of the deep model. It reduces the average image registration time by 42% on remote sensing image matching and registration. Additionally, LM-Net generalizes well on other multimodal remote sensing images. Dou Quan, Chonghua Lv, Shuang Wang 0001, Yi Li 0054, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | F3Net: Adaptive Frequency Feature Filtering Network for Multimodal Remote Sensing Image RegistrationabstractMultimodal remote sensing image registration is crucial for multimodal information fusion and applications. The significant nonlinear appearance difference between multimodal images caused by the various imaging mechanisms dramatically increases the challenge of image registration. This article proposes an adaptive frequency feature filtering network (F3Net) for cross-modal remote sensing image registration. On the one hand, F3Net explicitly explores the useful frequency components across modal images based on multilevel deep features. On the other hand, F3Net can take advantage of the nonlocal receptive fields by frequency modulation for feature learning and boosting image registration performances. F3Net inserts frequency feature filtering (F3) modules in multilevel deep features. Specifically, F3Net first performs the fast Fourier transform (FFT) for deep features. Then, F3Net designs a frequency attention (FA) module to adaptive enhance the shared and discriminative frequency features between multimodal images while suppressing the frequency components that hinder the cross-modal image registration. In addition, F3Net adopts multiscale frequency filtering fusion to facilitate discriminative feature learning, including global frequency feature filtering (GF3) based on the global image spectrum and local frequency feature filtering (LF3) based on the spectrum of stacked image regions. Experimental results on many remote sensing images have demonstrated the efficiency of the F3Net on multimodal image registration. Dou Quan, Shuang Wang 0001, Yunan Li 0001, Bo Ren 0001, Mengte Kang, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multi-View Feature Fusion and Visual Prompt for Remote Sensing Image CaptioningabstractRemote sensing image (RSI) captioning is a vision-language multimodal task concentrating on both image comprehension and sentence generation. Several studies suggest that encoder–decoder-based methods have achieved success in RSI captioning. However, existing encoder–decoder-based methods may not fully explore image representations for RSI captioning and suffer from a lack of additional prompt information for sentence generation. In this article, a novel multi-view feature fusion and prompt (MVP)-based model is proposed to obtain better RSI representations and enhance language model performance in RSI captioning. Specifically, we design an attention-based feature fusion module to dynamically fuse multi-view visual features, which are extracted from the fine-tuned vision-language pretraining (VLP) model and the vision-task pretraining (VP) model. Then, a flexible visual prefix mapping module is proposed to transform images into visual prefixes, providing semantic information for the subsequent sentence generation. Finally, a BERT-based caption generator is applied to generate accurate descriptions based on the fused visual features and the visual prefixes, which are both outputs from our designed modules. Extensive experiments are conducted on three well-known benchmark datasets, demonstrating that our method achieves state-of-the-art (SOTA) performance. The relevant code is available athttps://github.com/QiaoLing-Lin/MVP. Shuang Wang 0001, Qiaoling Lin, Xiutiao Ye, Yu Liao, Dou Quan, ZhongQian Jin, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Transcending Fusion: A Multiscale Alignment Method for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multiscale representations in image content and text vocabulary can enable the models to learn richer representations and enhance retrieval. Current multiscale RSITR approaches typically align multiscale fused image features with text features but overlook aligning image-text pairs at distinct scales separately. This oversight restricts their ability to learn joint representations suitable for effective retrieval. We introduce a novel multiscale alignment (MSA) method to overcome this limitation. Our method comprises three key innovations: 1) a multiscale cross-modal alignment transformer (MSCMAT), which computes cross-attention between single-scale image features and localized text features, integrating global textual context to derive a matching score matrix within a mini-batch; 2) a multiscale cross-modal semantic alignment loss (MSCMA loss) that enforces semantic alignment across scales; and 3) a cross-scale multimodal semantic consistency loss (CSMMC loss) that uses the matching matrix from the largest scale to guide alignment at smaller scales. We evaluated our method across multiple datasets, demonstrating its efficacy with various visual backbones and establishing its superiority over existing state-of-the-art methods. The GitHub URL for our project ishttps://github.com/yr666666/MSA. Rui Yang 0038, Shuang Wang 0001, Yingping Han, Yuanheng Li, Dong Zhao 0007, Dou Quan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Robust Optical and SAR Image Matching Using Attention-Enhanced Structural FeaturesabstractDue to the complementary nature of optical and SAR images, their alignment is of increasing interest. However, due to the significant radiometric differences between them, precise matching becomes a very challenging problem. Although current advanced structural features and deep learning-based methods have proposed feasible solutions, there is still much potential for improvement. In this paper, we propose a hybrid matching method using attention-enhanced structural features (namely AESF), which combines the advantages of both handcrafted-based and learning-based methods to improve the accuracy of optical and SAR image matching. It mainly consists of two modules: a novel effective multi-branch global attention (MBGA) module and a joint multi-cropping image matching loss function (MCTM) module. The MBGA module is designed to focus on shared information in structural feature descriptors of heterogeneous images across space and channel dimensions, significantly improving the expressive capacity of the classical structural features and generating more refined and robust image features. The MCTM module is constructed to fully exploit the association between global and local information of the input image, which can optimize the triple loss discriminator to discriminate positive and negative samples. To validate the effectiveness of the proposed method, it is compared with five state-of-the-art matching methods by using various optical and SAR datasets. The experimental results show that the matching accuracy at the 1-pixel threshold is improved by about 1.8%-8.7% compared with the most advanced deep learning method (OSMNet) and 6.5%-23% compared with the handcrafted description method (CFOG). Yuanxin Ye, Chao Yang 0028, Guoqing Gong, Peizhen Yang, Dou Quan, Jiayuan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Unified Deep Learning Network for Remote Sensing Image Registration and Change DetectionabstractImage registration and change detection are crucial for multitemporal remote sensing image analysis. The images should be registered before the change information detection. Existing deep learning methods have shown significant advantages in image registration and change detection tasks. They usually design two independent task-specific deep networks for image registration and change detection, respectively. These independent deep networks will learn from scratch and rely on many task-specific labeled training datasets. This article finds that image registration and change detection have similar learning mechanisms, which focus on extracting discriminative features. Inspired by this, we propose a Unified image Registration and Change detection Network (URCNet) that can perform image alignment and change information detection through a single network. Additionally, this article proposes various deep collaborative learning methods for URCNet optimization, which enforce that the URCNet can effectively support remote sensing image registration and change detection simultaneously. Extensive experiments demonstrate the effectiveness of the proposed URCNet for image registration and change detection, which can achieve comparable and better results with task-specific and more complex deep networks. The proposed URCNet can support multitasks based on the same scene images, different scene images, and even multimodal images. Moreover, URCNet shows significant advantages over other deep networks in change detection under limited labeled datasets. Rufan Zhou, Dou Quan, Shuang Wang 0001, Chonghua Lv, Xianwei Cao, Jocelyn Chanussot, Yi Li 0054, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | AUD-Net: A Unified Deep Detector for Multiple Hyperspectral Image Anomaly Detection via Relation and Few-Shot LearningabstractThis article addresses the problem of the building an out-of-the-box deep detector, motivated by the need to perform anomaly detection across multiple hyperspectral images (HSIs) without repeated training. To solve this challenging task, we propose a unified detector [anomaly detection network (AUD-Net)] inspired by few-shot learning. The crucial issues solved by AUD-Net include: how to improve the generalization of the model on various HSIs that contain different categories of land cover; and how to unify the different spectral sizes between HSIs. To achieve this, we first build a series of subtasks to classify the relations between the center and its surroundings in the dual window. Through relation learning, AUD-Net can be more easily generalized to unseen HSIs, as the relations of the pixel pairs are shared among different HSIs. Secondly, to handle different HSIs with various spectral sizes, we propose a pooling layer based on the vector of local aggregated descriptors, which maps the variable-sized features to the same space and acquires the fixed-sized relation embeddings. To determine whether the center of the dual window is an anomaly, we build a memory model by the transformer, which integrates the contextual relation embeddings in the dual window and estimates the relation embeddings of the center. By computing the feature difference between the estimated relation embeddings of the centers and the corresponding real ones, the centers with large differences will be detected as anomalies, as they are more difficult to be estimated by the corresponding surroundings. Extensive experiments on both the simulation dataset and 13 real HSIs demonstrate that this proposed AUD-Net has strong generalization for various HSIs and achieves significant advantages over the specific-trained detectors for each HSI. Ning Huyan, Xiangrong Zhang, Dou Quan, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Concurrent Multiscale Detector for End-to-End Image MatchingabstractThis article focuses on end-to-end image matching through joint key-point detection and descriptor extraction. To find repeatable and high discrimination key points, we improve the deep matching network from the perspectives of network structure and network optimization. First, we propose a concurrent multiscale detector (CS-det) network, which consists of several parallel convolutional networks to extract multiscale features and multilevel discriminative information for key-point detection. Moreover, we introduce an attention module to fuse the response maps of various features adaptively. Importantly, we propose two novel rank consistent losses (RC-losses) for network optimization, significantly improving image matching performances. On the one hand, we propose a score rank consistent loss (RC-S-loss) to ensure that the key points have high repeatability. Different from the score difference loss merely focusing on the absolute score of an individual key point, our proposed RC-S-loss pays more attention to the relative score of key points in the image. On the other hand, we propose a score-discrimination RC-loss to ensure that the key point has high discrimination, which can reduce the confusion from other key points in subsequent matching and then further enhance the accuracy of image matching. Extensive experimental results demonstrate that the proposed CS-det improves the mean matching result of deep detector by 1.4%-2.1%, and the proposed RC-losses can boost the matching performances by 2.7%-3.4% than score difference loss. Our source codes are available at https://github.com/iquandou/CS-Net. Dou Quan, Shuang Wang 0001, Ning Huyan, Yi Li 0054, Ruiqi Lei, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Select, Purify, and Exchange: A Multisource Unsupervised Domain Adaptation Method for Building ExtractionabstractAccurately extracting buildings from aerial images has essential research significance for timely understanding human intervention on the land. The distribution discrepancies between diversified unlabeled remote sensing images (changes in imaging sensor, location, and environment) and labeled historical images significantly degrade the generalization performance of deep learning algorithms. Unsupervised domain adaptation (UDA) algorithms have recently been proposed to eliminate the distribution discrepancies without re-annotating training data for new domains. Nevertheless, due to the limited information provided by a single-source domain, single-source UDA (SSUDA) is not an optimal choice when multitemporal and multiregion remote sensing images are available. We propose a multisource UDA (MSUDA) framework SPENet for building extraction, aiming at selecting, purifying, and exchanging information from multisource domains to better adapt the model to the target domain. Specifically, the framework effectively utilizes richer knowledge by extracting target-relevant information from multiple-source domains, purifying target domain information with low-level features of buildings, and exchanging target domain information in an interactive learning manner. Extensive experiments and ablation studies constructed on 12 city datasets prove the effectiveness of our method against existing state-of-the-art methods, e.g., our method achieves 59.1% intersection over union (IoU) on Austin and Kitsap → Potsdam, which surpasses the target domain supervised method by 2.2%. The code is available at https://github.com/QZangXDU/SPENet. Shuang Wang 0001, Qi Zang, Dong Zhao 0007, Chaowei Fang, Dou Quan, Yutong Wan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic SegmentationabstractUnsupervised domain adaptation (UDA) in semantic segmentation transfers the knowledge of the source domain to the target one to improve the adaptability of the segmentation model in the target domain. The need to access labeled source data makes UDA unable to handle adaptation scenarios involving privacy, property rights protection, and confidentiality. In this paper, we focus on unsupervised model adaptation (UMA), also called source-free domain adaptation, which adapts a source-trained model to the target domain without accessing source data. We find that the online self-training method has the potential to be deployed in UMA, but the lack of source domain loss will greatly weaken the stability and adaptability of the method. We analyze two reasons for the degradation of online self-training, i.e. inopportune updates of the teacher model and biased knowledge from the source-trained model. Based on this, we propose a dynamic teacher update mechanism and a training-consistency based resampling strategy to improve the stability and adaptability of online self-training. On multiple model adaptation benchmarks, our method obtains new state-of-the-art performance, which is comparable or even better than state-of-the-art UDA methods. The code is available at https://github.com/DZhaoXd/DT-ST. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Licheng Jiao |
CVPR | 4 |
| 2023 | Learning Pseudo-Relations for Cross-domain Semantic SegmentationabstractDomain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for category learning, while ignoring the useful relations between pixels for relation learning. In this paper, we propose a pseudo-relation learning framework, Relation Teacher (RTea), which can exploitable pixel relations to efficiently use unreliable pixels and learn generalized representations. In this framework, we build reasonable pseudo-relations on local grids and fuse them with low-level relations in the image space, which are motivated by the reliable local relations prior and available low-level relations prior. Then, we design a pseudo-relation learning strategy and optimize the class probability to meet the relation consistency by finding the optimal sub-graph division. In this way, the model’s certainty and consistency of prediction are enhanced on the target domain, and the cross-domain inadaptation is further eliminated. Extensive experiments on three datasets demonstrate the effectiveness of the proposed method. The code will be available at https://github.com/DZhaoXd/RTea. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Rui Yang 0038, Licheng Jiao |
ICCV | 4 |
| 2023 | Relational Image Patch Matching for Remote SensingabstractFeature descriptor-based methods have demonstrated remarkable performance in remote sensing image patch matching tasks and are usually optimized using contrastive loss and triplet loss. However, these optimization losses focus on calculating the distance between samples, ignoring the rich information of higher-order feature relationships between multiple image patches. The latter provides valuable information that can be used to improve task performance. Inspired by the superior performance of second-order relations in graph matching and clustering tasks, we aim to exploit the rich information available from high-order relations fully. This paper proposes a high-order relationship (HOR) learning method for remote sensing image patch matching. This method combines low-order feature relations between image patch pairs and high-order feature relations between multiple patches to enhance image matching performance. Extensive experimental results on a multimodel remote sensing image dataset, SEN 1-2, consisting of optical and SAR images, demonstrate that the proposed HOR learning method can improve the performance of remote sensing image patch matching. Xianwei Cao, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 2 |
| 2023 | A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge DistillationabstractWith the increasing development of remote sensing (RS) technology, remote sensing cross-modal image-text retrieval (RSCMITR) task has gradually attracted wide attention. At present, the large-scale pre-training model is brilliant in the field of natural images cross-modal retrieval, but the current RSCMITR models do not focus on it, resulting in less retrieval performance improvement. This paper proposes a lightweight network structure based on large-scale pre-training model and knowledge distillation, designing a lightweight model based on separable convolution and text convolution. Knowledge distillation technology is used to make the Light model learn the hidden knowledge of large-scale model CLIP-RS, which realizes fast and accurate retrieval. The proposed method achieves state-of-the-art performance on four commonly used RSCMITR datasets. Yu Liao, Rui Yang 0038, Hantong Xing, Dou Quan, Shuang Wang 0001, Biao Hou |
IGARSS | 5 |
| 2023 | Domain Distribution Alignment for Boosting Multi-Modal Remote Sensing Image MatchingabstractMulti-modal images can obtain complementary and rich information images, which are more widely used in various applications. However, due to the different imaging mechanisms of different sensors, there are significant domain distribution differences between multi-modal images. In multi-modal image matching, existing deep learning methods should deal with the image content difference caused by rotation transformation and the domain distribution difference caused by different sensors, which are very difficult for the deep network. To address this issue, we propose to combine an instance comparison and a batch comparison to deal with image content differences and domain distribution differences, respectively. We design a new domain distribution alignment method to explicitly constrain the sample domain distribution of the multi-modal images are consistent through the domain distribution alignment loss. Extensive multi-modal remote sensing image patch matching experiments have shown the effectiveness of the proposed method. Furthermore, the proposed multi-modal domain distribution alignment method has more obvious advantages when there are significant content differences and distribution differences. Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 2 |
| 2023 | Deep Continuous Matching Network for more Robust Multi-Modal Remote Sensing Image Patch MatchingabstractDue to the powerful feature extraction capabilities of deep neural networks, traditional approaches are gradually replaced by deep learning approaches for image matching tasks. For multi-modal image patch matching, the deep model should mainly learn the modality-invariant features. For multi-modal images with rotation transformation (RT), the deep model should learn the modality-invariant features and rotation-invariant features simultaneously. However, the performance of the latter trained model is degraded for the former task. The main reason is that the modality invariance of the features degenerates. This paper proposes a deep multi-modal remote sensing image matching network (DCMNet) that combines descriptor learning and continuous learning to solve this problem. Firstly, DCMNet is trained for learning modality-invariant features in multi-modal image patch matching. Then, DCMNet is optimized for multi-modal image patch matching with RT. In the later learning process, we reduce the change of important parameters for the modality-invariant features learning. Experiments demonstrate the effectiveness and robustness of DCMNet in alleviating the modal invariance degradation problem of features. Rufan Zhou, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 2 |
| 2023 | A Novel Coarse-to-Fine Deep Learning Registration Framework for Multimodal Remote Sensing ImagesabstractMulti-modal remote sensing images with large rotation transformation (RT) are challenging to be registered. It needs to deal with the global geometric deformation caused by great RT and significant local appearance differences caused by different imaging mechanisms. Existing deep learning methods mainly use a single deep descriptor learning (DDL) network to extract invariant features for identifying matching samples and discriminative feature descriptors for separating non-matching samples. However, it is difficult to extract local invariant feature descriptors to RT and modality change through a single DDL network. This paper proposes a novel coarse-to-fine deep learning image registration framework for multi-modal remote sensing images based on two task-specific deep models. Specifically, in the coarse registration stage, this paper designs an effective deep ordinal regression (DOR) network for rotation correction, which can reduce the difficulty of multi-modal image registration and boost image registration. The proposed DOR network transforms the rotation correction task into a rotation ordinal regression problem, which can exploit the potential relationship between the rotation ordinals to improve the accuracy of rotation estimation. In the fine registration stage, we adopt the DDL network to deal with the image modality change based on the rotation-corrected images. Extensive experimental results on multi-modal image datasets demonstrate the significant advantages of the proposed coarse-to-fine deep learning registration framework. The DOR network achieves higher rotation correction accuracy, which can significantly improve the multi-modal image registration performances. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Yu Gu 0015, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Deep Modality Independent Descriptor Learning for Optical and SAR Image Patch MatchingabstractDue to the complementary information between multi-modal images, they are widely used in various applications. However, there are significant differences in appearance caused by different imaging mechanisms, which bring great challenges to multi-modal image patch matching. To solve this problem, this paper proposes a deep modality independent descriptor learning network (DMID-Net) for multi-modal image patch matching. DMID-Net computes the self-similarity of deep features as the structure descriptor for image patch matching, which is independent of image modality and shared between multi-modal images. Thus, the acquired deep modality independent descriptor(DMID) can reduce the influence of significant differences between multi-modal images, further improving the matching performances. Experimental results on a large number of optical and SAR image-pairs demonstrate the effectiveness of DMID-Net on multi-modal image patch matching. Huiyuan Wei, Dou Quan, Ruiqi Lei, Baorui Duan, Shuang Wang 0001, Yi Li 0054, Biao Hou, Licheng Jiao |
IGARSS | 2 |
| 2022 | MANet: Multi-Scale Aware-Relation Network for Semantic Segmentation in Aerial ScenesabstractSemantic segmentation is an important yet unsolved problem in aerial scenes understanding. One of the major challenges is the intense variations of scenes and object scales. In this paper, we propose a novel multi-scale aware-relation network (MANet) to tackle this problem in remote sensing. Inspired by the process of human perception of multi-scale information, we explore discriminative and diverse multi-scale representations. For discriminative multi-scale representations, we propose an inter-class and intra-class region refinement method (IIRR) to reduce feature redundancy caused by fusion. IIRR utilizes the refinement maps with intra- and inter-class scale variation to guide multi-scale fine-grained features. Then, we propose multi-scale collaborative learning (MCL) to enhance the diversity of multi-scale feature representations. The MCL constrains the diversity of multi-scale feature network parameters to obtain diverse information. And the segmentation results are rectified according to the dispersion of the multi-level network predictions. In this way, MANet can learn multi-scale features by collaboratively exploiting the correlation among different scales. Extensive experiments on image and video datasets which have large scale variations have demonstrated the effectiveness of our proposed MANet. Pei He, Licheng Jiao, Ronghua Shang, Shuang Wang 0001, Xu Liu 0006, Dou Quan, Dong Zhao 0007 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Cluster-Memory Augmented Deep Autoencoder via Optimal Transportation for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection aims to detect objects significantly different from their surrounding background. Recently, many detectors based on autoencoder (AE) exhibited promising performances in hyperspectral anomaly detection tasks. However, the fundamental hypothesis of the AE-based detector that anomaly is more challenging to be reconstructed than background may not always be true in practice. We demonstrate that an autoencoder could well reconstruct anomalies even without anomalies for training. Because AE models mainly focus on the quality of sample reconstruction and do not care if the encoded features solely represent the background rather than anomalies. If more information is preserved than needed to reconstruct the background, the anomalies will be well reconstructed. This paper proposes a cluster-memory augmented autoencoder via deep optimal transportation clustering (OTCMA) for hyperspectral anomaly detection to solve this problem. The deep clustering method based on optimal transportation is proposed to enhance the features consistency of samples within the same categories and features discrimination of samples in different categories. The memory module stores the background’s consistent features, which are the cluster centers for each category background. We retrieve more consistent features from the memory module instead of reconstructing a sample utilizing its own encoded features. The network focuses more on consistent feature reconstruction by training AE with a memory module. This effectively restricts the reconstruction ability of AE and prevents reconstructing anomalies. Extensive experiments on the benchmark datasets demonstrate that our proposed OTCMA achieves state-of-the-art results. Besides, this paper presents further discussions about the effectiveness of our proposed memory module and different criterion for better anomaly detection. Ning Huyan, Xiangrong Zhang, Dou Quan, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deep Feature Correlation Learning for Multi-Modal Remote Sensing Image RegistrationabstractDeep descriptors have advantages over handcrafted descriptors on local image patch matching. However, due to the complex imaging mechanism of remote sensing images and the significant differences in appearance between multi-modal images, existing deep learning descriptors are unsuitable for multi-modal remote sensing image registration directly. To solve this problem, this paper proposes a deep feature correlation learning network (Cnet) for multi-modal remote sensing image registration. Firstly, Cnet builds a feature learning network based on the deep convolutional network with the attention learning module, to enhance the feature representation by focusing on meaningful features. Secondly, this paper designs a novel feature correlation loss function for Cnet optimization. It focuses on the relative feature correlation between matching and non-matching samples, which can improve the stability of network training and decrease the risk of overfitting. Additionally, the proposed feature correlation loss with a scale factor can further enhance the network training and accelerate the network convergence. Extensive experimental results on image patch matching (Brown, HPatches), cross-spectral image registration (VIS-NIR), multi-modal remote sensing image registration, and single-modal remote sensing image registration have demonstrated the effectiveness and robustness of the proposed method. Dou Quan, Shuang Wang 0001, Yu Gu 0015, Ruiqi Lei, Bowu Yang, Shaowei Wei, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Distillation Feature Learning Network for Optical and SAR Image RegistrationabstractOptical and SAR image registration is important for multi-modal remote sensing image information fusion. Recently, deep matching networks have shown better performances than traditional methods on image matching. However, due to significant differences between optical and SAR images, the performances of existing deep learning methods still need to be further improved. This paper proposes a self-distillation feature learning network (SDNet) for optical and SAR image registration, improving performance from network structure and network optimization. Firstly, we explore the impact of different weight-sharing strategies on optical and SAR image matching. Then, we design a partially unshared feature learning network for multi-modal image feature learning. It has fewer parameters than the fully unshared network and has more flexibility than the fully shared network. Additionally, the limited binary supervised information (matching or non-matching) is insufficient to train the deep matching networks for optical-SAR image registration. Thus, we propose a self-distillation feature learning method to exploit more similarity information for deep network optimization enhancing, such as the similarity ordering between a series of non-matching patch-pairs. The exploited rich similarity information will significantly enhance network training and improve matching accuracy. Finally, considering that existing deep learning methods brute-force constrain the features of the matching optical and SAR image patches are similar, which will be lost many discriminative information, degenerating matching performances. Thus, we build an auxiliary task reconstruction learning to optimize the feature learning network to keep more discriminative information. Extensive experiments demonstrate the effectiveness of our proposed method on multi-modal image registration. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Ruiqi Lei, Baorui Duan, Yi Li 0054, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Unsupervised Outlier Detection Using Memory and Contrastive LearningabstractOutlier detection is to separate anomalous data from inliers in the dataset. Recently, the most deep learning methods of outlier detection leverage an auxiliary reconstruction task by assuming that outliers are more difficult to recover than normal samples (inliers). However, it is not always true in deep auto-encoder (AE) based models. The auto-encoder based detectors may recover certain outliers even if outliers are not in the training data, because they do not constrain the feature learning. Instead, we think outlier detection can be done in the feature space by measuring the distance between outliers' features and the consistency feature of inliers. To achieve this, we propose an unsupervised outlier detection method using a memory module and a contrastive learning module (MCOD). The memory module constrains the consistency of features, which merely represent the normal data. The contrastive learning module learns more discriminative features, which boosts the distinction between outliers and inliers. Extensive experiments on four benchmark datasets show that our proposed MCOD performs well and outperforms eleven state-of-the-art methods. Ning Huyan, Dou Quan, Xiangrong Zhang, Xuefeng Liang, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Image Process. | 2 |
| 2022 | Element-Wise Feature Relation Learning Network for Cross-Spectral Image Patch MatchingabstractRecently, the majority of successful matching approaches are based on convolutional neural networks, which focus on learning the invariant and discriminative features for individual image patches based on image content. However, the image patch matching task is essentially to predict the matching relationship of patch pairs, that is, matching (similar) or non-matching (dissimilar). Therefore, we consider that the feature relation (FR) learning is more important than individual feature learning for image patch matching problem. Motivated by this, we propose an element-wise FR learning network for image patch matching, which transforms the image patch matching task into an image relationship-based pattern classification problem and dramatically improves generalization performances on image matching. Meanwhile, the proposed element-wise learning methods encourage full interaction between feature information and can naturally learn FR. Moreover, we propose to aggregate FR from multilevels, which integrates the multiscale FR for more precise matching. Experimental results demonstrate that our proposal achieves superior performances on cross-spectral image patch matching and single spectral image patch matching, and good generalization on image patch retrieval. Dou Quan, Shuang Wang 0001, Ning Huyan, Jocelyn Chanussot, Ruojing Wang, Xuefeng Liang, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | A Feature Decomposition Framework for Multi-Modal Image Patch MatchingabstractMulti-modal remote sensing images have complementary information which is conducive to enhancing the performance of various applications. Image patch matching plays a crucial role in the combination of multi-modal images. However, there are great differences in appearance and texture of multi-modal images, which brings great difficulties to image patching matching. To solve this problem, we propose a novel feature decomposition framework for multi-modal image patch matching. It aims to eliminate the hinder caused by the significant difference in multi-modal images. Specifically, this paper proposes to decompose the feature of images into common feature and modal private feature. Then, only the common feature is used for image patch matching, so as to improve the matching accuracy. Experimental results on optical and SAR images demonstrate that our proposed feature decomposition framework can significantly improve the performance of multi-modal image patch matching. Baorui Duan, Dou Quan, Yi Li 0054, Ruiqi Lei, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 2 |
| 2021 | Deep Global Feature-Based Template Matching for Fast Multi-Modal Image RegistrationabstractDue to the different imaging mechanisms, there is a significant non-line difference between multi-modal images, which brings difficulties to multi-modal image registration. The traditional methods based on grayscale and handcraft features are difficult with obtain common features between different source images. The performances of deep local features matching methods rely on the quality and quantity of the detected keypoints, which can be quite time-consuming to register images. To achieve fast and accurate multi-modal image registration, we propose a deep global feature-based template matching method (GFTM) which uses a deep convolutional network to extract common global deep features from multi-modal images. Then, fast template matching is performed on global deep features to search the position with maximal similarity. Additionally, we build a similarity label map and design three losses to optimize our network, including contrast loss, error loss and peak loss. Extensive experimental results on optical and SAR images demonstrated that our proposed method is effective on multi-modal image registration. Ruiqi Lei, Bowu Yang, Dou Quan, Yi Li 0054, Baorui Duan, Shuang Wang 0001, Huarong Jia, Biao Hou, Licheng Jiao |
IGARSS | 3 |
| 2021 | Multi-Relation Attention Network for Image Patch MatchingabstractDeep convolutional neural networks attract increasing attention in image patch matching. However, most of them rely on a single similarity learning model, such as feature distance and the correlation of concatenated features. Their performances will degenerate due to the complex relation between matching patches caused by various imagery changes. To tackle this challenge, we propose a multi-relation attention learning network (MRAN) for image patch matching. Specifically, we propose to fuse multiple feature relations (MR) for matching, which can benefit from the complementary advantages between different feature relations and achieve significant improvements on matching tasks. Furthermore, we propose a relation attention learning module to learn the fused relation adaptively. With this module, meaningful feature relations are emphasized and the others are suppressed. Extensive experiments show that our MRAN achieves best matching performances, and has good generalization on multi-modal image patch matching, multi-modal remote sensing image patch matching and image retrieval tasks. Dou Quan, Shuang Wang 0001, Yi Li 0054, Bowu Yang, Ning Huyan, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Image Process. | 1 |
| 2019 | AFD-Net: Aggregated Feature Difference Learning for Cross-Spectral Image Patch MatchingabstractImage patch matching across different spectral domains is more challenging than in a single spectral domain. We consider the reason is twofold: 1. the weaker discriminative feature learned by conventional methods; 2. the significant appearance difference between two images domains. To tackle these problems, we propose an aggregated feature difference learning network (AFD-Net). Unlike other methods that merely rely on the high-level features, we find the feature differences in other levels also provide useful learning information. Thus, the multi-level feature differences are aggregated to enhance the discrimination. To make features invariant across different domains, we introduce a domain invariant feature extraction network based on instance normalization (IN). In order to optimize the AFD-Net, we borrow the large margin cosine loss which can minimize intra-class distance and maximize inter-class distance between matching and non-matching samples. Extensive experiments show that AFD-Net largely outperforms the state-of-the-arts on the cross-spectral dataset, meanwhile, demonstrates a considerable generalizability on a single spectral dataset. Dou Quan, Xuefeng Liang, Shuang Wang 0001, Shaowei Wei, Ning Huyan, Licheng Jiao |
ICCV | 1 |
| 2019 | Better and Faster: Exponential Loss for Image Patch MatchingabstractRecent studies on image patch matching are paying more attention on hard sample learning, because easy samples do not contribute much to the network optimization. They have proposed various hard negative sample mining strategies, but very few addressed this problem from the perspective of loss functions. Our research shows that the conventional Siamese and triplet losses treat all samples linearly, thus make the training time consuming. Instead, we propose the exponential Siamese and triplet losses, which can naturally focus more on hard samples and put less emphasis on easy ones, meanwhile, speed up the optimization. To assist the exponential losses, we introduce the hard positive sample mining to further enhance the effectiveness. The extensive experiments demonstrate our proposal improves both metric and descriptor learning on several well accepted benchmarks, and outperforms the state-of-the-arts on the UBC dataset. Moreover, it also shows a better generalizability on cross-spectral image matching and image retrieval tasks. Shuang Wang 0001, Xuefeng Liang, Dou Quan, Bowu Yang, Shaowei Wei, Licheng Jiao |
ICCV | 4 |
| 2019 | An Improved Fully Convolutional Network for Learning Rich Building FeaturesabstractMany efficient approaches are proposed to detect building in remote sensing images. In this paper, in order to learning rich building features better, we propose a full convolutional network with dense connection. There contributions are made: 1) To strengthen feature propagation, an improved dense network is introduced to the full convolution network. 2) We have designed top-down short connections to facilitate the fusion of high and low feature information. 3) In addition, we add the weighted cross entropy edge loss function to make the network pay more attention to building edge in detail. Experiments show that the proposed method achieves excellent performance on the remote sensing image data taken by the QuickBird satellite. Shuang Wang 0001, Pei He, Dou Quan, Xuefeng Liang, Biao Hou |
IGARSS | 4 |
| 2018 | Cross-Spectral Image Patch Matching by Learning Features of the Spatially Connected Patches in a Shared Space
Dou Quan, Shuai Fang, Xuefeng Liang, Shuang Wang 0001, Licheng Jiao |
ACCV (2) | 1 |
| 2018 | A Two-Branch Network with Semi-Supervised Learning for Hyperspectral ClassificationabstractIn order to promote progress on fusion and analysis methodologies for multi-source remote sensing data, The Image Analysis and Data Fusion Technical Committee organized the 2018 IEEE GRSS Data Fusion contest. In this contest, we proposed a two-branch convolution network for hyperspectral image classification with a data re-sampling strategy and semi-supervised learning to address three existing problems, i.e. multi-scale feature learning, data imbalance, and small size of the dataset. The contest showed that our proposal achieved the best performance on two metrics: the overall accuracy of 77.39% and a kappa coefficient of 0.76 on the hyperspectral images provided by 2018 IEEE GRSS Data Fusion Contest. Shuai Fang, Dou Quan, Shuang Wang 0001 |
IGARSS | 2 |
| 2018 | Deep Generative Matching Network for Optical and SAR Image RegistrationabstractMultimodal remote sensing images contain complementary information, thus, could potentially benefit many remote sensing applications. To this end, the image registration is a common requirement for utilizing the multimodal images. However, due to the rather different imaging mechanisms, multimodal image registration becomes much more challenging than ordinary registration, particular for optical and synthetic aperture radar (SAR) images. In this work, we design a deep matching network to exploit the latent and coherent features between multimodal patch pairs for inferring their matching labels. But, the network requires immense data for training, which is not usually met. To address this issue, we propose a generative matching network (GMN) to generate the coupled optical and SAR images, hence, improve the quantity and diversity of the training data. The experimental results show that our proposal significantly improves the registration performance of optical and SAR image registration, and achieves subpixel or close to subpixel error. Dou Quan, Shuang Wang 0001, Xuefeng Liang, Ruojing Wang, Shuai Fang, Biao Hou, Licheng Jiao |
IGARSS | 1 |
| 2016 | Using deep neural networks for synthetic aperture radar image registrationabstractAt present, the performance of image registration mainly depends on the extracted features in feature-based image registration. However, due to the speckle noise, synthetic aperture radar (SAR) image registration will have a lower accuracy and less robustness. For this purpose, we design a deep neural network (DNN) for SAR image registration, using the DNN to learn the image features, automatically. The deep learning could learn the more essential features of the images, which are make the image registration to achieve more robust features and accurate matching. Moreover, this paper proposed a new strategy to remove the wrong matching points based on the RANSAC. The experimental results on SAR image registration show that this image registration method based on DNN have a better performance, and the new RANSAC strategy could eliminate many wrong matching points and get a good transformational model. Dou Quan, Shuang Wang 0001, Mengdan Ning, Licheng Jiao |
IGARSS | 1 |