EDBT 2026 Demo / reviewers in the wild / expert
Yunpeng Liu 0001
dblp:02/8137-1
· DBLP profile ↗
28ranked-venue papers
0as first author
24since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Physics prior adapter tuning for thermal infrared tracking
Jingyuan Guo, Qiao Liu 0001, Kanlun Tan, Di Yuan 0002, Yunpeng Liu 0001 |
Expert Syst. Appl. | 5 |
| 2026 | SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place RecognitionabstractRecent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a novel method to realize seamless adaptation of foundation models to VPR (SelaVPR). This method can produce both global and local features that focus on discriminative landmarks to recognize places for two-stage VPR by a parameter-efficient adaptation approach. Although SelaVPR has achieved competitive results, we argue that the previous adaptation is inefficient in training time and GPU memory usage, and the re-ranking paradigm is also costly in retrieval latency and storage usage. In pursuit of higher efficiency and better performance, we propose an extension of the SelaVPR, called SelaVPR++. Concretely, we first design a parameter-, time-, and memory-efficient adaptation method that uses lightweight multi-scale convolution (MultiConv) adapters to refine intermediate features from the frozen foundation backbone. This adaptation method does not back-propagate gradients through the backbone during training, and the MultiConv adapter facilitates feature interactions along the spatial axes and introduces proper local priors, thus achieving higher efficiency and better performance. Moreover, we propose an innovative re-ranking paradigm for more efficient VPR. Instead of relying on local features for re-ranking, which incurs huge overhead in latency and storage, we employ compact binary features for initial retrieval and robust floating-point (global) features for re-ranking. To obtain such binary features, we propose a similarity-constrained deep hashing method, which can be easily integrated into the VPR pipeline. Finally, we improve our training strategy and unify the training protocol of several common training datasets to merge them for better training of VPR models. Extensive experiments show that SelaVPR++ is highly efficient in training time, GPU memory usage, and retrieval latency (6000× faster than TransVPR), as well as outperforms the state-of-the-art methods by a large margin (ranks 1st on MSLS challenge leaderboard). Xiangyuan Lan, Yunpeng Liu 0001, Yaowei Wang 0001, Chun Yuan 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | PPIFuse: Physical Priors Injected Infrared and Visible Image FusionabstractExisting infrared and visible image fusion methods commonly use two structurally identical networks to extract deep features from source images, followed by a handcrafted or learnable feature fusion strategy. These methods overlook the modality-specific characteristics of the two image types, impairing the model’s ability to fully exploit their complementary information. Additionally, their fusion results often exhibit issues such as texture detail loss or unclear thermal targets. This is because the fusion rules they used are either too simple or too redundant. To address these challenges, we start from the infrared physics priors that are naturally complementary to visible images and incorporate the thermal diffusion equation and Stefan-Boltzmann Law into the image fusion architecture. Based on these two physical priors, we design a Thermal Diffusion Convolution (TDC) and a Stefan Thermal Attention (STA) to better extract infrared-specific features. Specifically, the TDC module leverages the anisotropic and isotropic characteristics of thermal diffusion adaptively to sharpen the edges of thermal targets and remove infrared noise, minimizing artifacts in the fused results. By decomposing the Stefan-Boltzmann Law, STA pays more attention on thermal features while suppressing redundant information, enabling more effective aggregation of complementary modality-specific details. To make full use of layer-wise complementary features, we propose an Interactive Injection Fusion framework(IIF) that hierarchically integrates these features, enhancing the richness of fused image content. Furthermore, an energy conservation constraint is designed to ensure the fused images adhere to physical principles. Extensive experimental results on five datasets demonstrate that our method sets a new state-of-the-art. Code is available at https://github.com/QiaoLiuHit/PPIFuse. Qianhong Zhang, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Why and How: Knowledge-Guided Learning for Cross-Spectral Image Patch MatchingabstractRecently, cross-spectral image patch matching based on feature relation learning has attracted extensive attention. However, existing methods focus on mining richer feature relations by building complex relation extraction structures. Meanwhile, performance bottlenecks have gradually emerged. To address this, we make the first attempt to explore a stable and efficient bridge between descriptor learning and metric learning, and construct a Knowledge-Guided Learning Network (KGL-Net), which achieves significant performance improvements while abandoning complex network structures. Specifically, we find that there is feature extraction consistency between metric learning based on feature difference learning and descriptor learning based on Euclidean distance. This provides the foundation for bridge building. To ensure the stability and efficiency of the constructed bridge, on the one hand, we conduct an in-depth exploration of 20 combined network architectures. On the other hand, a feature-guided loss is constructed to achieve mutual guidance of features. In addition, unlike existing methods, we consider that the feature mapping ability of the metric branch should receive more attention. Therefore, a hard negative sample mining for metric learning (HNSM-M) strategy is constructed. To the best of our knowledge, this is the first time that hard negative sample mining for metric networks has been implemented and brings significant performance gains. Extensive experimental results show that our KGL-Net achieves SOTA performance in multiple cross-spectral image patch matching datasets. Our code is available at https://github.com/YuChuang1205/KGL-Net. Chuang Yu 0003, Yunpeng Liu 0001, Jinmiao Zhao, Ao Chen 0003, Xiujun Shu, Bo Wang 0162, Zelin Shi, Xiangyu Yue 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Unsupervised Domain Adaptive Thermal Infrared TrackingabstractExisting deep Thermal InfraRed (TIR) trackers often use RGB datasets for training due to the lack of large-scale labeled TIR datasets. However, the performance of these methods on TIR image sequences is significantly degraded, because of the domain shift problem between the RGB and TIR datasets. To solve this problem, in this paper, we propose an unsupervised Dual-level Domain Adaptation TIR Tracking framework (DDAT), which can benefit from training on large-scale labeled RGB datasets and unlabeled TIR datasets. Specifically, to transfer the useful knowledge learned from RGB dataset to TIR tracking, we first propose an adversarial-based adaptation module on both the semantic-level and the feature-level. While the semantic-level adaptation can reduce the semantic gap between the TIR and RGB tracking tasks, the feature-level adaptation can learn domain-invariant features for more robust tracking. Second, we propose a partial domain adaptation module to alleviate the negative transfer problem because the RGB and TIR tracking domains have a non-identical class and feature spaces. Instead of aligning the entire feature space, this module adaptively selects partial similarity samples and features for alignment, thus getting more fine-grained aligned results. Third, we collect a currently largest-scale unlabeled TIR dataset to train the proposed framework. Extensive experiments on five TIR tracking benchmarks demonstrate the proposed method is effective and sets a new state-of-the-art. Qiao Liu 0001, Xin Li 0034, Jiatian Pi, Di Yuan 0002, Yunpeng Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Efficient Hierarchical Domain Adaptive Thermal Infrared TrackingabstractConstrained by the scarcity of labeled Thermal InfraRed (TIR) training data, current TIR trackers commonly rely on pre-trained RGB trackers. However, the domain discrepancy between TIR and RGB images limits effective utilization of RGB features, significantly degrades TIR tracking performance. To solve this challenge, we propose a hierarchical domain adaptation model to transfer useful pre-trained RGB features into TIR tracking more effective and efficient. Specifically, we first design a reflectance consistency network to learn style-invariant representations. Second, we present a target-aware adversarial network to align the target semantic features of the two domains. These two modules respectively narrow the distribution gap at the stylistic and semantic levels in a hierarchical manner. Third, to solve the inefficiency problem of domain adaptive training, we also propose a Bi-rank adapter side network to accelerate this process. While significantly reducing training time by 90%, our method achieves a new state-of-the-art on four TIR tracking benchmarks. Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001 |
ICASSP | 6 |
| 2025 | From Easy to Hard: Progressive Active Learning Framework for Infrared Small Target Detection with Single Point SupervisionabstractRecently, single-frame infrared small target (SIRST) detection with single point supervision has drawn wide-spread attention. However, the latest label evolution with single point supervision (LESPS) framework suffers from instability, excessive label evolution, and difficulty in exerting embedded network performance. Inspired by organisms gradually adapting to their environment and continuously accumulating knowledge, we construct an innovative Progressive Active Learning (PAL) framework, which drives the existing SIRST detection networks progressively and actively recognizes and learns harder samples. Specifically, to avoid the early low-performance model leading to the wrong selection of hard samples, we propose a model pre-start concept, which focuses on automatically selecting a portion of easy samples and helping the model have basic task-specific learning capabilities. Meanwhile, we propose a refined dual-update strategy, which can promote reasonable learning of harder samples and continuous refinement of pseudo-labels. In addition, to alleviate the risk of excessive label evolution, a decay factor is reasonably introduced, which helps to achieve a dynamic balance between the expansion and contraction of target annotations. Extensive experiments show that existing SIRST detection networks equipped with our PAL framework have achieved state-of-the-art (SOTA) results on multiple public datasets. Furthermore, our PAL framework can build an efficient and stable bridge between full supervision and single point supervision tasks. Our code is available at https://github.com/YuChuang1205/PAL Chuang Yu 0003, Jinmiao Zhao, Yunpeng Liu 0001, Sicheng Zhao, Yimian Dai, Xiangyu Yue 0001 |
ICCV | 3 |
| 2025 | Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraabstractVisual place recognition (VPR) is typically regarded as a specific image retrieval task, whose core lies in representing images as global descriptors. Over the past decade, dominant VPR methods (e.g., NetVLAD) have followed a paradigm that first extracts the patch features/tokens of the input image using a backbone, and then aggregates these patch features into a global descriptor via an aggregator. This backbone-plus-aggregator paradigm has achieved overwhelming dominance in the CNN era and remains widely used in transformer-based models. In this paper, however, we argue that a dedicated aggregator is not necessary in the transformer era, that is, we can obtain robust global descriptors only with the backbone. Specifically, we introduce some learnable aggregation tokens, which are prepended to the patch tokens before a particular transformer block. All these tokens will be jointly processed and interact globally via the intrinsic self-attention mechanism, implicitly aggregating useful information within the patch tokens to the aggregation tokens. Finally, we only take these aggregation tokens from the last output tokens and concatenate them as the global representation. Although implicit aggregation can provide robust global descriptors in an extremely simple manner, where and how to insert additional tokens, as well as the initialization of tokens, remains an open issue worthy of further exploration. To this end, we also propose the optimal token insertion strategy and token initialization method derived from empirical studies. Experimental results show that our method outperforms state-of-the-art methods on several VPR datasets with higher efficiency and ranks 1st on the MSLS challenge leaderboard. The code is available at https://github.com/lu-feng/image. Canming Ye, Xiangyuan Lan, Yunpeng Liu 0001, Chun Yuan 0003 |
NeurIPS | 5 |
| 2025 | High-Precision Remote Sensing Image Change Detection Based on Image Style Unification and Feature Extraction Optimization
Jinmiao Zhao, Zelin Shi, Chuang Yu 0003, Yunpeng Liu 0001 |
PRCV (15) | 4 |
| 2025 | Alignment-assisted Frequency Fusion Network for RGB-infrared vehicle detection
Zhenshuai Chen, Zhiyuan Lin 0005, Yunpeng Liu 0001, Zelin Shi |
Neurocomputing | 5 |
| 2025 | D2Fusion: Dual-domain feature decoupling for infrared and visible image fusion
Yan Fan 0006, Wei Ran, Kanlun Tan, Qiao Liu 0001, Di Yuan 0002, Xin Li 0034, Yunpeng Liu 0001 |
Knowl. Based Syst. | 7 |
| 2025 | Towards robust infrared small target detection: A feature-enhanced and sensitivity-tunable framework
Jinmiao Zhao, Zelin Shi, Chuang Yu 0003, Yunpeng Liu 0001, Yimian Dai |
Knowl. Based Syst. | 4 |
| 2025 | EDTformer: An Efficient Decoder Transformer for Visual Place RecognitionabstractVisual place recognition (VPR) aims to determine the general geographical location of a query image by retrieving visually similar images from a large geo-tagged database. To obtain a global representation for each place image, most approaches typically focus on the aggregation of deep features extracted from a backbone through using current prominent architectures (e.g., CNNs, MLPs, pooling layer, and transformer encoder), giving little attention to the transformer decoder. However, we argue that its strong capability to capture contextual dependencies and generate accurate features holds considerable potential for the VPR task. To this end, we propose an Efficient Decoder Transformer (EDTformer) for feature aggregation, which consists of several stacked simplified decoder blocks followed by two linear layers to directly produce robust and discriminative global representations. Specifically, we do this by formulating deep features as the keys and values, as well as a set of learnable parameters as the queries. Our EDTformer can fully utilize the contextual information within deep features, then gradually decode and aggregate the effective features into the learnable queries to output the global representations. Moreover, to provide more powerful deep features for EDTformer and further facilitate the robustness, we use the foundation model DINOv2 as the backbone and propose a Low-rank Parallel Adaptation (LoPA) method to enhance its performance in VPR, which can refine the intermediate features of the backbone progressively in a memory- and parameter-efficient way. As a result, our method not only outperforms single-stage VPR methods on multiple benchmark datasets, but also outperforms two-stage VPR methods which add a re-ranking with considerable cost. Code will be available at https://github.com/Tong-Jin01/EDTformer. Shuyu Hu, Chun Yuan 0003, Yunpeng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Query-Driven Feature Learning for Cross-View Geo-LocalizationabstractThe cross-view geo-localization task aims to accurately retrieve a location using images captured from different platforms, such as satellites and drones, which is particularly challenging due to a large variation in viewpoint. Current methods mainly focus on rigid strategies like partitioning or sorting local features, which may be ill-suited to accommodate the variance of viewpoint and distance scale in different camera perspectives. To address these issues, we propose a novel method called query-driven feature learning (QDFL) to query viewpoint-invariant feature vectors autonomously. Our method incorporates an adaptive query embedding unit (AQEU) and a feature fusion unit (FFU). AQEU adjusts feature map and implements a coarse query process to extract the contextual clues. FFU further refines feature map, fusing it at spatial and channel dimensions, tending to withdraw more fine-grained features. Subsequently, AQEU executes a fine query on salient landmarks in the fused feature map, enhancing the minutia descriptive power of query vectors. Additionally, we employ parameter-efficient transfer learning manner by integrating tunable adapters into the frozen pre-trained backbone, maintaining feature representation capabilities of foundation models while enabling seamless adaptation to cross-view geo-localization task. Extensive experiments show that our method achieves state-of-the-art performances on two well-known datasets, University-1652 and SUES-200. Moreover, our method exhibits an excellent generalizability compared with current state-of-the-art methods in cross-dataset experiments. The code is available at https://github.com/Shuyu-Hu/QDFL. Shuyu Hu, Zelin Shi, Yunpeng Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multi-Scale Direction-Aware Network for Infrared Small Target DetectionabstractInfrared small target detection faces the problem that it is difficult to effectively separate the background and the target. Existing deep learning-based methods focus on edge and shape features, but ignore the richer structural differences and detailed information embedded in high-frequency components from different directions, thereby failing to fully exploit the value of high-frequency directional features in target perception. To address this limitation, we propose a multi-scale direction-aware network (MSDA-Net), which is the first attempt to integrate the high-frequency directional features of infrared small targets as domain prior knowledge into neural networks. Specifically, to fully mine the high-frequency directional features, on the one hand, a high-frequency direction injection (HFDI) module without trainable parameters is constructed to inject the high-frequency directional information of the original image into the network. On the other hand, a multi-scale direction-aware (MSDA) module is constructed, which promotes the full extraction of local relations at different scales and the full perception of key features in different directions. In addition, considering the characteristics of infrared small targets, we construct a feature aggregation (FA) structure to address target disappearance in high-level feature maps, and a feature calibration fusion (FCF) module to alleviate feature bias during cross-layer feature fusion. Extensive experimental results show that our MSDA-Net achieves state-of-the-art (SOTA) results on multiple public datasets. The code can be available at https://github.com/YuChuang1205/MSDA-Net. Jinmiao Zhao, Zelin Shi, Chuang Yu 0003, Yunpeng Liu 0001, Xinyi Ying, Yimian Dai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao |
ICPR (34) | 27 |
| 2024 | Hierarchical Interactive Learning Network for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is a challenging task due to the small size and lack of intrinsic features. Meanwhile, small targets in the infrared spectrum often exhibit low contrast, which makes them difficult to distinguish from complex backgrounds. To address these challenges, we propose a novel hierarchical interactive learning network (HIL-Net). Specifically, we design a hierarchical interactive module (HIM), which realizes the hierarchical interaction between low-level and high-level features. By using deep layer information to enhance the expression of lower level features, we are able to better capture the characteristics of small targets. In addition, we introduce the local area enhancement attention (LAEA), which employs feature decomposition and reconstruction to perform fine-grained local contrast calculations, fully utilizing local information and effectively addressing the challenge of low contrast for small infrared targets. Extensive experiments prove that HIL-Net achieves the state-of-the-art results on the public NUAA-SIRST, NUDT-SIRST, IRSTD-1k, and TDSATUA datasets, with the mean intersection over union (mIoU) scores of 77.74%, 88.92%, 68.20%, and 55.18%, respectively. Junling Liu, Yunpeng Liu 0001, Huanliang Sun |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Prototype Contrastive Learning for Building Extraction From Remote Sensing ImagesabstractDeep learning has contributed to the rapid development of building extraction tasks from remote sensing(RS) images. Existing models typically leverage a segmentation-head to predict results, where multi-channel feature maps extracted by the network are directly output as single-channel predictions. However, it is rarely noticed that this process results in a loss of features, which can lead to incomplete extraction of smaller buildings. Besides, boundary-blurring is also a common problem in the task. Therefore, in this letter, we propose a Siamese Prototype Contrastive Learning Network (SPCL-Net) to address these two problems. In the network, a novel Prototype Contrastive Learning (PCL) module is proposed to alleviate feature loss problem by applying contrastive learning between prototype vectors. In addition, a Reverse Boundary Enhancement (RBE) module is proposed to facilitate the representation of building boundaries and mitigate the boundary-blurring problem. Experiments are conducted on two datasets, INRIA and WHU. Compared with existing models, the final results show that the proposed approach is better than theirs in terms of evaluation metrics IOU. Zhenshuai Chen, Zhiyuan Lin 0005, Chuang Yu 0003, Yunpeng Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Gradient-Guided Learning Network for Infrared Small Target DetectionabstractRecently, infrared small target detection has attracted extensive attention. However, due to the small size and the lack of intrinsic features of infrared small targets, the existing methods generally have the problem of inaccurate edge positioning and the target is easily submerged by the background. Therefore, we propose an innovative gradient-guided learning network (GGL-Net). Specifically, we are the first to explore the introduction of gradient magnitude images into the deep learning-based infrared small target detection method, which is conducive to emphasizing the edge details and alleviating the problem of inaccurate edge positioning of small targets. On this basis, we propose a novel dual-branch feature extraction network that utilizes the proposed gradient supplementary module (GSM) to encode raw gradient information into deeper network layers and embeds attention mechanisms reasonably to enhance feature extraction ability. In addition, we construct a two-way guidance fusion module (TGFM), which fully considers the characteristics of feature maps at different levels. It can facilitate the effective fusion of multi-scale feature maps and extract richer semantic information and detailed information through reasonable two-way guidance. Extensive experiments prove that GGL-Net has achieves state-of-the-art results on the public real NUAA-SIRST dataset and the public synthetic NUDT-SIRST dataset. Jinmiao Zhao, Chuang Yu 0003, Zelin Shi, Yunpeng Liu 0001, Yingdi Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Efficient Feature Relation Learning Network for Cross-Spectral Image Patch MatchingabstractRecently, cross-spectral image patch matching methods based on feature difference aggregation have achieved excellent performance, but they introduce a large number of parameters, limit matching speed and have poor scalability. At the same time, only using feature difference learning to extract differential features will lead to the loss of consistent features between cross-spectral image patches. Therefore, we construct a novel four-branch efficient feature relation learning network (EFR-Net) without feature difference aggregation. Specifically, a new four-branch feature relation learning strategy is proposed, which reasonably combines multiple feature relation learning to comprehensively and effectively extract the differential features and consistent features between image patches. At the same time, we construct an efficient local attention (ELA) module with negligible parameters, which can learn some global context information, enhance the interaction of local information and promote the extraction of discriminative features. In addition, a combined metric network is introduced to facilitate network optimization and improve network generalization. Furthermore, a public optical and SAR image patch matching dataset with a patch size of 64 × 64 pixels is constructed based on the OS dataset, which is called the OS patch dataset. We also establish an experimental benchmark on this new dataset. Extensive experimental results show that the proposed EFR-Net achieves excellent performance on cross-spectral image patch matching (OS patch dataset, VIS-NIR patch dataset) and single spectral image patch matching (Brown dataset). Chuang Yu 0003, Jinmiao Zhao, Yunpeng Liu 0001, Shuhang Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Feature Interaction Learning Network for Cross-Spectral Image Patch MatchingabstractRecently, feature relation learning has attracted extensive attention in cross-spectral image patch matching. However, most feature relation learning methods can only extract shallow feature relations and are accompanied by the loss of useful discriminative features or the introduction of disturbing features. Although the latest multi-branch feature difference learning network can relatively sufficiently extract useful discriminative features, the multi-branch network structure it adopts has a large number of parameters. Therefore, we propose a novel two-branch feature interaction learning network (FIL-Net). Specifically, a novel feature interaction learning idea for cross-spectral image patch matching is proposed, and a new feature interaction learning module is constructed, which can effectively mine common and private features between cross-spectral image patches, and extract richer and deeper feature relations with invariance and discriminability. At the same time, we re-explore the feature extraction network for the cross-spectral image patch matching task, and a new two-branch residual feature extraction network with stronger feature extraction capabilities is constructed. In addition, we propose a new multi-loss strong-constrained optimization strategy, which can facilitate reasonable network optimization and efficient extraction of invariant and discriminative features. Furthermore, a public VIS-LWIR patch dataset and a public SEN1-2 patch dataset are constructed. At the same time, the corresponding experimental benchmarks are established, which are convenient for future research while solving few existing cross-spectral image patch matching datasets. Extensive experiments show that the proposed FIL-Net achieves state-of-the-art performance in three different cross-spectral image patch matching scenarios. Chuang Yu 0003, Yunpeng Liu 0001, Jinmiao Zhao, Shuhang Wu, Zhuhua Hu |
IEEE Trans. Image Process. | 2 |
| 2022 | Pay Attention to Local Contrast Learning Networks for Infrared Small Target DetectionabstractInfrared small target suffers from the lack of intrinsic features, context and samples. Conventional detection methods are usually unable to sufficiently and effectively extract the features of infrared small targets. Therefore, we propose a novel attention-based local contrast learning network (ALCL-Net). Considering the scarcity of intrinsic features of infrared small targets, we propose ResNet32, which enhances the ability to extract infrared small target features and avoids the problem that the target features are overwhelmed by the background features due to too deep network. At the same time, we construct a simplified bilinear interpolation attention module (SBAM), which is used for fusion of hierarchical feature maps. It has fast inference speed and can focus on the feature of the target in the lack of context. Furthermore, local contrast learning (LCL) is introduced, which adopts the local contrast idea of non-deep learning methods. It can alleviate the dependence on dataset samples, thereby improving detection accuracy on datasets with few samples. Compared with the state-of-the-art methods, the proposed ALCL-Net achieves superior performance with an intersection-over-union (IoU) of 0.792 and normalized IoU (nIoU) of 0.771 on the public SIRST dataset. Chuang Yu 0003, Yunpeng Liu 0001, Shuhang Wu, Zhuhua Hu, Deyan Lan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multibranch Feature Difference Learning Network for Cross-Spectral Image Patch MatchingabstractCross-spectral image patch matching is still challenging due to significant nonlinear differences between image patches. Recently, image patch matching methods based on feature relation learning have attracted increasing attention and achieved good performance. However, we find that the metric learning methods based on feature difference cannot comprehensively and effectively extract useful discriminative information between image patch pairs by only adopting two branches network structure. Therefore, we propose a novel multi-branch feature difference learning network (MFD-Net). Specifically, we build a multi-branch parallel feature difference extraction network, which can capture richer and more discriminative feature difference information and achieve significant improvements on matching tasks. Furthermore, we propose a combined metric network composed of a master metric network module and multiple branch metric network modules, which promotes the forward update of network weights and reduces the similarity of features extracted by each feature difference extraction module with negligible increase in inference time. Extensive experimental results show that the proposed MFD-Net achieves superior performances on cross-spectral image patch matching and single spectral image patch matching. Chuang Yu 0003, Yunpeng Liu 0001, Tianci Liu 0001, Zhuhua Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Automatic lumbar spinal MRI image segmentation with a multi-scale attention network
Haixing Li, Zelin Shi, Chongnan Yan, Lanbo Wang, Yueming Mu, Yunpeng Liu 0001 |
Neural Comput. Appl. | 8 |
| 2020 | DAGN: A Real-Time UAV Remote Sensing Image Vehicle Detection FrameworkabstractReal-time small object detection from the remote sensing images taken by unmanned aerial vehicles (UAVs) is a challenging but fundamental problem for many UAV applications because of the complex scales, densities, and shapes of objects that are the result of the shooting angle of the UAV. In this letter, we focus on real-time small vehicle detection for UAV remote sensing images and propose a depthwise-separable attention-guided network (DAGN) based on YOLOv3. First, we combine the feature concatenation and attention block to provide the model with the excellent ability to distinguish important and inconsequential features. Then, we improve the loss function and candidate merging algorithm in YOLOv3. Through these strategies, the performance of vehicle detection is improved, while some detection speed is sacrificed. To accelerate our model, we replace some standard convolutions with depthwise-separable convolutions. Compared to YOLOv3 and other two-stage state-of-the-art models that are applied to Vehicle Detection in Aerial Imagery (VEDAI) data sets, DAGN has a detection accuracy of 0.671, which is 5.5% better than that of YOLOv3, and it achieves the same results as two-stage methods. In addition, DAGN achieves real-time detection using GeForce GTX 1080Ti. Yunpeng Liu 0001, Tianci Liu 0001, Zhiyuan Lin 0005, Sikui Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Supervised Dimensionality Reduction on Grassmannian for Image Set RecognitionabstractModeling videos and image sets by linear subspaces has achieved great success in various visual recognition tasks. However, subspaces constructed from visual data are always notoriously embedded in a high-dimensional ambient space, which limits the applicability of existing techniques. This letter explores the possibility of proposing a geometry-aware framework for constructing lower-dimensional subspaces with maximum discriminative power from high-dimensional subspaces in the supervised scenario. In particular, we make use of Riemannian geometry and optimization techniques on matrix manifolds to learn an orthogonal projection, which shows that the learning process can be formulated as an unconstrained optimization problem on a Grassmann manifold. With this natural geometry, any metric on the Grassmann manifold can theoretically be used in our model. Experimental evaluations on several data sets show that our approach results in significantly higher accuracy than other state-of-the-art algorithms. Tianci Liu 0001, Zelin Shi, Yunpeng Liu 0001 |
Neural Comput. | 3 |
| 2018 | Joint Normalization and Dimensionality Reduction on Grassmannian: A Generalized PerspectiveabstractThis letter proposes a generalized framework with joint normalization that learns lower dimensional subspaces with maximum discriminative power by using Riemannian geometry. We model the similarity/dissimilarity between subspaces using various metrics defined on Grassmannian and formulate dimensionality reduction as a nonlinear constraint optimization problem considering the orthogonalization. To obtain the linear mapping, we derive the components required to perform Riemannian optimization from the original Grassmannian through an orthonormal projection. We respect the Riemannian geometry of the Grassmann manifold and search for this projection directly from one Grassmann manifold to another face-to-face without any additional transformations. In this natural geometry-aware approach, any metric on the Grassmann manifold can theoretically reside in our model. We combine five metrics with our model, and the learning process is treated as an unconstrained optimization problem on a Grassmann manifold. Experiments on several datasets demonstrate that our approach leads to a significant accuracy gain over state-of-the-art methods. Tianci Liu 0001, Zelin Shi, Yunpeng Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Efficient and Robust Direct Image Registration Based on Joint Geometric and Photometric Lie AlgebraabstractThis paper considers the joint geometric and photometric image registration problem. The inverse compositional (IC) algorithm and the efficient second-order minimization (ESM) algorithm are two typical efficient methods applied to the geometric registration problem. Their efficiency stems from the utilization of the group structure of geometric transformations. To allow for photometric variations, the dual IC algorithm (DIC) proposed by Bartoli performs joint geometric and photometric image registration by extending the IC algorithm. The group structures of both geometric and photometric transformations are exploited. Despite the robustness to large photometric variations, DIC is vulnerable to large geometric deformations. The ESM algorithm is extended by Silveira et al. to address photometric variations. In their approach, the photometric transformations are modeled in Euclidean space. Their approach is robust to relatively large geometric and photometric transformations; however, it is not efficient for large photometric variations. We propose a new efficient and robust image registration method by exploiting the non-Euclidean Lie group structure of joint geometric and photometric transformations for both grayscale and color images. The image registration is formulated as a nonlinear least squares problem. In our method, the geometric and photometric transformations are jointly parameterized by their corresponding Lie algebras. Based on this parameterization approach, the second-order approximation strategy of ESM is employed to optimize the joint geometric and photometric parameters. The error function in the nonlinear least squares problem is approximated by a second-order Taylor expansion with respect to joint geometric and photometric parameters without computing the Hessian matrix. For further efficiency, independent convergence criteria for geometric and photometric parameters are used in the iterative optimization process. The superiority of our proposed method over the previous methods, in terms of efficiency, accuracy, and robustness, is demonstrated through extensive experiments on synthetic and real data. Zelin Shi, Yunpeng Liu 0001, Tianci Liu 0001 |
IEEE Trans. Image Process. | 3 |