EDBT 2026 Demo / reviewers in the wild / expert
Zaiping Lin
dblp:192/3836
· DBLP profile ↗
31ranked-venue papers
0as first author
25since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Triple-Directional Fusion Attention for Infrared Small Target DetectionabstractAttention mechanism has gained popularity due to its effectiveness. However, most existing mechanisms are designed for large-sized targets currently, with limited improvement in single-frame infrared small target (SIRST) detection tasks. In this letter, we propose a novel attention mechanism to enhance the extraction capacity of deep networks for infrared small targets, termed the triple-directional fusion attention module (TFAM). This module aggregates channel-, height-, and width-dimension into three independent directional perception attention vectors, preserving both accurate channel and spatial information. Through adaptive cross-direction interaction, TFAM establishes inter-directional dependencies essential for enhancing faint target signatures in deep layers. Notably, TFAM only requires minimal complexity for modeling and offers flexibility in integration. Experiments conducted on the NUDT-SIRST and NUAA-SIRST datasets demonstrate consistent improvements. Jun Chen 0007, Shipeng Zhu, Boyang Li 0007, Jianpeng Fan, Zaiping Lin, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2025 | CWIMamba: Cross-Scale Windowed Integration State Space Model for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) intends to detect potential anomalous targets hidden in the background of hyperspectral images (HSIs) and has garnered substantial attention in various remote sensing photography and surveying applications. Recent research advances in the HAD domain have highlighted the significance of deep convolutional networks (DCNs) and vision transformers (ViTs)-based formulas. However, DCNs are long-range dependency-limited networks, whereas ViTs bear the computational burden of quadratic complexity. Owing to their prominent nonlocal representations and linear complexity, Mamba-based approaches have drawn growing attention. Our study pioneers the integration of Mamba into HAD tasks, presenting CWIMamba, which introduces a novel cross-scale windowed integration state space model for considering the spatial distribution characteristics of the anomaly targets. Specifically, we devise a cross-scale windowed state space model (CSWSSM) to scan the spatial-spectral features based on the window-based bottleneck SSM with different scales. For better multiscale feature integration, a multiscale spatial-spectral feature adaptive integration (MS3FAI) method is explored to generate an intensified representation of multiscale feature interaction and fusion based on the elaborate adaptive spatial-spectral weighting scheme. Moreover, we also devised a Haar discrete wavelet transform convolution module (HDWTCM) to fully replenish the local informative representation and enhance the discriminative frequency characteristics between anomalies and background, introducing more inductive local features for accurate background reconstruction and anomaly suppression. Extensive experiments on five multifarious HAD datasets and seven indicators substantiate the state-of-the-art detection performance, demonstrating the effectiveness of CWIMamba. Wei An 0003, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Shilin Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Infrared Small Target Detection via Nonconvex Weighted Tensor Rank Minimization and Adaptive Spatial-Temporal ModelingabstractInfrared small target detection is of great significance for various applications. However, it is significantly challenged by complex backgrounds and low signal-to-clutter ratio. Although low-rank and sparse decomposition (LRSD)-based methods are widely employed, they are hampered by fixed temporal step sizes, transpose errors in tensor recovery, and the suboptimal approximation of sparsity using thel1norm. To tackle these problems, we propose an entropy-based adaptive spatial-temporal infrared tensor with nonconvex weighted average tensor rank (EASTIT-NWTAR) method. Firstly, we propose an adaptive spatial-temporal tensor construction approach that leverages tensor information entropy to dynamically adjust the temporal step size, ensuring an accurate representation of background changes and maintaining its low-rank property. Secondly, we propose a nonconvex weighted tensor norm combining the Laplace norm and weighted average tensor rank (WTAR) norm to effectively mitigate transpose errors and enhance low-rank recovery. Finally, we substitute thel1norm with the smoothly clipped absolute deviation (SCAD) norm to improve sparse target reconstruction accuracy. The proposed method is effectively solved using the alternating direction multiplier method (ADMM). Extensive experiments demonstrate that proposed method outperforms state-of-the-art methods in both target detection and background suppression. Yang Sun 0006, Zaiping Lin, Ting Liu 0017, Boyang Li 0007, Yimian Dai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Nonuniformity and Bad Pixel Correction Based on 3-D Network With Oversmoothing SuppressionabstractThe purpose of non-uniformity and blind pixel correction is to provide a more reliable foundation for subsequent image processing and target detection. Existing correction methods generally struggle to balance the contradiction between over-smoothing and residual noise. Particularly, over-smoothing can easily filter out texture details and dim small targets. Based on the multi-frame response model of infrared focal plane array detector, we propose a two-stage 3-D residual fully convolutional network for correction factor estimation, integrated with an over-smoothing suppression mechanism. The proposed method designs two 3-D sub-networks to estimate the gain correction factors and offset correction factors respectively. For the correction factor pre-estimation tensors outputted by the two sub-networks, an inter-frame averaging after outlier removal is applied to suppress over-smoothing. Ultimately, using multiplication and addition structures, the final estimated values of the gain and offset correction factors can be utilized to obtain the corrected images. Experimental results indicate that the proposed method exhibits substantial generalization capabilities towards different intensities non-uniformity pixel-wise fixed mode noise and can effectively correct the blind pixels of real infrared images while suppressing over-smoothing and maintaining the image details such as dim small targets well. Overall, as a method that combines the model-driven and the data-driven, our method possesses strong theoretical interpretability and superior performance. Teliang Wang, Wei An 0003, Zaiping Lin, Kun Li 0029 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Satellite Video Object Detection Based on Enhanced 3DTV Regularization and Gaussian PriorabstractSatellite videos have played important roles in many applications in recent years due to the advantages of continuous providing high temporal resolution remote sensing images. Although much progress has been achieved for moving object detection (MOD) in satellite videos, the low-rank characteristics of background and the intensity variations of moving objects across frames have not been fully exploited. In this article, we propose an efficient method for MOD in satellite videos, which models the background with enhanced 3-D total variation (E-3DTV) regularization and the moving objects with Gaussian prior. Specifically, considering that the gradient maps on the spatial and temporal dimensions exhibit different physical meanings, we model the background with different Laplacian sparsity priors for the gradient maps along the spatial and temporal dimensions for 3DTV regularization. Different from current methods, which model moving objects with sparsity characteristics in each frame alone, we utilize Gaussian prior to model intensity changing characteristics of moving objects across frames. After integrating background model and moving object model into low-rank sparse matrix factorization framework, the alternating direction method of multipliers (ADMM) is adopted to iteratively optimize the parameters of background and moving object models. We conduct experiments on VISO and SkySat datasets, and the results demonstrate that our method achieves superior MOD performance with high computational efficiency compared to state-of-the-art methods. Wei An 0003, Ting Liu 0017, Yang Sun 0006, Zaiping Lin, Yulan Guo, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao |
ICPR (34) | 17 |
| 2024 | Global-to-Local Spatial-Spectral Awareness Transformer Network for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is one of the momentous technologies in the field of Earth observation and remote sensing monitoring. Profiting from puissant deep feature extraction abilities, deep convolutional networks (DCN) perform excellently in the HAD domain. Nevertheless, limited by the restriction of unique local receptive fields, DCN-based detection methods struggle to catch the long-range dependence from a global perspective. In contrast, vision transformers (ViTs) perform better in global feature extraction but still disregard the local dependence properties. To this end, we proposed a novel method entitled the global-to-local spatial-spectral awareness transformer (G2LSSAT) network, in which the global transformer block (GTB) and local transformer block (LTB) are deployed in sequence to capture deep reconstruction characteristics from the global view to the local view in a spatial-spectral domain. In particular, the GTB is designed to explore the global spatial-spectral characteristics that are dependent on a crossbar-based global sparse attention module. Furthermore, the global glanced image is divided into multiple local patches and the LTB is devised to learn the local spatial-spectral features supported by a patch-based local self-invisible attention module. In addition, considering that the abnormal pixels always be unexpectedly reconstructed with the conventional self-attention module in ViTs, we introduce a invisible diagonal mask (IDM), which is embedded into the LTB module, to overshadow each pixel itself in the receptive field and reconstruct itself based on global and local dependent spatial-spectral features. Extensive experimental results on six datasets illustrate the superiority of the proposed G2LSSAT compared with other state-of-the-art detectors. Shilin Zhou 0001, Qiang Ling 0002, Zhaoxu Li, Zaiping Lin |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Mixed-Precision Network Quantization for Infrared Small Target SegmentationabstractNetwork quantization is leveraged to reduce the model size, memory footprint, and computational cost of deep neural networks. It is achieved by representing float weights and activations with lower bit counterparts, which is essential for model deployment on resource-limited devices. However, due to the extremely small size of infrared small targets in the feature map, low-bit quantization could lead to huge information loss of small targets and thus causes severe segmentation performance degradation. To achieve low-bit quantization while maintaining the segmentation performance, we first study the quantization sensitivity of small target segmentation network and observe the sensitivity heterogeneity of different layers in the network. Specifically, feature maps in shallow layers and encoder subnetwork are more vulnerable to information loss caused by quantization as compared to deep layers and decoder subnetwork. Based on these observations, we are motivated to assign a different bitwidth for each block according to their quantization sensitivity. A simple yet effective symmetrically progressive decreasing mixed-precision quantization (SPMix-Q) method is proposed to achieve high-performance segmentation under low-bit quantization (i.e., 2.42 bits for weights and 3.82 bits for activations). The experimental results show that our SPMix-Q achieves comparable accuracy with only 1/13 model size, 1/4.6 memory footprint, and 1/29 computational cost to the full-precision counterparts. Compared with the homogeneous low-bit quantization methods, our method achieves much better performance in terms of intersection of union (IoU) on the benchmark datasets. Our mobile-system-on-a-chip (SOC) (e.g., Kyrin 980, Snapdragon 660, and Dimensity 800U) deployable android application package (APK) is available at:https://github.com/YeRen123455/SIRST-Quantization-Deployment. Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Tianhao Wu 0014, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DDAug: Differentiable Data Augmentation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation(WSSS) with image-level labels has witnessed promising advances with the help ofclass activation maps(CAM). However, CAM is always confined to small discriminative seed regions due to its simple classification loss guided training manner. To handle this problem, recent works introduced specifically designed regularizations and modules to expand the CAM seed regions, serving as the final segmentation masks. In this paper, we surprisingly find that the classification loss could suppress the gains from these regularization and modules in the late training phase, thereby limiting the further growth of CAM, which we call as theexplicit supervision disturb(ESD) issue. Interestingly, we find that specificdata augmentation(DA) operations (e.g., CutMix) can relieve such ESD issue, and the benefits introduced by different DA operations vary a lot. To maximize the benefits, we proposedifferentiable data augmentation(DDAug) to automatically search for the proper DA policy. Specifically, we design amulti-level search spaceto sequentially sample DA operations with different properties. Extensive experiments demonstrate that the proposed DDAug can alleviate the ESD issue and introduce consistent improvements to various popular WSSS methods, achieving the state-of-the-art performance on the MS COCO 2014 and PASCAL VOC 2012 datasets. Boyang Li 0007, Fei Zhang 0016, Longguang Wang, Yingqian Wang 0002, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Multim. | 6 |
| 2023 | Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point SupervisionabstractTraining a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we make the first attempt to achieve infrared small target detection with point-level supervision. Interestingly, during the training phase supervised by point labels, we discover that CNNs first learn to segment a cluster of pixels near the targets, and then gradually converge to predict groundtruth point labels. Motivated by this “mapping degeneration” phenomenon, we propose a label evolution framework named label evolution with single point supervision (LESPS) to progressively expand the point label by leveraging the intermediate predictions of CNNs. In this way, the network predictions can finally approximate the updated pseudo labels, and a pixel-level target mask can be obtained to train CNNs in an end-to-end manner. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Experimental results show that CNNs equipped with LESPS can well recover the target masks from corresponding point labels, and can achieve over 70% and 95% of their fully supervised performance in terms of pixel-level intersection over union (IoU) and object-level probability of detection (Pd), respectively. Code is available at https://github.com/XinyiYing/LESPS. Xinyi Ying, Li Liu 0002, Yingqian Wang 0002, Ruojing Li, Zaiping Lin, Weidong Sheng, Shilin Zhou 0001 |
CVPR | 6 |
| 2023 | Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-level annotations. To reduce the annotation burden, we propose the first method to achieve SIRST detection with single-point supervision. The core idea of this work is to recover the per-pixel mask of each target from the given single point label by using clustering approaches, which looks simple but is indeed challenging since targets are always insalient and accompanied with background clutters. To handle this issue, we introduce randomness to the clustering process by adding noise to the input images, and then obtain much more reliable pseudo masks by averaging the clustered results. Thanks to this "Monte Carlo" clustering approach, our method can accurately recover pseudo masks and thus turn arbitrary fully supervised SIRST detection networks into weakly supervised ones with only single point annotation. Experiments on four datasets demonstrate that our method can be applied to existing SIRST detection networks to achieve comparable performance with their fully-supervised counterparts, which reveals that single-point supervision is strong enough for SIRST detection. Our code will be available at: https://github.com/YeRen123455/SIRST-Single-Point-Supervision. Boyang Li 0007, Yingqian Wang 0002, Longguang Wang, Fei Zhang 0016, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
ICCV | 6 |
| 2023 | Exploring Fine-Grained Sparsity in Convolutional Neural Networks for Efficient InferenceabstractNeural networks contain considerable redundant computation, which drags down the inference efficiency and hinders the deployment on resource-limited devices. In this paper, we study the sparsity in convolutional neural networks and propose a generic sparse mask mechanism to improve the inference efficiency of networks. Specifically, sparse masks are learned in both data and channel dimensions to dynamically localize and skip redundant computation at a fine-grained level. Based on our sparse mask mechanism, we develop SMPointSeg, SMSR, and SMStereo for point cloud semantic segmentation, single image super-resolution, and stereo matching tasks, respectively. It is demonstrated that our sparse masks are well compatible to different model components and network architectures to accurately localize redundant computation, with computational cost being significantly reduced for practical speedup. Extensive experiments show that our SMPointSeg, SMSR, and SMStereo achieve state-of-the-art performance on benchmark datasets in terms of both accuracy and efficiency. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Anomaly Detection for Hyperspectral Imagery via Tensor Low-Rank Approximation With Multiple Subspace LearningabstractHyperspectral anomaly detection (HAD) is regarded as an indispensable, pivotal technology in remote sensing and earth science domains. Nevertheless, most existing detection approaches for anomaly targets flatten 3-D hyperspectral images (HSIs) with spatial and spectral information into 2-D spectral vector data, which virtually breaks up the internal spatial structure in HSIs and degenerates the detection performance. To this end, we directly consider the HSI data cube as a 3-D tensor and develop a novel tensor low-rank approximation (TLRA) detection algorithm to separate the sparse anomalous component from the background with low-rank characteristics. Then, in light of the multi-subspace structure in heterogeneous backgrounds, we utilize multiple subspace learning (MSL) theory to encode the background tensor with a coefficient tensor and corresponding dictionary tensor. In addition, considering that different singular values indicate different information quantities and should be penalized to different extents, we introduce a tighter tensor rank surrogate named the ϵ-shrinkage tensor nuclear norm (ϵ-TNN) to recover the low-rank component more accurately. Meanwhile, concerning the sparse anomaly target, thel2,1constraint is incorporated to represent the group sparsity of the abnormal component. Finally, an effective iterative optimization algorithm based on the alternating direction method of multipliers (ADMM) is devised to solve the proposed TLRA-MSL model. We conduct extensive experiments on six hyperspectral datasets to prove the effectiveness and robustness of our method. The experimental results illustrate that better detection performance is obtained using the proposed model compared with other state-of-the-art algorithms. Qiang Ling 0002, Zhaoxu Li, Zaiping Lin, Shilin Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | You Only Train Once: Learning a General Anomaly Enhancement Network With Random Masks for Hyperspectral Anomaly DetectionabstractIn this paper, we introduce a new approach to address the challenge of generalization in hyperspectral anomaly detection (AD). Our method eliminates the need for adjusting parameters or retraining on new test scenes as required by most existing methods. Employing an image-level training paradigm, we achieve a general anomaly enhancement network for hyperspectral AD that only needs to be trained once. Trained on a set of anomaly-free hyperspectral images with random masks, our network can learn the spatial context characteristics between anomalies and background in an unsupervised way. Additionally, a plug-and-play model selection module is proposed to search for a spatial-spectral transform domain that is more suitable for AD task than the original data. To establish a unified benchmark to comprehensive evaluate our method and existing methods, we develop a large-scale hyperspectral AD dataset (HAD100) that includes 100 real test scenes with diverse anomaly targets. In comparison experiments, we combine our network with a parameter-free detector, and achieve the optimal balance between detection accuracy and inference speed among state-of-the-art AD methods. Experimental results also show that our method still achieves competitive performance when the training and test set are captured by different sensor devices. Our code is available at https://github.com/ZhaoxuLi123/AETNet. Zhaoxu Li, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Dense Nested Attention Network for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds. With the advances of deep learning, CNN-based methods have yielded promising results in generic object detection due to their powerful modeling capability. However, existing CNN-based methods cannot be directly applied to infrared small targets since pooling layers in their networks could lead to the loss of targets in deep layers. To handle this problem, we propose a dense nested attention network (DNA-Net) in this paper. Specifically, we design a dense nested interactive module (DNIM) to achieve progressive interaction among high-level and low-level features. With the repetitive interaction in DNIM, the information of infrared small targets in deep layers can be maintained. Based on DNIM, we further propose a cascaded channel and spatial attention module (CSAM) to adaptively enhance multi-level features. With our DNA-Net, contextual information of small targets can be well incorporated and fully exploited by repetitive fusion and enhancement. Moreover, we develop an infrared small target dataset (namely, NUDT-SIRST) and propose a set of evaluation metrics to conduct comprehensive performance evaluation. Experiments on both public and our self-developed datasets demonstrate the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of probability of detection (${P}_{d}$), false-alarm rate (${F}_{a}$), and intersection of union ($IoU$). Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 5 |
| 2022 | Detecting Dim Small Target in Infrared Images via Subpixel Sampling Cuneate NetworkabstractInfrared dim small target detection is regarded as a critical technology for the interpretation of space-based remote sensing images. In recent years, driven by deep learning technology and the surge of data, remarkable effects have been achieved for dim small target detection in infrared images. Nevertheless, the intrinsic feature scarcity and low signal-to-clutter ratio (SCR) characteristics pose tremendous challenges to deep learning-based detection methods. In this letter, we present a novel sub-pixel sampling cuneate network (SPSCNet) to detect dim small targets in infrared images. The overall model architecture is based on an end-to-end cuneate network with multiple groups of parallel high-to-low resolution subnetworks. Specifically, we design a multi-scale feature reweighted fusion (MSFRF) module to effectively fuse multi-scale feature maps which contain both low-level detail features and high-level semantics information. In addition, considering that the pooling operation may lose dim small targets with low SCR, we also exploit a sub-pixel sampling scheme to greatly retain the features of small targets. Moreover, to better test and verify the performance of the proposed method, we also develop an infrared dim small target (IDST) dataset to conduct more comparative experiments. Extensive experiments on the SIRST and IDST datasets illustrate that the proposed SPSCNet yields state-of-the-art performance in comparison with other detection algorithms. Qiang Ling 0002, Zaiping Lin, Shilin Zhou 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Moving Object Detection in Satellite Videos via Spatial-Temporal Tensor Model and Weighted Schatten p-Norm MinimizationabstractLow-rank matrix decomposition approaches have achieved significant progress in small and dim object detection in satellite videos. However, it is still challenging to achieve robust performance and fast processing under complex and highly heterogeneous backgrounds since satellite video data can neither adequately fit the foreground structure nor the background model in the existing matrix decomposition models. In this letter, we propose a novel object detection method based on a spatial–temporal tensor data structure. First, we construct a tensor data structure to exploit the inner spatial and temporal correlation within a satellite video. Second, we extend the decomposition formulation with bounded noise to achieve robust performance under complex backgrounds. This formulation integrates low-rank background, structured sparse foreground, and their noises into a tensor decomposition problem. For background separation, a weighted Schatten$p$-norm is incorporated to provide adaptive threshold to obtain the singular value of the background tensor. Finally, the proposed model is solved using the alternative direction method of multipliers (ADMM) scheme. Experimental results on various real scenes demonstrate the superiority of the proposed method against the compared approaches. Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Parallax Attention for Unsupervised Stereo Correspondence LearningabstractStereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Spectral-Spatial Deep Support Vector Data Description for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to distinguish anomalies from background-by-background modeling. Deep learning has been applied to HAD and achieves promising detection results. However, there exist several issues that need to be addressed: 1) unrealistic Gaussian assumption on the latent representations may limit its application; 2) deep features are not well-suited to anomaly detection due to the separation between feature learning and anomaly detection; 3) lack of adequate exploitation of spectral-spatial features; 4) negative effect caused by spectral band redundancy. In this article, we propose an end-to-end trainable deep one-class classification network for HAD. Specifically, a minimal enclosing hypersphere is trained to involve the deep features of background samples. These background samples are selected by a density clustering-based method. In this way, feature learning and anomaly detection are incorporated into a unified framework. Meanwhile, there is no explicit Gaussian assumption on the background features. Moreover, due to the complementarity of spectral and spatial features, a novel feature fusion strategy is proposed to fuse spectral and spatial features extracted by a two-stream deep convolutional autoencoder network. Finally, a band attention module is used to automatically learn small weights for redundant bands and thus reduce the negative effect caused by redundant bands. Experimental results on five public datasets demonstrate the superiority of the proposed method compared to several state-of-the-art HAD methods in the detection performance. Kun Li 0029, Qiang Ling 0002, Yao Qin 0002, Yingqian Wang 0002, Yaoming Cai, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A BenchmarkabstractSatellite video cameras can provide continuous observation for a large-scale area, which is important for many remote sensing applications. However, achieving moving object detection and tracking in satellite videos remains challenging due to the insufficient appearance information of objects and lack of high-quality datasets. In this article, we first build a large-scale satellite video dataset with rich annotations for the task of moving object detection and tracking. This dataset is collected by the Jilin-1 satellite constellation and composed of 47 high-quality videos with 1 646 038 instances of interest for object detection and 3711 trajectories for object tracking. We then introduce a motion modeling baseline to improve the detection rate and reduce false alarms based on accumulative multiframe differencing and robust matrix completion. Finally, we establish the first public benchmark for moving object detection and tracking in satellite videos and extensively evaluate the performance of several representative approaches on our dataset. Comprehensive experimental analyses and insightful conclusions are also provided. The dataset is available athttps://github.com/QingyongHu/VISO. Qingyong Hu, Hao Liu 0061, Feng Zhang 0046, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Exploring Sparsity in Image Super-Resolution for Efficient InferenceabstractCurrent CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for those flat regions. Therefore, existing CNN-based methods involve redundant computation in flat regions, which increases their computational cost and limits their applications on mobile devices. In this paper, we explore the sparsity in image SR to improve inference efficiency of SR networks. Specifically, we develop a Sparse Mask SR (SMSR) network to learn sparse masks to prune redundant computation. Within our SMSR, spatial masks learn to identify "important" regions while channel masks learn to mark redundant channels in those "unimportant" regions. Consequently, redundant computation can be accurately localized and skipped while maintaining comparable performance. It is demonstrated that our SMSR achieves state-of-the-art performance with 41%/33%/27% FLOPs being reduced for ×2/3/4 SR. Code is available at: https://github.com/LongguangWang/SMSR. Longguang Wang, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003, Yulan Guo |
CVPR | 5 |
| 2021 | Learning A Single Network for Scale-Arbitrary Super-ResolutionabstractRecently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., ×2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to learn a scale-arbitrary image SR network from scale-specific networks. Specifically, we develop a plug-in module for existing SR networks to perform scale-arbitrary SR, which consists of multiple scale-aware feature adaption blocks and a scale-aware upsampling layer. Moreover, conditional convolution is used in our plug-in module to generate dynamic scale-aware filters, which enables our network to adapt to arbitrary scale factors. Our plug-in module can be easily adapted to existing networks to realize scale-arbitrary SR with a single model. These networks plugged with our module can produce promising results for non-integer and asymmetric SR while maintaining state-of-the-art performance for SR with integer scale factors. Besides, the additional computational and memory cost of our module is very small. Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
ICCV | 3 |
| 2021 | Segmentation-Based Weighting Strategy for Hyperspectral Anomaly DetectionabstractBackground information extraction and modeling have always been the cores of hyperspectral anomaly detection (AD) algorithms. In this letter, a simple and efficient weighting strategy based on image segmentation is proposed. The strategy integrates detection results with spatial information from segmentation, and the background is suppressed in the fusion results. First, an improved graph-based image-segmentation algorithm is adopted to isolate potential anomaly targets from the background. The segmentation is based on the spectral similarity of adjacent pixels and can extract the potential target well even if the global background is complex. Then, a weight matrix is constructed according to the segmentation result, and two types of background, narrow boundaries and large homogeneous areas, are suppressed and assigned small weights. Finally, the normalized weight matrix is combined with the detection results of AD algorithms. Experiments conducted on different data sets show that the proposed strategy is efficient and robust and can improve the detection performance and robustness of AD algorithms. Zhaoxu Li, Qiang Ling 0002, Zaiping Lin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Deep Video Super-Resolution Using HR Optical Flow EstimationabstractVideo super-resolution (SR) aims at generating a sequence of high-resolution (HR) frames with plausible and temporally consistent details from their low-resolution (LR) counterparts. The key challenge for video SR lies in the effective exploitation of temporal dependency between consecutive frames. Existing deep learning based methods commonly estimate optical flows between LR frames to provide temporal dependency. However, the resolution conflict between LR optical flows and HR outputs hinders the recovery of fine details. In this paper, we propose an end-to-end video SR network to super-resolve both optical flows and images. Optical flow SR from LR frames provides accurate temporal dependency and ultimately improves video SR performance. Specifically, we first propose an optical flow reconstruction network (OFRnet) to infer HR optical flows in a coarse-to-fine manner. Then, motion compensation is performed using HR optical flows to encode temporal dependency. Finally, compensated LR inputs are fed to a super-resolution network (SRnet) to generate SR results. Extensive experiments have been conducted to demonstrate the effectiveness of HR optical flows for SR performance improvement. Comparative results on the Vid4 and DAVIS-10 datasets show that our network achieves the state-of-the-art performance. Longguang Wang, Yulan Guo, Li Liu 0002, Zaiping Lin, Xinpu Deng, Wei An 0003 |
IEEE Trans. Image Process. | 4 |
| 2019 | Learning Parallax Attention for Stereo Image Super-ResolutionabstractStereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets. Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 4 |
| 2019 | A Constrained Sparse Representation Model for Hyperspectral Anomaly DetectionabstractIn this paper, we propose a novel sparsity-based algorithm for anomaly detection in hyperspectral imagery. The algorithm is based on the concept that a background pixel can be approximately represented as a sparse linear combination of its spatial neighbors while an anomaly pixel cannot if the anomalies are removed from its neighborhood. To be physically meaningful, the sum-to-one and nonnegativity constraints are imposed to abundance vector based on the linear mixture model, and the upper bound constraint on sparsity level is removed for better recovery of the test pixel. First, the proposed method utilizes the redundant background information to automatically remove anomalies from the background dictionary. Then, the reconstruction error obtained by the new background dictionary is directly used for anomaly detection. Moreover, a kernel version of the proposed method is also derived to completely exploit the nonlinear feature of hyperspectral data. An important advantage of the proposed methods is their capability to adaptively model the background even when some anomaly pixels are involved. Extensive experiments have been conducted on three real hyperspectral data sets. It is demonstrated that the proposed detectors achieve a promising detection performance with a relatively low computational cost. Qiang Ling 0002, Yulan Guo, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Learning for Video Super-Resolution Through HR Optical Flow Estimation
Longguang Wang, Yulan Guo, Zaiping Lin, Xinpu Deng, Wei An 0003 |
ACCV (1) | 3 |
| 2018 | Infrared Small Target Detection Using Multiscale Gray and Variance Difference
Jinyan Gao, Yulan Guo, Zaiping Lin, Wei An 0003 |
PRCV (4) | 3 |
| 2017 | FSVO: Semi-direct monocular visual odometry using fixed mapsabstractWe propose a fixed-map semi-direct visual odometry (FSVO) algorithm for Micro Aerial Vehicles (MAVs). The proposed approach does not need computationally expensive feature extraction and matching techniques for motion estimation at each frame. Instead, we extract and match ORiented Brief (ORB) features between keyframes and assist-frames. We replace the incremental map generation step in traditional algorithms with fixed map generation at keyframe and assistframe only in our algorithm, resulting in reduced storage memory and higher flexibility for relocalization. Based on the fixed-map, we design a new keyframe selection criterion and a relocalization step. Our algorithm has no limit on the orientation of the camera and reduces drifting effectively. Experimental results on the EuRoC and KITTI datasets show that our algorithm achieves higher precision and robustness than the SVO algorithm. Zhiheng Fu, Yulan Guo, Zaiping Lin, Wei An 0003 |
ICIP | 3 |