VLDB 2026 Research / reviewers in the wild / expert
Boyang Li 0007
dblp:70/1211-7
· DBLP profile ↗
24ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-4479-9008ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structured Grouping Collaborative Decorrelated Regularization for Model Pruning in Infrared Small Target Detection
Yonghao Li, Jun Chen 0005, Boyang Li 0007, Yulan Guo, Longguang Wang, Siyi Deng |
Pattern Recognit. | 3 |
| 2025 | Triple-Directional Fusion Attention for Infrared Small Target DetectionabstractAttention mechanism has gained popularity due to its effectiveness. However, most existing mechanisms are designed for large-sized targets currently, with limited improvement in single-frame infrared small target (SIRST) detection tasks. In this letter, we propose a novel attention mechanism to enhance the extraction capacity of deep networks for infrared small targets, termed the triple-directional fusion attention module (TFAM). This module aggregates channel-, height-, and width-dimension into three independent directional perception attention vectors, preserving both accurate channel and spatial information. Through adaptive cross-direction interaction, TFAM establishes inter-directional dependencies essential for enhancing faint target signatures in deep layers. Notably, TFAM only requires minimal complexity for modeling and offers flexibility in integration. Experiments conducted on the NUDT-SIRST and NUAA-SIRST datasets demonstrate consistent improvements. Jun Chen 0007, Shipeng Zhu, Boyang Li 0007, Jianpeng Fan, Zaiping Lin, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Graph Laplacian regularization for fast infrared small target detection
Ting Liu 0017, Yongxian Liu, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 4 |
| 2025 | Differentiable Prior-Driven Data Augmentation for Sensor-Based Human Activity RecognitionabstractSensor-based human activity recognition (HAR) usually suffers from the problem of insufficient annotated data, due to the difficulty in labeling the intuitive signals of wearable sensors. To this end, recent advances have adopted handcrafted operations or generative models for data augmentation. The handcrafted operations are driven by some physical priors of human activities, e.g., action distortion and strength fluctuations. However, these approaches may face challenges in maintaining semantic data properties. Although the generative models have better data adaptability, it is difficult for them to incorporate important action priors into data generation. This article proposes a differentiable prior-driven data augmentation framework for HAR. First, we embed the handcrafted augmentation operations into a differentiable module, which adaptively selects and optimizes the operations to be combined together. Then, we construct a generative module to add controllable perturbations to the data derived by the handcrafted operations and further improve the diversity of data augmentation. By integrating the handcrafted operation module and the generative module into one learnable framework, the generalization performance of the recognition models is enhanced effectively. Extensive experimental results with three different classifiers on five public datasets demonstrate the effectiveness of the proposed framework. Project page:https://github.com/crocodilegogogo/DriveData-Under-Review. Ye Zhang 0037, Qing Gao 0002, Qingtang Ding, Boyang Li 0007, Yulan Guo |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Infrared Small Target Detection via Nonconvex Weighted Tensor Rank Minimization and Adaptive Spatial-Temporal ModelingabstractInfrared small target detection is of great significance for various applications. However, it is significantly challenged by complex backgrounds and low signal-to-clutter ratio. Although low-rank and sparse decomposition (LRSD)-based methods are widely employed, they are hampered by fixed temporal step sizes, transpose errors in tensor recovery, and the suboptimal approximation of sparsity using thel1norm. To tackle these problems, we propose an entropy-based adaptive spatial-temporal infrared tensor with nonconvex weighted average tensor rank (EASTIT-NWTAR) method. Firstly, we propose an adaptive spatial-temporal tensor construction approach that leverages tensor information entropy to dynamically adjust the temporal step size, ensuring an accurate representation of background changes and maintaining its low-rank property. Secondly, we propose a nonconvex weighted tensor norm combining the Laplace norm and weighted average tensor rank (WTAR) norm to effectively mitigate transpose errors and enhance low-rank recovery. Finally, we substitute thel1norm with the smoothly clipped absolute deviation (SCAD) norm to improve sparse target reconstruction accuracy. The proposed method is effectively solved using the alternating direction multiplier method (ADMM). Extensive experiments demonstrate that proposed method outperforms state-of-the-art methods in both target detection and background suppression. Yang Sun 0006, Zaiping Lin, Ting Liu 0017, Boyang Li 0007, Yimian Dai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Motion and Appearance Decoupling Representation for Event CamerasabstractEvent cameras, with high temporal resolution and high dynamic range, have shown great potential under extreme scenarios such as high-speed movement and low illumination. However, previous event representation methods typically aggregate event data into a single dense tensor, often overlooking the dynamic changes of events within a given time unit. This limitation can introduce historical artifacts and semantic inconsistencies, ultimately degrading model performance. Inspired by human visual prior, we propose a motion and appearance decoupling (MAD) event representation to disentangle the mixed spatial-temporal event tensor into two independent branches. This bio-inspired design helps the network extract discriminative temporal (i.e., motion) and spatial (i.e., appearance) information, thus reducing the network's learning burden toward complex high-level interpretation tasks. In our method, the event motion guided attention module (EMGA) is designed to achieve temporal and spatial feature interaction and fusion sequentially. Based on EMGA, three specially designed decoder heads are proposed for several representative event-based tasks (i.e., object detection, semantic segmentation, and human pose estimation). Experimental results demonstrate that our method achieves state-of-the-art performance on the above three tasks, which reveals that our method is an easy-to-implement replacement for currently event-based methods. Our code is available at: https://github.com/ChenYichen9527/MAD-representation. Boyang Li 0007, Yingqian Wang 0002, Xinyi Ying, Longguang Wang, Chushu Zhang, Yulan Guo, Wei An 0003 |
IEEE Trans. Image Process. | 2 |
| 2025 | Direction-Coded Temporal U-Shape Module for Multiframe Infrared Small Target DetectionabstractInfrared small target (IRST) detection aims at separating targets from cluttered background. Although many deep learning-based single-frame IRST (SIRST) detection methods have achieved promising detection performance, they cannot deal with extremely dim targets while suppressing the clutters since the targets are spatially indistinctive. Multiframe IRST (MIRST) detection can well handle this problem by fusing the temporal information of moving targets. However, the extraction of motion information is challenging since general convolution is insensitive to motion direction. In this article, we propose a simple yet effective direction-coded temporal U-shape module (DTUM) for MIRST detection. Specifically, we build a motion-to-data mapping to distinguish the motion of targets and clutters by indexing different directions. Based on the motion-to-data mapping, we further design a direction-coded convolution block (DCCB) to encode the motion direction into features and extract the motion information of targets. Our DTUM can be equipped with most single-frame networks to achieve MIRST detection. Moreover, in view of the lack of MIRST datasets, including dim targets, we build a multiframe infrared small and dim target dataset (namely, NUDT-MIRSDT) and propose several evaluation metrics. The experimental results on the NUDT-MIRSDT dataset demonstrate the effectiveness of our method. Our method achieves the state-of-the-art performance in detecting infrared small and dim targets and suppressing false alarms. Our codes will be available at https://github.com/TinaLRJ/Multi-frame-infrared-small-target-detection-DTUM. Ruojing Li, Wei An 0003, Boyang Li 0007, Yingqian Wang 0002, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao |
ICPR (34) | 1 |
| 2024 | Learning Remote Sensing Object Detection With Single Point SupervisionabstractPointly Supervised Object Detection (PSOD) has attracted considerable interests due to its lower labeling cost as compared to box-level supervised object detection. However, the complex scenes, densely packed and dynamic-scale objects in Remote Sensing (RS) images hinder the development of PSOD methods in RS field. In this paper, we make the first attempt to achieve RS object detection with single point supervision, and propose a PSOD method tailored for RS images. Specifically, we design a point label upgrader (PLUG) to generate pseudo box labels from single point labels, and then use the pseudo boxes to supervise the optimization of existing detectors. Moreover, to handle the challenge of the densely packed objects in RS images, we propose a sparse feature guided semantic prediction module which can generate high-quality semantic maps by fully exploiting informative cues from sparse objects. Extensive ablation studies on the DOTA dataset have validated the effectiveness of our method. Our method can achieve significantly better performance as compared to state-of-the-art image-level and point-level supervised detection methods, and reduce the performance gap between PSOD and box-level supervised object detection. Code is available at https://github.com/heshitian/PLUG. Shitian He, Huanxin Zou, Yingqian Wang 0002, Boyang Li 0007, Ning Jing |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MCGC: A Multiscale Chain Growth Clustering Algorithm for Generating Infrared Small Target Mask Under Single-Point SupervisionabstractDue to the lack of color and texture information and the fuzzy boundary of infrared (IR) small targets, the pixel-level mask annotation process consumes a lot of manual cost and is difficult to achieve accurate annotation. To further reduce the annotation burden, we propose an IR small target mask generation algorithm based on single-point supervised multi-scale chain growth clustering (MCGC). The core of this work is the adaptive generation of IR small-target pseudo mask maps under the supervision of randomly given single-point labels, sequentially through the strategies of multi-scale chain growth, Euclidean coefficient decay, K-Means clustering, and eight-neighborhood clustering. On the four public datasets, ablation experiments, qualitative and quantitative comparison experiments demonstrate that the MCGC algorithm has an efficient and accurate IR small target pseudo mask generation capability, which can be adapted to different numbers, scales, shapes, and intensities of targets in complex backgrounds. In addition, IR-Labelmask, an IR small target mask annotation software designed based on the MCGC algorithm, is publicly available on kourenke/IR-Labelmask-software (github.com). To our knowledge, this is the first mask annotation software designed for IR small target. Renke Kou, Chunping Wang 0001, Qiang Fu 0017, Zhanwu Li, Ying Luo 0001, Boyang Li 0007, Wei Li 0032, Zhenming Peng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Mixed-Precision Network Quantization for Infrared Small Target SegmentationabstractNetwork quantization is leveraged to reduce the model size, memory footprint, and computational cost of deep neural networks. It is achieved by representing float weights and activations with lower bit counterparts, which is essential for model deployment on resource-limited devices. However, due to the extremely small size of infrared small targets in the feature map, low-bit quantization could lead to huge information loss of small targets and thus causes severe segmentation performance degradation. To achieve low-bit quantization while maintaining the segmentation performance, we first study the quantization sensitivity of small target segmentation network and observe the sensitivity heterogeneity of different layers in the network. Specifically, feature maps in shallow layers and encoder subnetwork are more vulnerable to information loss caused by quantization as compared to deep layers and decoder subnetwork. Based on these observations, we are motivated to assign a different bitwidth for each block according to their quantization sensitivity. A simple yet effective symmetrically progressive decreasing mixed-precision quantization (SPMix-Q) method is proposed to achieve high-performance segmentation under low-bit quantization (i.e., 2.42 bits for weights and 3.82 bits for activations). The experimental results show that our SPMix-Q achieves comparable accuracy with only 1/13 model size, 1/4.6 memory footprint, and 1/29 computational cost to the full-precision counterparts. Compared with the homogeneous low-bit quantization methods, our method achieves much better performance in terms of intersection of union (IoU) on the benchmark datasets. Our mobile-system-on-a-chip (SOC) (e.g., Kyrin 980, Snapdragon 660, and Dimensity 800U) deployable android application package (APK) is available at:https://github.com/YeRen123455/SIRST-Quantization-Deployment. Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Tianhao Wu 0014, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Single-Frame Infrared Small Target Detection via Gaussian Curvature Inspired NetworkabstractSingle-frame infrared small target detection (SIRSTD) is in urgent demand for many practical tasks, such as fire rescue and urban management systems, benefiting from the excellent performance of infrared (IR) imaging in harsh climates and low-light environments. SIRSTD strives to segment small targets from the background as accurately as possible. However, in a real-world application, complex background environments with high brightness and strong edges have similar physical characteristics to small IR targets, which makes it extremely difficult to separate small targets. To address this challenge, we propose a novel Gaussian Curvature Inspired Network (GCI-Net). Inspired by the well-known Gaussian curvature, we develop a Gaussian curvature-based branch (GCB) to eliminate the smoothing noise and preserve the target structure texture information. In addition, we design a complementary patch-group attention (PGA) module that relies on the complementary relationship between low-level and high-level features to provide accurate guidance for GCB. The curvature information generated by the GCB is continuously optimized under the constraint of the curvature information of the ground truth. The proposed GCI-Net provides a reliable guarantee for accurate separation of small targets from the background. We conduct extensive experiments on the public IRSTD-1k and SIRST datasets. The experimental results demonstrate that the proposed GCI-Net outperforms the state-of-the-art (SOTA) methods. Mingjin Zhang, Ke Yue, Boyang Li 0007, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | DDAug: Differentiable Data Augmentation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation(WSSS) with image-level labels has witnessed promising advances with the help ofclass activation maps(CAM). However, CAM is always confined to small discriminative seed regions due to its simple classification loss guided training manner. To handle this problem, recent works introduced specifically designed regularizations and modules to expand the CAM seed regions, serving as the final segmentation masks. In this paper, we surprisingly find that the classification loss could suppress the gains from these regularization and modules in the late training phase, thereby limiting the further growth of CAM, which we call as theexplicit supervision disturb(ESD) issue. Interestingly, we find that specificdata augmentation(DA) operations (e.g., CutMix) can relieve such ESD issue, and the benefits introduced by different DA operations vary a lot. To maximize the benefits, we proposedifferentiable data augmentation(DDAug) to automatically search for the proper DA policy. Specifically, we design amulti-level search spaceto sequentially sample DA operations with different properties. Extensive experiments demonstrate that the proposed DDAug can alleviate the ESD issue and introduce consistent improvements to various popular WSSS methods, achieving the state-of-the-art performance on the MS COCO 2014 and PASCAL VOC 2012 datasets. Boyang Li 0007, Fei Zhang 0016, Longguang Wang, Yingqian Wang 0002, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Multim. | 1 |
| 2023 | Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-level annotations. To reduce the annotation burden, we propose the first method to achieve SIRST detection with single-point supervision. The core idea of this work is to recover the per-pixel mask of each target from the given single point label by using clustering approaches, which looks simple but is indeed challenging since targets are always insalient and accompanied with background clutters. To handle this issue, we introduce randomness to the clustering process by adding noise to the input images, and then obtain much more reliable pseudo masks by averaging the clustered results. Thanks to this "Monte Carlo" clustering approach, our method can accurately recover pseudo masks and thus turn arbitrary fully supervised SIRST detection networks into weakly supervised ones with only single point annotation. Experiments on four datasets demonstrate that our method can be applied to existing SIRST detection networks to achieve comparable performance with their fully-supervised counterparts, which reveals that single-point supervision is strong enough for SIRST detection. Our code will be available at: https://github.com/YeRen123455/SIRST-Single-Point-Supervision. Boyang Li 0007, Yingqian Wang 0002, Longguang Wang, Fei Zhang 0016, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
ICCV | 1 |
| 2023 | Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationabstractThis paper studies the problem of weakly open-vocabulary semantic segmentation (WOVSS), which learns to segment objects of arbitrary classes using mere image-text pairs. Existing works turn to enhance the vanilla vision transformer by introducing explicit grouping recognition, i.e., employing several group tokens/centroids to cluster the image tokens and perform the group-text alignment. Nevertheless, these methods suffer from a granularity inconsistency regarding the usage of group tokens, which are aligned in the all-to-one v.s. one-to-one manners during the training and inference phases, respectively. We argue that this discrepancy arises from the lack of elaborate supervision for each group token. To bridge this granularity gap, this paper explores explicit supervision for the group tokens from the prototypical knowledge. To this end, this paper proposes the non-learnable prototypical regularization (NPR) where non-learnable prototypes are estimated from source features to serve as supervision and enable contrastive matching of the group tokens. This regularization encourages the group tokens to segment objects with less redundancy and capture more comprehensive semantic regions, leading to increased compactness and richness. Based on NPR, we propose the prototypical guidance segmentation network (PGSeg) that incorporates multi-modal regularization by leveraging prototypical sources from both images and texts at different levels, progressively enhancing the segmentation capability with diverse prototypical patterns. Experimental results show that our proposed method achieves state-of-the-art performance on several benchmark datasets. Fei Zhang 0016, Tianfei Zhou, Boyang Li 0007, Chaofan Ma, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001 |
NeurIPS | 3 |
| 2023 | Infrared Small Target Detection via Nonconvex Tensor Tucker Decomposition With Factor PriorabstractInfrared small target detection in complex scenes is an important but challenging research hotspot in infrared early warning fields. Previous studies have proved that low-rank Tucker decomposition (TD) achieves good detection performance in complex scenes. However, a key limitation of existing low-rank TD methods is that the rank needs to be set in advance, and an inaccurate predefined rank can lead to performance degradation. Inspired by the theorem that n-rank is upper bounded by the rank of each Tucker factor matrix, we propose a nonconvex tensor TD model with factor prior for infrared small target detection. In our method, we use a logdet-based function to constrain the latent factors of low-rank TD, which avoids empirical rank selection and sufficiently uses the latent data structure information in the factor matrix. Meanwhile, performing singular value decomposition (SVD) calculations on small factor matrices can reduce computational complexity. Then, group sparsity regularized total variation is used to better exploit the shared sparse pattern of difference images, which helps better remove background clutter and obtain better detection results. Finally, the proposed method is efficiently solved by the well-designed alternating direction method of multipliers (ADMM). Extensive experimental results demonstrate that our method is more effective and robust in complex scenes than other state-of-the-art methods. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Representative Coefficient Total Variation for Efficient Infrared Small Target DetectionabstractLow-rank and sparse decomposition based models are powerful and robust tools for infrared small target detection. However, due to the calculation of singular value decomposition (SVD) and the optimization of complex regularization terms, existing low-rank models often suffer from high computational complexity. To solve this problem, based on the theorem that representative coefficient matrix obtained by orthogonal transformation of data matrix can inherit the spatial structure of data matrix, we propose a representative coefficient total variation (RCTV) method for efficient infrared small target detection. In our method, we use total variational to constraint representative coefficient matrix instead of data matrix to describe local smooth prior, which helps remove noise and reduce computational complexity. Meanwhile, we control the number of columns in the representative coefficient matrix to maintain the low-rank characteristics of background, which avoids SVD calculation and improves detection efficiency. Therefore, the RCTV regularization can simultaneously describe local smooth prior and low-rank prior. Moreover, to better enhance the sparsity of targets and distinguish sparse non-target points, we use the log-sum function to adaptively assign weights to targets. It helps obtain more accurate detection performance. The proposed model is efficiently solved by the alternating direction multiplier method (ADMM). A large number of experiments show that the proposed method is superior to existing low-rank methods in both detection accuracy and efficiency. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | MTU-Net: Multilevel TransUNet for Space-Based Infrared Tiny Ship DetectionabstractSpace-based infrared tiny ship detection aims at separating tiny ships from the images captured by Earth-orbiting satellites. Due to the extremely large image coverage area (e.g., thousands of square kilometers), candidate targets in these images are much smaller, dimer, and more changeable than those targets observed by aerial- and land-based imaging devices. Existing short imaging distance-based infrared datasets and target detection methods cannot be well adopted to the space-based surveillance task. To address these problems, we develop a space-based infrared tiny ship detection dataset (namely, NUDT-SIRST-Sea) with 48 space-based infrared images and$17\,598$pixel-level tiny ship annotations. Each image covers about$10\,000$km2of area with$10 \ 000\,\, \times \ 10 \ 000$pixels. Considering the extreme characteristics (e.g., small, dim, and changeable) of those tiny ships in such challenging scenes, we propose a multilevel TransUNet (MTU-Net) in this article. Specifically, we design a vision Transformer (ViT) convolutional neural network (CNN) hybrid encoder to extract multilevel features. Local feature maps are first extracted by several convolution layers and then fed into the multilevel feature extraction module [multilevel ViT module (MVTM)] to capture long-distance dependency. We further propose a copy–rotate–resize–paste (CRRP) data augmentation approach to accelerate the training phase, which effectively alleviates the issue of sample imbalance between targets and background. Besides, we design a FocalIoU loss to achieve both target localization and shape description. Experimental results on the NUDT-SIRST-Sea dataset show that our MTU-Net outperforms traditional and existing deep learning-based single-frame infrared small target (SIRST) methods in terms of probability of detection, false alarm rate, and intersection over union. Our code is available athttps://github.com/TianhaoWu16/Multi-level-TransUNet-for-Space-based-Infrared-Tiny-ship-Detection Tianhao Wu 0014, Boyang Li 0007, Yihang Luo, Yingqian Wang 0002, Ting Liu 0017, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Dense Nested Attention Network for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds. With the advances of deep learning, CNN-based methods have yielded promising results in generic object detection due to their powerful modeling capability. However, existing CNN-based methods cannot be directly applied to infrared small targets since pooling layers in their networks could lead to the loss of targets in deep layers. To handle this problem, we propose a dense nested attention network (DNA-Net) in this paper. Specifically, we design a dense nested interactive module (DNIM) to achieve progressive interaction among high-level and low-level features. With the repetitive interaction in DNIM, the information of infrared small targets in deep layers can be maintained. Based on DNIM, we further propose a cascaded channel and spatial attention module (CSAM) to adaptively enhance multi-level features. With our DNA-Net, contextual information of small targets can be well incorporated and fully exploited by repetitive fusion and enhancement. Moreover, we develop an infrared small target dataset (namely, NUDT-SIRST) and propose a set of evaluation metrics to conduct comprehensive performance evaluation. Experiments on both public and our self-developed datasets demonstrate the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of probability of detection (${P}_{d}$), false-alarm rate (${F}_{a}$), and intersection of union ($IoU$). Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 1 |
| 2022 | Gated Recurrent Multiattention Network for VHR Remote Sensing Image ClassificationabstractWith the advances of deep learning, many recent CNN-based methods have yielded promising results for image classification. In very high-resolution (VHR) remote sensing images, the contributions of different regions to image classification can vary significantly, because informative areas are generally limited and scattered throughout the whole image. Therefore, how to pay more attention to these informative areas and better incorporate them over long distances are two main challenges to be addressed. In this article, we propose a gated recurrent multiattention neural network (GRMA-Net) to address these problems. Because informative features generally occur at multiple stages in a network (i.e., local texture features at shallow layers and global profile features at deep layers), we use multilevel attention modules to focus on informative regions to extract more discriminative features. Then, these features are arranged as spatial sequences and fed into a deep-gated recurrent unit (GRU) to capture long-range dependency and contextual relationship. We evaluate our method on the UC Merced (UCM), Aerial Image dataset (AID), NWPU-RESISC (NWPU), and Optimal-31 (Optimal) datasets. Experimental results have demonstrated the superior performance of our method as compared to other state-of-the-art methods. Boyang Li 0007, Yulan Guo, Jun-Gang Yang, Longguang Wang, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Nonconvex Tensor Low-Rank Approximation for Infrared Small Target DetectionabstractInfrared small target detection is an important fundamental task in the infrared system. Therefore, many infrared small target detection methods have been proposed, in which the low-rank model has been used as a powerful tool. However, most low-rank-based methods assign the same weights for different singular values, which will lead to inaccurate background estimation. Considering that different singular values have different importance and should be treated discriminatively, in this article, we propose a nonconvex tensor low-rank approximation (NTLA) method for infrared small target detection. In our method, NTLA regularization adaptively assigns different weights to different singular values for accurate background estimation. Based on the proposed NTLA, we propose asymmetric spatial–temporal total variation (ASTTV) regularization to achieve more accurate background estimation in complex scenes. Compared with the traditional total variation approach, ASTTV exploits different smoothness intensities for spatial and temporal regularization. We design an efficient algorithm to find the optimal solution for our method. Compared with some state-of-the-art methods, the proposed method achieves an improvement in terms of various evaluation metrics. Extensive experimental results in various complex scenes demonstrate that our method has strong robustness and a low false-alarm rate. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yang Sun 0006, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Further Exploring Convolutional Neural Networks' Potential for Land-Use Scene ClassificationabstractRecently, with the success of deep convolutional neural networks (CNNs), many end-to-end learning algorithms have yielded excellent results. However, in the field of land use and land cover (LULC), very deep CNNs cannot be driven with even tens of thousands of images. In contrast to transferring methods that only employ a model pretrained with an irrelevant data set (e.g., ImageNet) and directly inherit parameters without refining, we explore an approach for effectively driving a deep CNN with a small data capacity. We propose a novel concept called the best activation model (BAM) in the end-to-end process for LULC image classification. BAM theoretically represents the best activation status for end-to-end networks with a small data set, taking both the data-capacity limitation and target-scene specificity into full consideration. The proposed method overcomes the problem of under-fitting and has optimal scene specificity for LULC scenes. Our approach greatly improves the time efficiency and yields excellent performance compared with state-of-the-art methods, obtaining averages of 99.0%, 98.8%, and 96.1% on the UC Merced Land-Use, WHU-RS19 data sets, and the Google data set of SIRI_WHU, respectively. Boyang Li 0007, Weihua Su, Ruihao Li 0001, Jiacheng Wei |
IEEE Geosci. Remote. Sens. Lett. | 1 |