VLDB 2026 Research / reviewers in the wild / expert
Wei An 0003
dblp:74/8455-3
· DBLP profile ↗
73ranked-venue papers
0as first author
54since 2021 · last 2026
0000-0001-8319-2105ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 16 since 2021Artificial intelligence and machine learning · 28 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 23 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Lookup NetworkabstractConvolutional neural networks are constructed with massive operations with different types and are highly computationally intensive. Among these operations, multiplication operation is higher in computational complexity and usually requires more energy consumption with longer inference time than other operations, which hinders the deployment of convolutional neural networks on mobile devices. In many resource-limited edge devices, complicated operations can be calculated via lookup tables to reduce computational cost. Motivated by this, in this paper, we introduce a generic and efficient lookup operation which can be used as a basic operation for the construction of neural networks. Instead of calculating the multiplication of weights and activation values, simple yet efficient lookup operations are adopted to compute their responses. To enable end-to-end optimization of the lookup operation, we construct the lookup tables in a differentiable manner and propose several training strategies to promote their convergence. By replacing computationally expensive multiplication operations with our lookup operations, we develop lookup networks for the image classification, image super-resolution, and point cloud classification tasks. It is demonstrated that our lookup networks can benefit from the lookup operations to achieve higher efficiency in terms of energy consumption and inference speed while maintaining competitive performance to vanilla convolutional networks. Extensive experiments show that our lookup networks produce state-of-the-art performance on different tasks (both classification and regression tasks) and different data types (both images and point clouds). Yulan Guo, Longguang Wang, Wendong Mao, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Probing Deep Into Temporal Profile Makes the Infrared Small Target Detector Much BetterabstractInfrared small target (IRST) detection is challenging in simultaneously achieving precise, robust, and efficient performance due to extremely dim targets and strong interference. Current learning-based methods attempt to leverage "more" information from both the spatial and the short-term temporal domains, but suffer from unreliable performance under complex conditions while incurring computational redundancy. In this paper, we explore the "more essential" information from a more crucial domain for the detection. Through theoretical analysis, we reveal that the global temporal saliency and correlation information in the temporal profile demonstrate significant superiority in distinguishing target signals from other signals. To investigate whether such superiority is preferentially leveraged by well-trained networks, we built the first prediction attribution tool in this field and verified the importance of the temporal profile information. Inspired by the above conclusions, we remodel the IRST detection task as a one-dimensional signal anomaly detection task, and propose an efficient deep temporal probe network (DeepPro) that only performs calculations in the time dimension for IRST detection. We conducted extensive experiments to fully validate the effectiveness of our method. The experimental results are exciting, as our DeepPro outperforms existing state-of-the-art IRST detection methods on widely-used benchmarks with extremely high efficiency, and achieves a significant improvement on dim targets and in complex scenarios. We provide a new modeling domain, a new insight, a new method, and a new performance, which can promote the development of IRST detection. Ruojing Li, Wei An 0003, Yingqian Wang 0002, Xinyi Ying, Yimian Dai, Longguang Wang, Yulan Guo, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Diving Into Epipolar Transformers for Light Field Super-Resolution and Disparity EstimationabstractLight field (LF) cameras capture the light rays of a 3D scene from multiple views simultaneously, and thus provide a more immersive experience of the real world as compared to traditional cameras. Although significant progress has been made in various LF image processing tasks, it remains challenging to effectively model the non-local spatial-angular correlations inherent in LF images, particularly when dealing with complex disparity variations. In this paper, we focus on orthogonal epipolar geometry of LF images and propose a generic Epipolar Transformer mechanism that incorporates geometrically meaningful correlations along the epipolar lines. Our Epipolar Transformer mechanism enjoys the following benefits: learning effective and diverse LF feature representations, delivering satisfactory results without redundant architectural designs, and enabling flexible extension to various LF-related tasks with simple adaptations. For LF spatial and angular super-resolution, our methods not only achieve state-of-the-art performance on benchmark datasets, but also demonstrate superior and robust performance on large disparity variations. For disparity estimation, we explore the use of geometry information encoded in our Epipolar Transformer to directly regress the disparity results, effectively avoiding the limitation of a fixed maximum disparity. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Yulan Guo, Li Liu 0002, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Dynamic High-Frequency Convolution for Infrared Small Target DetectionabstractInfrared small targets are typically tiny and locally salient, which belong to high-frequency components (HFCs) in images. Single-frame infrared small target (SIRST) detection is challenging, since there are many HFCs along with targets, such as bright corners, broken clouds, and other clutters. Current learning-based methods rely on the powerful capabilities of deep networks, but neglect explicit modeling and discriminative representation learning of various HFCs, which is important to distinguish targets from other HFCs. To address the aforementioned issues, we propose a dynamic high-frequency convolution (DHiF) to translate the discriminative modeling process into the generation of a dynamic local filter bank. Especially, DHiF is sensitive to HFCs, owing to the dynamic parameters of its generated filters being symmetrically adjusted within a zero-centered range according to Fourier transformation properties. Combining with standard convolution operations, DHiF can adaptively and dynamically process different HFC regions and capture their distinctive grayscale variation characteristics for discriminative representation learning. DHiF functions as a drop-in replacement for standard convolution and can be used in arbitrary SIRST detection networks without significant decrease in computational efficiency. To validate the effectiveness of our DHiF, we conducted extensive experiments across different SIRST detection networks on real-scene datasets. Compared to other state-of-the-art convolution operations, DHiF exhibits superior detection performance with promising improvement. Codes are available at https://github.com/TinaLRJ/DHiF. Ruojing Li, Wei An 0003, Xinyi Ying, Yingqian Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Event-Based Tiny Object Detection: A Benchmark Dataset and BaselineabstractSmall object detection (SOD) in anti-UAV task is a challenging problem due to the small size of UAVs and complex backgrounds. Traditional frame-based cameras struggle to detect small objects in complex environments due to their low frame rates, limited dynamic range, and data redundancy. Event cameras, with microsecond temporal resolution and high dynamic range, provide a more effective solution for SOD. However, existing event-based object detection datasets are limited in scale, feature large targets size, and lack diverse backgrounds, making them unsuitable for SOD benchmarks. In this paper, we introduce a Event-based Small object detection (EVSOD) dataset (namely EV-UAV), the first large-scale, highly diverse benchmark for anti-UAV tasks. It includes 147 sequences with over 2.3 million event-level annotations, featuring extremely small targets (averaging 6.8 $\times$ 5.4 pixels) and diverse scenarios such as urban clutter and extreme lighting conditions. Furthermore, based on the observation that small moving targets form continuous curves in spatiotemporal event point clouds, we propose Event based Sparse Segmentation Network (EV-SpSegNet), a novel baseline for event segmentation in point cloud space, along with a Spatiotemporal Correlation (STC) loss that leverages motion continuity to guide the network in retaining target events. Extensive experiments on the EV-UAV dataset demonstrate the superiority of our method and provide a benchmark for future research in EVSOD. The dataset and code are at https://github.com/ChenYichen9527/Ev-UAV. Yimian Dai, Shiman He, Wei An 0003 |
ICCV | 6 |
| 2025 | Triple-Directional Fusion Attention for Infrared Small Target DetectionabstractAttention mechanism has gained popularity due to its effectiveness. However, most existing mechanisms are designed for large-sized targets currently, with limited improvement in single-frame infrared small target (SIRST) detection tasks. In this letter, we propose a novel attention mechanism to enhance the extraction capacity of deep networks for infrared small targets, termed the triple-directional fusion attention module (TFAM). This module aggregates channel-, height-, and width-dimension into three independent directional perception attention vectors, preserving both accurate channel and spatial information. Through adaptive cross-direction interaction, TFAM establishes inter-directional dependencies essential for enhancing faint target signatures in deep layers. Notably, TFAM only requires minimal complexity for modeling and offers flexibility in integration. Experiments conducted on the NUDT-SIRST and NUAA-SIRST datasets demonstrate consistent improvements. Jun Chen 0007, Shipeng Zhu, Boyang Li 0007, Jianpeng Fan, Zaiping Lin, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 9 |
| 2025 | Unsupervised Degradation Representation Learning for Unpaired Restoration of Images and Point CloudsabstractRestoration tasks in low-level vision aim to restore high-quality (HQ) data from their low-quality (LQ) observations. To circumvents the difficulty of acquiring paired data in real scenarios, unpaired approaches that aim to restore HQ data solely on unpaired data are drawing increasing interest. Since restoration tasks are tightly coupled with the degradation model, unknown and highly diverse degradations in real scenarios make learning from unpaired data quite challenging. In this paper, we propose a degradation representation learning scheme to address this challenge. By learning to distinguish various degradations in the representation space, our degradation representations can extract implicit degradation information in an unsupervised manner. Moreover, to handle diverse degradations, we develop degradation-aware (DA) convolutions with flexible adaption to various degradations to fully exploit the degrdation information in the learned representations. Based on our degradation representations and DA convolutions, we introduce a generic framework for unpaired restoration tasks. Based on our framework, we propose UnIRnet and UnPRnet for unpaired image and point cloud restoration tasks, respectively. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on unpaired image and point cloud restoration tasks show that our UnIRnet and UnPRnet achieve state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Graph Laplacian regularization for fast infrared small target detection
Ting Liu 0017, Yongxian Liu, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 6 |
| 2025 | Event-Based Motion Deblurring With Blur-Aware Reconstruction FilterabstractEvent-based motion deblurring aims at reconstructing a sharp image from a single blurry image and its corresponding events triggered during the exposure time. Existing methods learn the spatial distribution of blur from blurred images, then treat events as temporal residuals and learn blurred temporal features from them, and finally restore clear images through spatio-temporal interaction of the two features. However, due to the high coupling of detailed features such as the texture and contour of the scene with blur features, it is difficult to directly learn effective blur spatial distribution from the original blurred image. In this paper, we provide a novel perspective, i.e., employing the blur indication provided by events, to instruct the network in spatially differentiated image reconstruction. Due to the consistency between event spatial distribution and image blur, event spatial indication can learn blur spatial features more simply and directly, and serve as a complement to temporal residual guidance to improve deblurring performance. Based on the above insight, we propose an event-based motion deblurring network consisting of a Multi-Scale Event-based Double Integral (MS-EDI) module designed from temporal residual guidance, and a Blur-Aware Filter Prediction (BAFP) module to conduct filter processing directed by spatial blur indication. The network, after incorporating spatial residual guidance, has significantly enhanced its generalization ability, surpassing the best-performing image-based and event-based methods on both synthetic, semi-synthetic, and real-world datasets. In addition, our method can be extended to blurry image super-resolution and achieves impressive performance. Our code is available at:https://github.com/ChenYichen9527/MBNetnow. Chushu Zhang, Wei An 0003, Longguang Wang, Qiang Ling 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Fixed Relative Pose Prior for Camera Array Self-CalibrationabstractCamera arrays have unique advantages in various computer vision tasks, such as 3D scene reconstruction and depth estimation. For these tasks, precise calibration of sub-cameras is crucial. Since the baselines of sub-cameras are usually small, it is challenging to calibrate the camera array through a single recording of the scene. Consequently, the majority of existing calibration methods address this issue by recording a scene at different spatial locations. However, this approach neglects the prior that the relative pose of the sub-cameras remains unchanged across different locations, which leads to an increase in cumulative reprojection errors. In this letter, we propose to incorporate this fixed relative pose prior to precisely calibrate the camera array. Specifically, we first capture dual-array frames by recording a scene at two spatial locations. Then, we incorporate the fixed relative pose prior to the camera array calibration process by integrating the linear constraint into the organization of sub-aperture images (SAIs). Our method maintains the minimum necessary degrees of freedom for the calibration model, and reduces cumulative reprojection error. Moreover, we develop a real-world light field dataset for comprehensive performance evaluation. Experimental results demonstrate that our method can achieve higher calibration accuracy as compared to existing methods. Our code and dataset are available athttps://github.com/Zhangyaning-NUDT/Fixed-relative-pose-prior-for-camera-array-self-calibration. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Learning Rotation-Invariant Neighbor Consensus for Mismatch Removal
Fanzhi Cao, Wei An 0003, Tianxin Shi, Yingqian Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | CWIMamba: Cross-Scale Windowed Integration State Space Model for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) intends to detect potential anomalous targets hidden in the background of hyperspectral images (HSIs) and has garnered substantial attention in various remote sensing photography and surveying applications. Recent research advances in the HAD domain have highlighted the significance of deep convolutional networks (DCNs) and vision transformers (ViTs)-based formulas. However, DCNs are long-range dependency-limited networks, whereas ViTs bear the computational burden of quadratic complexity. Owing to their prominent nonlocal representations and linear complexity, Mamba-based approaches have drawn growing attention. Our study pioneers the integration of Mamba into HAD tasks, presenting CWIMamba, which introduces a novel cross-scale windowed integration state space model for considering the spatial distribution characteristics of the anomaly targets. Specifically, we devise a cross-scale windowed state space model (CSWSSM) to scan the spatial-spectral features based on the window-based bottleneck SSM with different scales. For better multiscale feature integration, a multiscale spatial-spectral feature adaptive integration (MS3FAI) method is explored to generate an intensified representation of multiscale feature interaction and fusion based on the elaborate adaptive spatial-spectral weighting scheme. Moreover, we also devised a Haar discrete wavelet transform convolution module (HDWTCM) to fully replenish the local informative representation and enhance the discriminative frequency characteristics between anomalies and background, introducing more inductive local features for accurate background reconstruction and anomaly suppression. Extensive experiments on five multifarious HAD datasets and seven indicators substantiate the state-of-the-art detection performance, demonstrating the effectiveness of CWIMamba. Wei An 0003, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Shilin Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Nonuniformity and Bad Pixel Correction Based on 3-D Network With Oversmoothing SuppressionabstractThe purpose of non-uniformity and blind pixel correction is to provide a more reliable foundation for subsequent image processing and target detection. Existing correction methods generally struggle to balance the contradiction between over-smoothing and residual noise. Particularly, over-smoothing can easily filter out texture details and dim small targets. Based on the multi-frame response model of infrared focal plane array detector, we propose a two-stage 3-D residual fully convolutional network for correction factor estimation, integrated with an over-smoothing suppression mechanism. The proposed method designs two 3-D sub-networks to estimate the gain correction factors and offset correction factors respectively. For the correction factor pre-estimation tensors outputted by the two sub-networks, an inter-frame averaging after outlier removal is applied to suppress over-smoothing. Ultimately, using multiplication and addition structures, the final estimated values of the gain and offset correction factors can be utilized to obtain the corrected images. Experimental results indicate that the proposed method exhibits substantial generalization capabilities towards different intensities non-uniformity pixel-wise fixed mode noise and can effectively correct the blind pixels of real infrared images while suppressing over-smoothing and maintaining the image details such as dim small targets well. Overall, as a method that combines the model-driven and the data-driven, our method possesses strong theoretical interpretability and superior performance. Teliang Wang, Wei An 0003, Zaiping Lin, Kun Li 0029 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Sparsity-Aware Global Channel Pruning for Infrared Small-Target Detection NetworksabstractFor infrared small-target detection, convolutional neural network (CNN)-based methods have demonstrated promising performance. However, due to the small size of the targets, existing infrared small-target detection methods necessitate intricate structures with intensive computation to maintain distinctive features of targets in deep layers, which poses great challenges for deployment on edge devices with constrained resources. The current pruning methods predominantly focus on the optimization of classification networks, with an emphasis on semantic information. Nevertheless, the spatial details crucial for infrared small-target detection are excessively pruned in the shallow layers, resulting in a significant degradation of detection performance. In this article, based on the sparse distribution of small targets in infrared images, we propose a sparsity-aware global channel pruning (SAGCP) framework to optimize infrared small-target detection networks. Specifically, sparse modeling is used to code the target region for the first time, and sparse priors can be induced into the feature map for the identification of redundant channels. Without the need for extra structures or intricate criteria to identify redundant channels, our pruning method can leverage the inherent properties of infrared small targets to extract more robust features and obtain more compact models. When SAGCP is applied to the existing infrared small-target detection methods, the pruned network is superior in model efficiency and detection performance. For example, when applying our method to DNA-Net, the pruned model can achieve a 72.34% reduction in parameters, and a 57.49% decrease in floating point operations (FLOPs), but a 2.02% increase in intersection over union (IoU). Shuanglin Wu, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Satellite Video Object Detection Based on Enhanced 3DTV Regularization and Gaussian PriorabstractSatellite videos have played important roles in many applications in recent years due to the advantages of continuous providing high temporal resolution remote sensing images. Although much progress has been achieved for moving object detection (MOD) in satellite videos, the low-rank characteristics of background and the intensity variations of moving objects across frames have not been fully exploited. In this article, we propose an efficient method for MOD in satellite videos, which models the background with enhanced 3-D total variation (E-3DTV) regularization and the moving objects with Gaussian prior. Specifically, considering that the gradient maps on the spatial and temporal dimensions exhibit different physical meanings, we model the background with different Laplacian sparsity priors for the gradient maps along the spatial and temporal dimensions for 3DTV regularization. Different from current methods, which model moving objects with sparsity characteristics in each frame alone, we utilize Gaussian prior to model intensity changing characteristics of moving objects across frames. After integrating background model and moving object model into low-rank sparse matrix factorization framework, the alternating direction method of multipliers (ADMM) is adopted to iteratively optimize the parameters of background and moving object models. We conduct experiments on VISO and SkySat datasets, and the results demonstrate that our method achieves superior MOD performance with high computational efficiency compared to state-of-the-art methods. Wei An 0003, Ting Liu 0017, Yang Sun 0006, Zaiping Lin, Yulan Guo, Hanyun Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | Motion and Appearance Decoupling Representation for Event CamerasabstractEvent cameras, with high temporal resolution and high dynamic range, have shown great potential under extreme scenarios such as high-speed movement and low illumination. However, previous event representation methods typically aggregate event data into a single dense tensor, often overlooking the dynamic changes of events within a given time unit. This limitation can introduce historical artifacts and semantic inconsistencies, ultimately degrading model performance. Inspired by human visual prior, we propose a motion and appearance decoupling (MAD) event representation to disentangle the mixed spatial-temporal event tensor into two independent branches. This bio-inspired design helps the network extract discriminative temporal (i.e., motion) and spatial (i.e., appearance) information, thus reducing the network's learning burden toward complex high-level interpretation tasks. In our method, the event motion guided attention module (EMGA) is designed to achieve temporal and spatial feature interaction and fusion sequentially. Based on EMGA, three specially designed decoder heads are proposed for several representative event-based tasks (i.e., object detection, semantic segmentation, and human pose estimation). Experimental results demonstrate that our method achieves state-of-the-art performance on the above three tasks, which reveals that our method is an easy-to-implement replacement for currently event-based methods. Our code is available at: https://github.com/ChenYichen9527/MAD-representation. Boyang Li 0007, Yingqian Wang 0002, Xinyi Ying, Longguang Wang, Chushu Zhang, Yulan Guo, Wei An 0003 |
IEEE Trans. Image Process. | 9 |
| 2025 | Direction-Coded Temporal U-Shape Module for Multiframe Infrared Small Target DetectionabstractInfrared small target (IRST) detection aims at separating targets from cluttered background. Although many deep learning-based single-frame IRST (SIRST) detection methods have achieved promising detection performance, they cannot deal with extremely dim targets while suppressing the clutters since the targets are spatially indistinctive. Multiframe IRST (MIRST) detection can well handle this problem by fusing the temporal information of moving targets. However, the extraction of motion information is challenging since general convolution is insensitive to motion direction. In this article, we propose a simple yet effective direction-coded temporal U-shape module (DTUM) for MIRST detection. Specifically, we build a motion-to-data mapping to distinguish the motion of targets and clutters by indexing different directions. Based on the motion-to-data mapping, we further design a direction-coded convolution block (DCCB) to encode the motion direction into features and extract the motion information of targets. Our DTUM can be equipped with most single-frame networks to achieve MIRST detection. Moreover, in view of the lack of MIRST datasets, including dim targets, we build a multiframe infrared small and dim target dataset (namely, NUDT-MIRSDT) and propose several evaluation metrics. The experimental results on the NUDT-MIRSDT dataset demonstrate the effectiveness of our method. Our method achieves the state-of-the-art performance in detecting infrared small and dim targets and suppressing false alarms. Our codes will be available at https://github.com/TinaLRJ/Multi-frame-infrared-small-target-detection-DTUM. Ruojing Li, Wei An 0003, Boyang Li 0007, Yingqian Wang 0002, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Real-World Light Field Image Super-Resolution Via Degradation ModulationabstractRecent years have witnessed the great advances of deep neural networks (DNNs) in light field (LF) image super-resolution (SR). However, existing DNN-based LF image SR methods are developed on a single fixed degradation (e.g., bicubic downsampling), and thus cannot be applied to super-resolve real LF images with diverse degradation. In this article, we propose a simple yet effective method for real-world LF image SR. In our method, a practical LF degradation model is developed to formulate the degradation process of real LF images. Then, a convolutional neural network is designed to incorporate the degradation prior into the SR process. By training on LF images using our formulated degradation, our network can learn to modulate different degradation while incorporating both spatial and angular information in LF images. Extensive experiments on both synthetically degraded and real-world LF images demonstrate the effectiveness of our method. Compared with existing state-of-the-art single and LF image SR methods, our method achieves superior SR performance under a wide range of degradation, and generalizes better to real LF images. Codes and models are available at https://yingqianwang.github.io/LF-DMnet/. Yingqian Wang 0002, Zhengyu Liang, Longguang Wang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite VideosabstractMoving object detection in satellite videos (SVMOD) is a challenging task due to the extremely dim and small target characteristics. Current learning-based methods extract spatio-temporal information from multi-frame dense representation with labor-intensive manual labels to tackle SVMOD, which needs high annotation costs and contains tremendous computational redundancy due to the severe imbalance between foreground and background regions. In this paper, we propose a highly efficient unsupervised framework for SVMOD. Specifically, we propose a generic unsupervised framework for SVMOD, in which pseudo labels generated by a traditional method can evolve with the training process to promote detection performance. Furthermore, we propose a highly efficient and effective sparse convolutional anchor-free detection network by sampling the dense multi-frame image form into a sparse spatio-temporal point cloud representation and skipping the redundant computation on background regions. Coping these two designs, we can achieve both high efficiency (label and computation efficiency) and effectiveness. Extensive experiments demonstrate that our method can not only process 98.8 frames per second on 1024 ×1024 images but also achieve state-of-the-art performance. Wei An 0003, Yifan Zhang 0030, Zhuo Su 0002, Weidong Sheng, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Mixed-Precision Network Quantization for Infrared Small Target SegmentationabstractNetwork quantization is leveraged to reduce the model size, memory footprint, and computational cost of deep neural networks. It is achieved by representing float weights and activations with lower bit counterparts, which is essential for model deployment on resource-limited devices. However, due to the extremely small size of infrared small targets in the feature map, low-bit quantization could lead to huge information loss of small targets and thus causes severe segmentation performance degradation. To achieve low-bit quantization while maintaining the segmentation performance, we first study the quantization sensitivity of small target segmentation network and observe the sensitivity heterogeneity of different layers in the network. Specifically, feature maps in shallow layers and encoder subnetwork are more vulnerable to information loss caused by quantization as compared to deep layers and decoder subnetwork. Based on these observations, we are motivated to assign a different bitwidth for each block according to their quantization sensitivity. A simple yet effective symmetrically progressive decreasing mixed-precision quantization (SPMix-Q) method is proposed to achieve high-performance segmentation under low-bit quantization (i.e., 2.42 bits for weights and 3.82 bits for activations). The experimental results show that our SPMix-Q achieves comparable accuracy with only 1/13 model size, 1/4.6 memory footprint, and 1/29 computational cost to the full-precision counterparts. Compared with the homogeneous low-bit quantization methods, our method achieves much better performance in terms of intersection of union (IoU) on the benchmark datasets. Our mobile-system-on-a-chip (SOC) (e.g., Kyrin 980, Snapdragon 660, and Dimensity 800U) deployable android application package (APK) is available at:https://github.com/YeRen123455/SIRST-Quantization-Deployment. Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Tianhao Wu 0014, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | DDAug: Differentiable Data Augmentation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation(WSSS) with image-level labels has witnessed promising advances with the help ofclass activation maps(CAM). However, CAM is always confined to small discriminative seed regions due to its simple classification loss guided training manner. To handle this problem, recent works introduced specifically designed regularizations and modules to expand the CAM seed regions, serving as the final segmentation masks. In this paper, we surprisingly find that the classification loss could suppress the gains from these regularization and modules in the late training phase, thereby limiting the further growth of CAM, which we call as theexplicit supervision disturb(ESD) issue. Interestingly, we find that specificdata augmentation(DA) operations (e.g., CutMix) can relieve such ESD issue, and the benefits introduced by different DA operations vary a lot. To maximize the benefits, we proposedifferentiable data augmentation(DDAug) to automatically search for the proper DA policy. Specifically, we design amulti-level search spaceto sequentially sample DA operations with different properties. Extensive experiments demonstrate that the proposed DDAug can alleviate the ESD issue and introduce consistent improvements to various popular WSSS methods, achieving the state-of-the-art performance on the MS COCO 2014 and PASCAL VOC 2012 datasets. Boyang Li 0007, Fei Zhang 0016, Longguang Wang, Yingqian Wang 0002, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Multim. | 7 |
| 2024 | Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T VideosabstractTracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on singlemodality scenes, and overlook the complementary characteristics of tiny objects captured by multiple remote sensors. To enhance tracking performance by integrating complementary information from multiple sources, we propose a novel framework called HGT-Track (Heterogeneous Graph Transformer based Multi-Tiny-Object Tracking). Specifically, we first employ a Transformer-based encoder to embed images from different modalities. Subsequently, we utilize Heterogeneous Graph Transformer to aggregate spatial and temporal information from multiple modalities to generate detection and tracking features. Additionally, we introduce a target re-detection module (ReDet) to ensure tracklet continuity by maintaining consistency across different modalities. Furthermore, this paper introduces the first benchmark VT-Tiny-MOT (Visible-Thermal Tiny MultiObject Tracking) for RGB-T fused multiple tiny object tracking. Extensive experiments are conducted on VT-Tiny-MOT, and the results have demonstrated the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of MOTA (Multiple-Object Tracking Accuracy) and ID-F1 score. The code and dataset will be made available at https://github.com/xuqingyu26/HGTMT Longguang Wang, Weidong Sheng, Yingqian Wang 0002, Chao Ma 0014, Wei An 0003 |
IEEE Trans. Multim. | 7 |
| 2023 | Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-level annotations. To reduce the annotation burden, we propose the first method to achieve SIRST detection with single-point supervision. The core idea of this work is to recover the per-pixel mask of each target from the given single point label by using clustering approaches, which looks simple but is indeed challenging since targets are always insalient and accompanied with background clutters. To handle this issue, we introduce randomness to the clustering process by adding noise to the input images, and then obtain much more reliable pseudo masks by averaging the clustered results. Thanks to this "Monte Carlo" clustering approach, our method can accurately recover pseudo masks and thus turn arbitrary fully supervised SIRST detection networks into weakly supervised ones with only single point annotation. Experiments on four datasets demonstrate that our method can be applied to existing SIRST detection networks to achieve comparable performance with their fully-supervised counterparts, which reveals that single-point supervision is strong enough for SIRST detection. Our code will be available at: https://github.com/YeRen123455/SIRST-Single-Point-Supervision. Boyang Li 0007, Yingqian Wang 0002, Longguang Wang, Fei Zhang 0016, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
ICCV | 7 |
| 2023 | RDFM: Robust Deep Feature Matching for Multimodal Remote-Sensing ImagesabstractRobust feature matching for multimodal remote sensing images remains challenging due to the significant nonlinear radiation difference (NRD) caused by modality variations. In this letter, we present a novel feature-matching method for multimodal remote sensing images, called RDFM, which exploits only deep features extracted by a pre-trained VGG network to achieve competitive performance. It is shown that template matching of these pre-trained features is robust to NRD for various multi-modal remote sensing images, and no additional training is required to improve the matching performance. In order to extract as many correspondences as possible, we use dense template matching to obtain point correspondences and introduce a 4D convolution-based implementation of dense template matching for the sake of computational efficiency. RDFM consists of two main steps. First, enormous coarse correspondences are extracted by applying dense template matching at the deep layer of the pre-trained network, and then a coarse-to-fine hierarchical refinement is performed to obtain high-quality correspondences. To verify the effectiveness of RDFM, six different types of multimodal image datasets are used in our experiments, including day-night, depth-optical, infrared-optical, map-optical, optical-optical, and SAR-optical datasets. The comprehensive experimental results show that RDFM is able to overcome the problem caused by NRD and achieves a better performance than the state-of-the-art methods for multimodal remote sensing image matching. The code of RDFM is publicly available at https://github.com/Fans2017/RDFM. Fanzhi Cao, Tianxin Shi, Kaiyang Han, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Disentangling Light Fields for Super-Resolution and Disparity EstimationabstractLight field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been proposed for various LF image processing tasks. However, it is challenging for CNNs to effectively process LF images since the spatial and angular information are highly inter-twined with varying disparities. In this paper, we propose a generic mechanism to disentangle these coupled information for LF image processing. Specifically, we first design a class of domain-specific convolutions to disentangle LFs from different dimensions, and then leverage these disentangled features by designing task-specific modules. Our disentangling mechanism can well incorporate the LF structure prior and effectively handle 4D LF data. Based on the proposed mechanism, we develop three networks (i.e., DistgSSR, DistgASR and DistgDisp) for spatial super-resolution, angular super-resolution and disparity estimation. Experimental results show that our networks achieve state-of-the-art performance on all these three tasks, which demonstrates the effectiveness, efficiency, and generality of our disentangling mechanism. Project page: https://yingqianwang.github.io/DistgLF/. Yingqian Wang 0002, Longguang Wang, Gaochang Wu, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Exploring Fine-Grained Sparsity in Convolutional Neural Networks for Efficient InferenceabstractNeural networks contain considerable redundant computation, which drags down the inference efficiency and hinders the deployment on resource-limited devices. In this paper, we study the sparsity in convolutional neural networks and propose a generic sparse mask mechanism to improve the inference efficiency of networks. Specifically, sparse masks are learned in both data and channel dimensions to dynamically localize and skip redundant computation at a fine-grained level. Based on our sparse mask mechanism, we develop SMPointSeg, SMSR, and SMStereo for point cloud semantic segmentation, single image super-resolution, and stereo matching tasks, respectively. It is demonstrated that our sparse masks are well compatible to different model components and network architectures to accurately localize redundant computation, with computational cost being significantly reduced for practical speedup. Extensive experiments show that our SMPointSeg, SMSR, and SMStereo achieve state-of-the-art performance on benchmark datasets in terms of both accuracy and efficiency. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Combining Deep Denoiser and Low-rank Priors for Infrared Small Target DetectionabstractMany existing low-rank methods have achieved good detection performance in uniform scenes, but they suffer from a high false alarm rate in complex noisy scenes. Therefore, it is important to improve the detection performance of low-rank models in noisy scenes. In this paper, we first formulate an implicit regularizer by plugging a denoising neural network (termed as deep denoiser), which can learn deep image priors from a large number of natural images. Then, we use the weighted sum of weighted tensor nuclear norm for more accurate background estimation. Finally, alternating direction multiplier method is used to solve the model under the plug-and-play framework. By integrating low-rank prior with deep denoiser prior, our model achieves higher accuracy. Experiments on different scenes demonstrate that our method achieves an improved performance in terms of visual effects and quantitative metrics. Specially, the overall accuracy of AUC value (AUCOA) achieved by the proposed method on Sequences 1-6 are 1.24%, 1.16%, 0.63%, 1.9%, 0.82%, 2.06% higher than those achieved by the second top performing methods, respectively. Ting Liu 0017, Jun-Gang Yang, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 5 |
| 2023 | Learning scalable dynamic filter in convolutional networks
Shuanglin Wu, Xinyi Ying, Longguang Wang, Jun-Gang Yang, Wei An 0003 |
Pattern Recognit. Lett. | 6 |
| 2023 | You Only Train Once: Learning a General Anomaly Enhancement Network With Random Masks for Hyperspectral Anomaly DetectionabstractIn this paper, we introduce a new approach to address the challenge of generalization in hyperspectral anomaly detection (AD). Our method eliminates the need for adjusting parameters or retraining on new test scenes as required by most existing methods. Employing an image-level training paradigm, we achieve a general anomaly enhancement network for hyperspectral AD that only needs to be trained once. Trained on a set of anomaly-free hyperspectral images with random masks, our network can learn the spatial context characteristics between anomalies and background in an unsupervised way. Additionally, a plug-and-play model selection module is proposed to search for a spatial-spectral transform domain that is more suitable for AD task than the original data. To establish a unified benchmark to comprehensive evaluate our method and existing methods, we develop a large-scale hyperspectral AD dataset (HAD100) that includes 100 real test scenes with diverse anomaly targets. In comparison experiments, we combine our network with a parameter-free detector, and achieve the optimal balance between detection accuracy and inference speed among state-of-the-art AD methods. Experimental results also show that our method still achieves competitive performance when the training and test set are captured by different sensor devices. Our code is available at https://github.com/ZhaoxuLi123/AETNet. Zhaoxu Li, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Infrared Small Target Detection via Nonconvex Tensor Tucker Decomposition With Factor PriorabstractInfrared small target detection in complex scenes is an important but challenging research hotspot in infrared early warning fields. Previous studies have proved that low-rank Tucker decomposition (TD) achieves good detection performance in complex scenes. However, a key limitation of existing low-rank TD methods is that the rank needs to be set in advance, and an inaccurate predefined rank can lead to performance degradation. Inspired by the theorem that n-rank is upper bounded by the rank of each Tucker factor matrix, we propose a nonconvex tensor TD model with factor prior for infrared small target detection. In our method, we use a logdet-based function to constrain the latent factors of low-rank TD, which avoids empirical rank selection and sufficiently uses the latent data structure information in the factor matrix. Meanwhile, performing singular value decomposition (SVD) calculations on small factor matrices can reduce computational complexity. Then, group sparsity regularized total variation is used to better exploit the shared sparse pattern of difference images, which helps better remove background clutter and obtain better detection results. Finally, the proposed method is efficiently solved by the well-designed alternating direction method of multipliers (ADMM). Extensive experimental results demonstrate that our method is more effective and robust in complex scenes than other state-of-the-art methods. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Representative Coefficient Total Variation for Efficient Infrared Small Target DetectionabstractLow-rank and sparse decomposition based models are powerful and robust tools for infrared small target detection. However, due to the calculation of singular value decomposition (SVD) and the optimization of complex regularization terms, existing low-rank models often suffer from high computational complexity. To solve this problem, based on the theorem that representative coefficient matrix obtained by orthogonal transformation of data matrix can inherit the spatial structure of data matrix, we propose a representative coefficient total variation (RCTV) method for efficient infrared small target detection. In our method, we use total variational to constraint representative coefficient matrix instead of data matrix to describe local smooth prior, which helps remove noise and reduce computational complexity. Meanwhile, we control the number of columns in the representative coefficient matrix to maintain the low-rank characteristics of background, which avoids SVD calculation and improves detection efficiency. Therefore, the RCTV regularization can simultaneously describe local smooth prior and low-rank prior. Moreover, to better enhance the sparsity of targets and distinguish sparse non-target points, we use the log-sum function to adaptively assign weights to targets. It helps obtain more accurate detection performance. The proposed model is efficiently solved by the alternating direction multiplier method (ADMM). A large number of experiments show that the proposed method is superior to existing low-rank methods in both detection accuracy and efficiency. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | MTU-Net: Multilevel TransUNet for Space-Based Infrared Tiny Ship DetectionabstractSpace-based infrared tiny ship detection aims at separating tiny ships from the images captured by Earth-orbiting satellites. Due to the extremely large image coverage area (e.g., thousands of square kilometers), candidate targets in these images are much smaller, dimer, and more changeable than those targets observed by aerial- and land-based imaging devices. Existing short imaging distance-based infrared datasets and target detection methods cannot be well adopted to the space-based surveillance task. To address these problems, we develop a space-based infrared tiny ship detection dataset (namely, NUDT-SIRST-Sea) with 48 space-based infrared images and$17\,598$pixel-level tiny ship annotations. Each image covers about$10\,000$km2of area with$10 \ 000\,\, \times \ 10 \ 000$pixels. Considering the extreme characteristics (e.g., small, dim, and changeable) of those tiny ships in such challenging scenes, we propose a multilevel TransUNet (MTU-Net) in this article. Specifically, we design a vision Transformer (ViT) convolutional neural network (CNN) hybrid encoder to extract multilevel features. Local feature maps are first extracted by several convolution layers and then fed into the multilevel feature extraction module [multilevel ViT module (MVTM)] to capture long-distance dependency. We further propose a copy–rotate–resize–paste (CRRP) data augmentation approach to accelerate the training phase, which effectively alleviates the issue of sample imbalance between targets and background. Besides, we design a FocalIoU loss to achieve both target localization and shape description. Experimental results on the NUDT-SIRST-Sea dataset show that our MTU-Net outperforms traditional and existing deep learning-based single-frame infrared small target (SIRST) methods in terms of probability of detection, false alarm rate, and intersection over union. Our code is available athttps://github.com/TianhaoWu16/Multi-level-TransUNet-for-Space-based-Infrared-Tiny-ship-Detection Tianhao Wu 0014, Boyang Li 0007, Yihang Luo, Yingqian Wang 0002, Ting Liu 0017, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | RepISD-Net: Learning Efficient Infrared Small-Target Detection Network via Structural Re-ParameterizationabstractInfrared small target detection is a challenging task for deep learning-based methods because targets tend to disappear in the deep layers. To handle this problem, existing deep neural networks usually apply various dense and skip connections for feature maintenance. Although these well-designed networks have achieved good detection performance, the complex network structures reduce their efficiency. In this paper, we propose a simple yet efficient network (RepISD-Net) for infrared small target detection. The core of our RepISD-Net is to use different network architectures but equivalent model parameters for training and inference, respectively. Specifically, in the training phase, we design a parallel multi-branch edge compensation block (ECB) to enhance the local salient features and capture finer contour characteristic of infrared small targets. In the inference phase, the multi-branch topology structures are merged into a single branch with only cascaded 3×3 convolutions for fast inference. We conduct extensive experiments on several public datasets to validate the effectiveness of our method. Experimental results demonstrate that our RepISD-Net can achieve comparable or even better detection performance with significant acceleration in inference speed as compared to state-of-the-art infrared small target detection methods. Code is submitted for review and will be released upon acceptance. Shuanglin Wu, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Incorporating Deep Background Prior Into Model-Based Method for Unsupervised Moving Vehicle Detection in Satellite VideosabstractBackground reconstruction is a key step of moving object detection in satellite videos. Most existing model-based methods exploit low-rank prior to recover background, which have achieved good performance but suffered degradation under complex and dynamic scenes. In this paper, we introduce a deep background prior into model-based methods for moving vehicle detection in satellite videos. Our deep background prior is obtained by a background reconstruction network, which can learn to reconstruct background from consecutive frames. By applying our deep background prior into model-based methods, a closed-form solution can be obtained via alternating direction method of multipliers (ADMM) and then detection results can be acquired through iterative optimization. More importantly, our background reconstruction network can be trained in an unsupervised way by introducing specifically designed loss, thus relieving the dependence on large-scale labeled dataset. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed method. Ting Liu 0017, Xinyi Ying, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Dense Nested Attention Network for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds. With the advances of deep learning, CNN-based methods have yielded promising results in generic object detection due to their powerful modeling capability. However, existing CNN-based methods cannot be directly applied to infrared small targets since pooling layers in their networks could lead to the loss of targets in deep layers. To handle this problem, we propose a dense nested attention network (DNA-Net) in this paper. Specifically, we design a dense nested interactive module (DNIM) to achieve progressive interaction among high-level and low-level features. With the repetitive interaction in DNIM, the information of infrared small targets in deep layers can be maintained. Based on DNIM, we further propose a cascaded channel and spatial attention module (CSAM) to adaptively enhance multi-level features. With our DNA-Net, contextual information of small targets can be well incorporated and fully exploited by repetitive fusion and enhancement. Moreover, we develop an infrared small target dataset (namely, NUDT-SIRST) and propose a set of evaluation metrics to conduct comprehensive performance evaluation. Experiments on both public and our self-developed datasets demonstrate the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of probability of detection (${P}_{d}$), false-alarm rate (${F}_{a}$), and intersection of union ($IoU$). Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 7 |
| 2022 | Occlusion-Aware Cost Constructor for Light Field Depth EstimationabstractMatching cost construction is a key step in light field (LF) depth estimation, but was rarely studied in the deep learning era. Recent deep learning-based LF depth estimation methods construct matching cost by sequentially shifting each sub-aperture image (SAI) with a series of pre-defined offsets, which is complex and time-consuming. In this paper, we propose a simple and fast cost constructor to construct matching cost for LF depth estimation. Our cost constructor is composed by a series of convolutions with specifically designed dilation rates. By applying our cost constructor to SAI arrays, pixels under predefined disparities can be integrated and matching cost can be constructed without using any shifting operation. More importantly, the proposed cost constructor is occlusion-aware and can handle occlusions by dynamically modulating pixels from different views. Based on the proposed cost constructor, we develop a deep network for LF depth estimation. Our network ranks first on the commonly used 4D LF benchmark in terms of the mean square error (MSE), and achieves a faster running time than other state-of-the-art methods. Yingqian Wang 0002, Longguang Wang, Zhengyu Liang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 5 |
| 2022 | Learnable Lookup Table for Neural Network QuantizationabstractNeural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(·) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use predefined functions (e.g., exponential function) with learnable parameters to build the quantizer for joint optimization. However, these complicated quantizers introduce considerable computational overhead during inference since activation quantization should be conducted online. In this paper, we formulate the quantization process as a simple lookup operation and propose to learn lookup tables as quantizers. Specifically, we develop differentiable lookup tables and introduce several training strategies for optimization. Our lookup tables can be trained with the network in an end-to-end manner to fit the distributions in different layers and have very small additional computational cost. Comparison with previous methods show that quantized networks using our lookup tables achieve state-of-the-art performance on image classification, image super-resolution, and point cloud classification tasks. Longguang Wang, Yingqian Wang 0002, Li Liu 0002, Wei An 0003, Yulan Guo |
CVPR | 5 |
| 2022 | RSMOT: Remote Sensing Multi-Object Tracking Network with Local Motion Prior for Objects in Satellite VideosabstractMulti-object tracking (MOT) in satellite videos is a new and challenging task. The difficulties stem from the extremely small objects and the low contrast between objects and background. To tackle the challenges of MOT in satellite videos, a multi-object tracking method is proposed in this paper to incorporate the local motion prior into the network. Specifically, we design a local cost volume construction module to obtain tracking offsets between adjacent frames. Based on the tracking offsets, features of previous frames can be propagated to the current frame to incorporate spatio-temporal information. We conduct extensive experiments on videos from Jilin-1 satellite, and the results demonstrate the effectiveness of the proposed method. Shuanglin Wu, Yingqian Wang 0002, Wei An 0003 |
IGARSS | 5 |
| 2022 | DSFNet: Dynamic and Static Fusion Network for Moving Object Detection in Satellite VideosabstractMoving object detection (MOD) in satellite videos remains challenging due to the extremely small size of the interested targets and the highly complex background. Both the intra-frame (static) and inter-frame (dynamic) information are of great importance to MOD. In this letter, we propose a two-stream detection network named dynamic and static fusion network (DSFNet) to tackle the MOD problem in satellite videos. Specifically, the DSFNet is composed of a 2-D backbone to extract static context information from a single frame and a lightweight 3-D backbone to extract dynamic motion cues from consecutive frames. Then the extracted static and dynamic features are fused and fed into the detection head to detect the moving targets in satellite videos. We conduct extensive experiments on videos collected from Jilin-1 satellite and the results have demonstrated the effectiveness and robustness of the proposed DSFNet. Experimental results show that our DSFNet achieves the-state-of-the-art performance. Xinyi Ying, Ruojing Li, Shuanglin Wu, Li Liu 0002, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2022 | Moving Object Detection in Satellite Videos via Spatial-Temporal Tensor Model and Weighted Schatten p-Norm MinimizationabstractLow-rank matrix decomposition approaches have achieved significant progress in small and dim object detection in satellite videos. However, it is still challenging to achieve robust performance and fast processing under complex and highly heterogeneous backgrounds since satellite video data can neither adequately fit the foreground structure nor the background model in the existing matrix decomposition models. In this letter, we propose a novel object detection method based on a spatial–temporal tensor data structure. First, we construct a tensor data structure to exploit the inner spatial and temporal correlation within a satellite video. Second, we extend the decomposition formulation with bounded noise to achieve robust performance under complex backgrounds. This formulation integrates low-rank background, structured sparse foreground, and their noises into a tensor decomposition problem. For background separation, a weighted Schatten$p$-norm is incorporated to provide adaptive threshold to obtain the singular value of the background tensor. Finally, the proposed model is solved using the alternative direction method of multipliers (ADMM) scheme. Experimental results on various real scenes demonstrate the superiority of the proposed method against the compared approaches. Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Parallax Attention for Unsupervised Stereo Correspondence LearningabstractStereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Coarse-to-fine pseudo supervision guided meta-task optimization for few-shot object classification
Yawen Cui, Qing Liao 0001, Dewen Hu, Wei An 0003, Li Liu 0002 |
Pattern Recognit. | 4 |
| 2022 | Dense Dual-Attention Network for Light Field Image Super-ResolutionabstractLight field (LF) images can be used to improve the performance of image super-resolution (SR) because both angular and spatial information is available. It is challenging to incorporate distinctive information from different views for LF image SR. Moreover, the long-term information from the previous layers can be weakened as the depth of network increases. In this paper, we propose a dense dual-attention network for LF image SR. Specifically, we design a view attention module to adaptively capture discriminative features across different views and a channel attention module to selectively focus on informative information across all channels. These two modules are fed to two branches and stacked separately in a chain structure for adaptive fusion of hierarchical features and distillation of valid information. Meanwhile, a dense connection is used to fully exploit multi-level information. Extensive experiments demonstrate that our dense dual-attention mechanism can capture informative information across views and channels to improve SR performance. Comparative results show the advantage of our method over state-of-the-art methods on public datasets. Yu Mo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Gated Recurrent Multiattention Network for VHR Remote Sensing Image ClassificationabstractWith the advances of deep learning, many recent CNN-based methods have yielded promising results for image classification. In very high-resolution (VHR) remote sensing images, the contributions of different regions to image classification can vary significantly, because informative areas are generally limited and scattered throughout the whole image. Therefore, how to pay more attention to these informative areas and better incorporate them over long distances are two main challenges to be addressed. In this article, we propose a gated recurrent multiattention neural network (GRMA-Net) to address these problems. Because informative features generally occur at multiple stages in a network (i.e., local texture features at shallow layers and global profile features at deep layers), we use multilevel attention modules to focus on informative regions to extract more discriminative features. Then, these features are arranged as spatial sequences and fed into a deep-gated recurrent unit (GRU) to capture long-range dependency and contextual relationship. We evaluate our method on the UC Merced (UCM), Aerial Image dataset (AID), NWPU-RESISC (NWPU), and Optimal-31 (Optimal) datasets. Experimental results have demonstrated the superior performance of our method as compared to other state-of-the-art methods. Boyang Li 0007, Yulan Guo, Jun-Gang Yang, Longguang Wang, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Spectral-Spatial Deep Support Vector Data Description for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to distinguish anomalies from background-by-background modeling. Deep learning has been applied to HAD and achieves promising detection results. However, there exist several issues that need to be addressed: 1) unrealistic Gaussian assumption on the latent representations may limit its application; 2) deep features are not well-suited to anomaly detection due to the separation between feature learning and anomaly detection; 3) lack of adequate exploitation of spectral-spatial features; 4) negative effect caused by spectral band redundancy. In this article, we propose an end-to-end trainable deep one-class classification network for HAD. Specifically, a minimal enclosing hypersphere is trained to involve the deep features of background samples. These background samples are selected by a density clustering-based method. In this way, feature learning and anomaly detection are incorporated into a unified framework. Meanwhile, there is no explicit Gaussian assumption on the background features. Moreover, due to the complementarity of spectral and spatial features, a novel feature fusion strategy is proposed to fuse spectral and spatial features extracted by a two-stream deep convolutional autoencoder network. Finally, a band attention module is used to automatically learn small weights for redundant bands and thus reduce the negative effect caused by redundant bands. Experimental results on five public datasets demonstrate the superiority of the proposed method compared to several state-of-the-art HAD methods in the detection performance. Kun Li 0029, Qiang Ling 0002, Yao Qin 0002, Yingqian Wang 0002, Yaoming Cai, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Nonconvex Tensor Low-Rank Approximation for Infrared Small Target DetectionabstractInfrared small target detection is an important fundamental task in the infrared system. Therefore, many infrared small target detection methods have been proposed, in which the low-rank model has been used as a powerful tool. However, most low-rank-based methods assign the same weights for different singular values, which will lead to inaccurate background estimation. Considering that different singular values have different importance and should be treated discriminatively, in this article, we propose a nonconvex tensor low-rank approximation (NTLA) method for infrared small target detection. In our method, NTLA regularization adaptively assigns different weights to different singular values for accurate background estimation. Based on the proposed NTLA, we propose asymmetric spatial–temporal total variation (ASTTV) regularization to achieve more accurate background estimation in complex scenes. Compared with the traditional total variation approach, ASTTV exploits different smoothness intensities for spatial and temporal regularization. We design an efficient algorithm to find the optimal solution for our method. Compared with some state-of-the-art methods, the proposed method achieves an improvement in terms of various evaluation metrics. Extensive experimental results in various complex scenes demonstrate that our method has strong robustness and a low false-alarm rate. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yang Sun 0006, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A BenchmarkabstractSatellite video cameras can provide continuous observation for a large-scale area, which is important for many remote sensing applications. However, achieving moving object detection and tracking in satellite videos remains challenging due to the insufficient appearance information of objects and lack of high-quality datasets. In this article, we first build a large-scale satellite video dataset with rich annotations for the task of moving object detection and tracking. This dataset is collected by the Jilin-1 satellite constellation and composed of 47 high-quality videos with 1 646 038 instances of interest for object detection and 3711 trajectories for object tracking. We then introduce a motion modeling baseline to improve the detection rate and reduce false alarms based on accumulative multiframe differencing and robust matrix completion. Finally, we establish the first public benchmark for moving object detection and tracking in satellite videos and extensively evaluate the performance of several representative approaches on our dataset. Comprehensive experimental analyses and insightful conclusions are also provided. The dataset is available athttps://github.com/QingyongHu/VISO. Qingyong Hu, Hao Liu 0061, Feng Zhang 0046, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | Exploring Sparsity in Image Super-Resolution for Efficient InferenceabstractCurrent CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for those flat regions. Therefore, existing CNN-based methods involve redundant computation in flat regions, which increases their computational cost and limits their applications on mobile devices. In this paper, we explore the sparsity in image SR to improve inference efficiency of SR networks. Specifically, we develop a Sparse Mask SR (SMSR) network to learn sparse masks to prune redundant computation. Within our SMSR, spatial masks learn to identify "important" regions while channel masks learn to mark redundant channels in those "unimportant" regions. Consequently, redundant computation can be accurately localized and skipped while maintaining comparable performance. It is demonstrated that our SMSR achieves state-of-the-art performance with 41%/33%/27% FLOPs being reduced for ×2/3/4 SR. Code is available at: https://github.com/LongguangWang/SMSR. Longguang Wang, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003, Yulan Guo |
CVPR | 6 |
| 2021 | Unsupervised Degradation Representation Learning for Blind Super-ResolutionabstractMost existing CNN-based super-resolution (SR) methods are developed based on an assumption that the degradation is fixed and known (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation is different from their assumption. To handle various unknown degradations in real-world applications, previous methods rely on degradation estimation to reconstruct the SR image. Nevertheless, degradation estimation methods are usually time-consuming and may lead to SR failure due to large estimation errors. In this paper, we propose an unsupervised degradation representation learning scheme for blind SR without explicit degradation estimation. Specifically, we learn abstract representations to distinguish various degradations in the representation space rather than explicit estimation in the pixel space. Moreover, we introduce a Degradation-Aware SR (DASR) network with flexible adaption to various degradations based on the learned representations. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on both synthetic and real images show that our network achieves state-of-the-art performance for the blind SR task. Code is available at: https://github.com/LongguangWang/DASR. Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 6 |
| 2021 | Learning A Single Network for Scale-Arbitrary Super-ResolutionabstractRecently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., ×2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to learn a scale-arbitrary image SR network from scale-specific networks. Specifically, we develop a plug-in module for existing SR networks to perform scale-arbitrary SR, which consists of multiple scale-aware feature adaption blocks and a scale-aware upsampling layer. Moreover, conditional convolution is used in our plug-in module to generate dynamic scale-aware filters, which enables our network to adapt to arbitrary scale factors. Our plug-in module can be easily adapted to existing networks to realize scale-arbitrary SR with a single model. These networks plugged with our module can produce promising results for non-integer and asymmetric SR while maintaining state-of-the-art performance for SR with integer scale factors. Besides, the additional computational and memory cost of our module is very small. Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
ICCV | 5 |
| 2021 | Infrared Dim and Small Target Detection via Multiple Subspace Learning and Spatial-Temporal Patch-Tensor ModelabstractRobust detection of infrared small and dim targets with highly heterogeneous backgrounds plays an indispensable role in infrared search and tracking (IRST) system, which is still a challenging problem. In this article, a novel method based on multisubspace learning and spatial-temporal tensor data structure is aimed to solve this problem. First, a tensor data structure is constructed to use inner correlation in spatial and temporal domain of the infrared image sequence. Second, in consideration of the complex and heterogeneous backgrounds in infrared images, the proposed method promotes the multisubspace property to tensor domain to separate the target and the background more accurately. Finally, an efficient and effective optimization algorithm based on alternating direction method of multipliers (ADMM) is designed to solve this problem. Experimental results on various and real scenes demonstrate the superiority of the proposed method compared to other five baseline methods. Yang Sun 0006, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Light Field Image Super-Resolution Using Deformable ConvolutionabstractLight field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among LF images. In this paper, we propose a deformable convolution network (i.e., LF-DFnet) to handle the disparity problem for LF image SR. Specifically, we design an angular deformable alignment module (ADAM) for feature-level alignment. Based on ADAM, we further propose a collect-and-distribute approach to perform bidirectional alignment between the center-view feature and each side-view feature. Using our approach, angular information can be well incorporated and encoded into features of each view, which benefits the SR reconstruction of all LF images. Moreover, we develop a baseline-adjustable LF dataset to evaluate SR performance under different disparity variations. Experiments on both public and our self-developed datasets have demonstrated the superiority of our method. Our LF-DFnet can generate high-resolution images with more faithful details and achieve state-of-the-art reconstruction accuracy. Besides, our LF-DFnet is more robust to disparity variations, which has not been well addressed in literature. Yingqian Wang 0002, Jun-Gang Yang, Longguang Wang, Xinyi Ying, Tianhao Wu 0014, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 6 |
| 2020 | Spatial-Angular Interaction for Light Field Image Super-Resolution
Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
ECCV (23) | 4 |
| 2020 | DeOccNet: Learning to See Through Foreground Occlusions in Light FieldsabstractBackground objects occluded in some views of a light field (LF) camera can be seen by other views. Consequently, occluded surfaces are possible to be reconstructed from LF images. In this paper, we handle the LF de-occlusion (LF-DeOcc) problem using a deep encoder-decoder network (namely, DeOccNet). In our method, sub-aperture images (SAIs) are first given to the encoder to incorporate both spatial and angular information. The encoded representations are then used by the decoder to render an occlusion-free center-view SAI. To the best of our knowledge, DeOccNet is the first deep learning-based LF-DeOcc method. To handle the insufficiency oftraining data, we propose an LF synthesis approach to embed selected occlusion masks into existing LF images. Besides, several synthetic and real-world LFs are developed for performance evaluation. Experimental results show that, after training on the generated data, our DeOccNet can effectively remove foreground occlusions and achieves superior performance as compared to other state-of-the-art methods. Source codes are available at: https://github.com/YingqianWang/DeOccNet. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Longguang Wang, Wei An 0003, Yulan Guo |
WACV | 5 |
| 2020 | High-precision refocusing method with one interpolation for camera array imagesabstractCamera array image refocusing can change the in‐focus region so that objects lying on a specified plane are in focus, whereas objects lying off this plane are blurred. Existing refocusing methods for camera array or light field images usually contain two interpolations. Since interpolation brings distortion, especially on the sharp edge in the images, existing methods are not sufficiently precise. In order to improve the quality of the refocusing result, the authors propose a high‐precision method to refocus camera array images. They first back‐project the pixel coordinates to a corresponding location on the focal plane in the world coordinate. Then they reproject the world coordinates to the pixel coordinates by using the parameters of each camera in the array. After that, they align the images with the focal plane by employing interpolation according to the acquired pixel coordinates. Finally, they get the synthetic image refocused on the focal plane by averaging the resulted images. In the proposed method, only one interpolation is used. So that it alleviates the quality degradation of the refocused image compared to the existing methods. Experiments on real‐world scenes (captured by their self‐developed light field devices) demonstrate that their method can yield better results than the existing methods. Jun-Gang Yang, Yingqian Wang 0002, Chengjin An, Wei An 0003 |
IET Image Process. | 5 |
| 2020 | A Stereo Attention Module for Stereo Image Super-ResolutionabstractIn stereo image super-resolution (SR), exploiting both intra-view and cross-view information is significant but challenging. As existing single image SR (SISR) methods are powerful in intra-view information exploitation, in this letter, we propose a generic stereo attention module (SAM) to extend arbitrary SISR networks for stereo image SR. Specifically, we apply two identical pretrained SISR networks to stereo images. The extracted stereo features at different stages are fed to SAMs to interact cross-view information. Finally, the intra-view and cross-view information is incorporated by SISR networks for stereo image SR. Experiments on the KITTI2012, KITTI2015 and Middlebury datasets have demonstrated the effectiveness of our scheme. Using SAM, we can exploit cross-view information while maintaining the superiority of intra-view information exploitation, resulting in notable performance gain to SISR networks. Moreover, SRResNet equipped with our SAM outperforms the state-of-the-art stereo SR methods. Source code is available at https://github.com/XinyiYing/SAM. Xinyi Ying, Yingqian Wang 0002, Longguang Wang, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 5 |
| 2020 | Deformable 3D Convolution for Video Super-ResolutionabstractThe spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal motion compensation are usually performed sequentially. In this paper, we propose a deformable 3D convolution network (D3Dnet) to incorporate spatio-temporal information from both spatial and temporal dimensions for video SR. Specifically, we introduce deformable 3D convolution (D3D) to integrate deformable convolution with 3D convolution, obtaining both superior spatio-temporal modeling capability and motion-aware modeling flexibility. Extensive experiments have demonstrated the effectiveness of D3D in exploiting spatio-temporal information. Comparative results show that our network achieves state-of-the-art SR performance. Code is available at: https://github.com/XinyiYing/D3Dnet. Xinyi Ying, Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 5 |
| 2020 | Deep Video Super-Resolution Using HR Optical Flow EstimationabstractVideo super-resolution (SR) aims at generating a sequence of high-resolution (HR) frames with plausible and temporally consistent details from their low-resolution (LR) counterparts. The key challenge for video SR lies in the effective exploitation of temporal dependency between consecutive frames. Existing deep learning based methods commonly estimate optical flows between LR frames to provide temporal dependency. However, the resolution conflict between LR optical flows and HR outputs hinders the recovery of fine details. In this paper, we propose an end-to-end video SR network to super-resolve both optical flows and images. Optical flow SR from LR frames provides accurate temporal dependency and ultimately improves video SR performance. Specifically, we first propose an optical flow reconstruction network (OFRnet) to infer HR optical flows in a coarse-to-fine manner. Then, motion compensation is performed using HR optical flows to encode temporal dependency. Finally, compensated LR inputs are fed to a super-resolution network (SRnet) to generate SR results. Extensive experiments have been conducted to demonstrate the effectiveness of HR optical flows for SR performance improvement. Comparative results on the Vid4 and DAVIS-10 datasets show that our network achieves the state-of-the-art performance. Longguang Wang, Yulan Guo, Li Liu 0002, Zaiping Lin, Xinpu Deng, Wei An 0003 |
IEEE Trans. Image Process. | 6 |
| 2019 | Learning Parallax Attention for Stereo Image Super-ResolutionabstractStereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets. Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 6 |
| 2019 | Selective Light Field Refocusing for Camera Arrays Using Bokeh Rendering and SuperresolutionabstractCamera arrays provide spatial and angular information within a single snapshot. With refocusing methods, focal planes can be altered after exposure. In this letter, we propose a light field refocusing method to improve the imaging quality of camera arrays. In our method, the disparity is first estimated. Then, the unfocused region (bokeh) is rendered by using a depth-based anisotropic filter. Finally, the refocused image is produced by a reconstruction-based superresolution approach where the bokeh image is used as a regularization term. Our method can selectively refocus images with focused region being superresolved and bokeh being esthetically rendered. Our method also enables postadjustment of depth of field. We conduct experiments on both public and self-developed datasets. Our method achieves superior visual performance with acceptable computational cost as compared to the other state-of-the-art methods. Yingqian Wang 0002, Jun-Gang Yang, Yulan Guo, Wei An 0003 |
IEEE Signal Process. Lett. | 5 |
| 2019 | A Constrained Sparse Representation Model for Hyperspectral Anomaly DetectionabstractIn this paper, we propose a novel sparsity-based algorithm for anomaly detection in hyperspectral imagery. The algorithm is based on the concept that a background pixel can be approximately represented as a sparse linear combination of its spatial neighbors while an anomaly pixel cannot if the anomalies are removed from its neighborhood. To be physically meaningful, the sum-to-one and nonnegativity constraints are imposed to abundance vector based on the linear mixture model, and the upper bound constraint on sparsity level is removed for better recovery of the test pixel. First, the proposed method utilizes the redundant background information to automatically remove anomalies from the background dictionary. Then, the reconstruction error obtained by the new background dictionary is directly used for anomaly detection. Moreover, a kernel version of the proposed method is also derived to completely exploit the nonlinear feature of hyperspectral data. An important advantage of the proposed methods is their capability to adaptively model the background even when some anomaly pixels are involved. Extensive experiments have been conducted on three real hyperspectral data sets. It is demonstrated that the proposed detectors achieve a promising detection performance with a relatively low computational cost. Qiang Ling 0002, Yulan Guo, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Learning Multi-View Representation With LSTM for 3-D Shape Recognition and RetrievalabstractShape representation for 3-D models is an important topic in computer vision, multimedia analysis, and computer graphics. Recent multiview-based methods demonstrate promising performance for 3-D shape recognition and retrieval. However, most multiview-based methods ignore the correlations of multiple views or suffer from high computional cost. In this paper, we propose a novel multiview-based network architecture for 3-D shape recognition and retrieval. Our network combines convolutional neural networks (CNNs) with long short-term memory (LSTM) to exploit the correlative information from multiple views. Well-pretrained CNNs with residual connections are first used to extract a low-level feature of each view image rendered from a 3-D shape. Then, a LSTM and a sequence voting layer are employed to aggregate these features into a shape descriptor. The highway network and a three-step training strategy are also adopted to boost the optimization of the deep network. Experimental results on two public datasets demonstrate that the proposed method achieves promising performance for 3-D shape recognition and the state-of-the-art performance for the 3-D shape retrieval. Chao Ma 0014, Yulan Guo, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Multim. | 4 |
| 2018 | Learning for Video Super-Resolution Through HR Optical Flow Estimation
Longguang Wang, Yulan Guo, Zaiping Lin, Xinpu Deng, Wei An 0003 |
ACCV (1) | 5 |
| 2018 | Simultaneous Context Feature Learning and Hashing for Large Scale Loop Closure DetectionabstractVisual loop closure is important in pose tracking and relocalization in many robotics and Argument Reality (AR) systems. For large and highly repetitive environments, sparse keypoint-based methods face several challenges, especially the discriminability of descriptors. In this paper, we propose an augmented descriptor by combining ORB feature and the context descriptor to increase its discriminability and matching performance. An end-to-end network is adopted to perform simultaneous feature learning and code hashing for the context. In addition, feature position clustering is used to reduce the number of contexts. Besides, hash mapping is adopted to reduce the dimensionality of ORB features. Finally, the context descriptors and ORB features with dimensionality reduction are stacked. Experimental results on the NewCollege and TUM datasets demonstrate that our algorithm achieves higher precision/recall and faster speed than the original algorithm proposed by Antonio et al. [1]. Zhiheng Fu, Yulan Guo, Wei An 0003 |
ICPR | 3 |
| 2018 | Infrared Small Target Detection Using Multiscale Gray and Variance Difference
Jinyan Gao, Yulan Guo, Zaiping Lin, Wei An 0003 |
PRCV (4) | 4 |
| 2018 | Semi-Online Multiple Object Tracking Using Graphical Tracklet AssociationabstractOnline multiple object tracking (MOT) is highly challenging when multiple objects have similar appearance or under long occlusion. In this letter, we propose a semi-online MOT method using online discriminative appearance learning and tracklet association with a sliding window. We connect similar detections of neighboring frames in a temporal window, and improve the performance of appearance feature by online discriminative appearance learning. Then, tracklet association is performed by minimizing a subgraph decomposition cost. Occlusions and missing detections are recovered after tracklet stitching. Our method has been tested on two public datasets. Experimental results have demonstrated the significant performance improvement of our method. Specifically, the proposed method is improved by 8.31% and 12.38% in terms of Multiple Object Tracking Accuracy and Multiple Object Tracking Precision, respectively, as compared to the baseline. Yulan Guo, Xing Tang 0003, Qingyong Hu, Wei An 0003 |
IEEE Signal Process. Lett. | 5 |
| 2017 | Correlation Filter Tracking: Beyond an Open-loop System
Qingyong Hu, Yulan Guo, Yunjin Chen, Wei An 0003 |
BMVC | 5 |
| 2017 | BV-CNNs: Binary Volumetric Convolutional Networks for 3D Object Recognition
Chao Ma 0014, Wei An 0003, Yinjie Lei, Yulan Guo |
BMVC | 2 |
| 2017 | State estimation with incomplete linear constraintabstractA problem of state estimation with destination constraint is considered in this paper. An anti-radiation missile (ARM) often moves towards the target along a trajectory which is almost linear in the X-Y plane. The linear constraint for trajectory and target position are known as priori and can be used to enhance the performance of a tracking filter. In this paper, a destination constrained Kalman filter (DCKF) is first revised for our problem. Then, two methods are proposed to incorporate the prior knowledge by estimating the slope of the trajectory. In the first method, the slope is estimated directly at each time using the point estimated by a unconstrained Kalman filter and the destination point. In the second method, a least square method is used to estimate the slope from all measurements. Several effective linear equality constrained state estimation methods can be used to exploit the estimated slop and the destination point. A typical ARM tracking scenario is established to test the proposed Kalman filter. A comprehensive comparison to recent work is also presented, including unconstrained nonlinear filtering methods and the Posterior Cramer-Rao Lower Bound (PCRLB). Monte-Carlo simulation results are presented to illustrate the effectiveness of the proposed methods for state estimation with destination constraint. Yuan Huang 0006, Xueying Wang 0001, Yulan Guo, Wei An 0003 |
FUSION | 4 |
| 2017 | FSVO: Semi-direct monocular visual odometry using fixed mapsabstractWe propose a fixed-map semi-direct visual odometry (FSVO) algorithm for Micro Aerial Vehicles (MAVs). The proposed approach does not need computationally expensive feature extraction and matching techniques for motion estimation at each frame. Instead, we extract and match ORiented Brief (ORB) features between keyframes and assist-frames. We replace the incremental map generation step in traditional algorithms with fixed map generation at keyframe and assistframe only in our algorithm, resulting in reduced storage memory and higher flexibility for relocalization. Based on the fixed-map, we design a new keyframe selection criterion and a relocalization step. Our algorithm has no limit on the orientation of the camera and reduces drifting effectively. Experimental results on the EuRoC and KITTI datasets show that our algorithm achieves higher precision and robustness than the SVO algorithm. Zhiheng Fu, Yulan Guo, Zaiping Lin, Wei An 0003 |
ICIP | 4 |
| 2017 | An Attitude Jitter Correction Method for Multispectral Parallax Imagery Based on Compressive SensingabstractAttitude jitter is a common problem for high-resolution earth-observation satellites and can diminish the geo-positioning and mapping performance of observed images. It is especially necessary to address this problem when high-performance attitude measurements are unavailable. Therefore, an attitude jitter correction method for multispectral parallax imagery that utilizes the compressive-sensing technology is proposed in this letter. In the proposed method, the attitude jitter is estimated from the parallax disparities of different band images, and then the image displacement caused by attitude jitter can be corrected. Using the normalized cross correlation method and compressive-sensing technology, the proposed method can deal with the condition of texture-feature deficiency in the partial image. The multispectral images of the Terra and ZY-3 satellites are used as experimental data to evaluate the proposed method. The registration errors of different bands are greatly reduced in both the cross- and along-track directions, and the experiment results indicate that the proposed method is effective for correcting the attitude jitter of both satellites. Jun Chen 0007, Jun-Gang Yang, Wei An 0003, Zhi-Jie Chen |
IEEE Geosci. Remote. Sens. Lett. | 3 |