VLDB 2026 Research / reviewers in the wild / expert
Yingqian Wang 0002
dblp:88/284-2
· DBLP profile ↗
68ranked-venue papers
7as first author
61since 2021 · last 2026
0000-0002-9081-6227ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 27 · 4 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 22 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A mutual information-based framework for generalized image fusion via common-unique decoupling
Liyuan Pan, Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002 |
Knowl. Based Syst. | 8 |
| 2026 | Deep Lookup NetworkabstractConvolutional neural networks are constructed with massive operations with different types and are highly computationally intensive. Among these operations, multiplication operation is higher in computational complexity and usually requires more energy consumption with longer inference time than other operations, which hinders the deployment of convolutional neural networks on mobile devices. In many resource-limited edge devices, complicated operations can be calculated via lookup tables to reduce computational cost. Motivated by this, in this paper, we introduce a generic and efficient lookup operation which can be used as a basic operation for the construction of neural networks. Instead of calculating the multiplication of weights and activation values, simple yet efficient lookup operations are adopted to compute their responses. To enable end-to-end optimization of the lookup operation, we construct the lookup tables in a differentiable manner and propose several training strategies to promote their convergence. By replacing computationally expensive multiplication operations with our lookup operations, we develop lookup networks for the image classification, image super-resolution, and point cloud classification tasks. It is demonstrated that our lookup networks can benefit from the lookup operations to achieve higher efficiency in terms of energy consumption and inference speed while maintaining competitive performance to vanilla convolutional networks. Extensive experiments show that our lookup networks produce state-of-the-art performance on different tasks (both classification and regression tasks) and different data types (both images and point clouds). Yulan Guo, Longguang Wang, Wendong Mao, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Probing Deep Into Temporal Profile Makes the Infrared Small Target Detector Much BetterabstractInfrared small target (IRST) detection is challenging in simultaneously achieving precise, robust, and efficient performance due to extremely dim targets and strong interference. Current learning-based methods attempt to leverage "more" information from both the spatial and the short-term temporal domains, but suffer from unreliable performance under complex conditions while incurring computational redundancy. In this paper, we explore the "more essential" information from a more crucial domain for the detection. Through theoretical analysis, we reveal that the global temporal saliency and correlation information in the temporal profile demonstrate significant superiority in distinguishing target signals from other signals. To investigate whether such superiority is preferentially leveraged by well-trained networks, we built the first prediction attribution tool in this field and verified the importance of the temporal profile information. Inspired by the above conclusions, we remodel the IRST detection task as a one-dimensional signal anomaly detection task, and propose an efficient deep temporal probe network (DeepPro) that only performs calculations in the time dimension for IRST detection. We conducted extensive experiments to fully validate the effectiveness of our method. The experimental results are exciting, as our DeepPro outperforms existing state-of-the-art IRST detection methods on widely-used benchmarks with extremely high efficiency, and achieves a significant improvement on dim targets and in complex scenarios. We provide a new modeling domain, a new insight, a new method, and a new performance, which can promote the development of IRST detection. Ruojing Li, Wei An 0003, Yingqian Wang 0002, Xinyi Ying, Yimian Dai, Longguang Wang, Yulan Guo, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Diving Into Epipolar Transformers for Light Field Super-Resolution and Disparity EstimationabstractLight field (LF) cameras capture the light rays of a 3D scene from multiple views simultaneously, and thus provide a more immersive experience of the real world as compared to traditional cameras. Although significant progress has been made in various LF image processing tasks, it remains challenging to effectively model the non-local spatial-angular correlations inherent in LF images, particularly when dealing with complex disparity variations. In this paper, we focus on orthogonal epipolar geometry of LF images and propose a generic Epipolar Transformer mechanism that incorporates geometrically meaningful correlations along the epipolar lines. Our Epipolar Transformer mechanism enjoys the following benefits: learning effective and diverse LF feature representations, delivering satisfactory results without redundant architectural designs, and enabling flexible extension to various LF-related tasks with simple adaptations. For LF spatial and angular super-resolution, our methods not only achieve state-of-the-art performance on benchmark datasets, but also demonstrate superior and robust performance on large disparity variations. For disparity estimation, we explore the use of geometry information encoded in our Epipolar Transformer to directly regress the disparity results, effectively avoiding the limitation of a fixed maximum disparity. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Yulan Guo, Li Liu 0002, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Dynamic High-Frequency Convolution for Infrared Small Target DetectionabstractInfrared small targets are typically tiny and locally salient, which belong to high-frequency components (HFCs) in images. Single-frame infrared small target (SIRST) detection is challenging, since there are many HFCs along with targets, such as bright corners, broken clouds, and other clutters. Current learning-based methods rely on the powerful capabilities of deep networks, but neglect explicit modeling and discriminative representation learning of various HFCs, which is important to distinguish targets from other HFCs. To address the aforementioned issues, we propose a dynamic high-frequency convolution (DHiF) to translate the discriminative modeling process into the generation of a dynamic local filter bank. Especially, DHiF is sensitive to HFCs, owing to the dynamic parameters of its generated filters being symmetrically adjusted within a zero-centered range according to Fourier transformation properties. Combining with standard convolution operations, DHiF can adaptively and dynamically process different HFC regions and capture their distinctive grayscale variation characteristics for discriminative representation learning. DHiF functions as a drop-in replacement for standard convolution and can be used in arbitrary SIRST detection networks without significant decrease in computational efficiency. To validate the effectiveness of our DHiF, we conducted extensive experiments across different SIRST detection networks on real-scene datasets. Compared to other state-of-the-art convolution operations, DHiF exhibits superior detection performance with promising improvement. Codes are available at https://github.com/TinaLRJ/DHiF. Ruojing Li, Wei An 0003, Xinyi Ying, Yingqian Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Revisiting Subspace Disentangling for Light Field Spatial Super-Resolution
Yingqian Wang 0002, Xueying Wang 0001, Zhengyu Liang, Longguang Wang, Lvli Tian, Jun-Gang Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | DuaDiff: Dual-Conditional Diffusion Model for Guided Thermal Image Super-ResolutionabstractThermal imaging offers valuable properties, but suffers from inherently low spatial resolution, which can be enhanced using a high-resolution (HR) visible image as guidance. However, the substantial modality differences between thermal and visible images, coupled with significant resolution gaps, pose challenges to existing guided super-resolution (SR) approaches. In this article, we present dual-conditional diffusion (DuaDiff), an innovative diffusion model featuring a dual-conditioning mechanism to enhance guided thermal image SR. Unlike typical conditional diffusion models, DuaDiff integrates a learnable Laplacian pyramid to extract high-frequency details from the visible image, serving as one of the conditioning inputs. By capturing multiscale high-frequency components, DuaDiff effectively focuses on intricate textures and edges in the HR visible images, significantly enhancing thermal image fidelity. Furthermore, we project both thermal and visible images into a semantic latent space, constructing another conditioning input. Leveraging these complementary conditions, DuaDiff employs a multimodal latent feature cross-attention module to facilitate effective interaction between noise, thermal, and visible latent representations. Extensive experiments on the FLIR-ADAS and CATS datasets for $4\times $ and $8\times $ guided SR demonstrate that combining learnable Laplacian conditioning with semantic latent conditioning enables DuaDiff to surpass state-of-the-art methods in both visual quality and metric evaluation, particularly in scenarios with a large resolution gap. Besides, the applications to downstream tasks further confirm the capability of DuaDiff to recover high-fidelity semantic information. The code will be released. Linrui Shi, Gaochang Wu, Yingqian Wang 0002, Yebin Liu, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Exploring cross-branch information for semi-supervised remote sensing object detection
Shitian He, Huanxin Zou, Yingqian Wang 0002, Hao Chen 0046, Ning Jing |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Multimodal image generation and fusion through content-style hybrid disentanglementabstract• Research highlight 1: We propose a novel cross-task hybrid training methodology for multimodal images, offering a simple yet unified solution that simultaneously addresses both image generation and fusion tasks. • Research highlight 2: Building upon mutual-supervised multimodal image pairs, we innovatively integrate single-modality self-supervision to develop a hybrid-supervised decoupling framework with a dedicated loss function, achieving robust separation of content-style representations. • Research highlight 3: Extensive experiments spanning on four modalities and seven popular datasets demonstrate our method’s consistent superiority and impressive cross-task capability. Ablation studies further reveal that our framework learns generalized representations transferable across different image processing tasks. Multimodal image fusion and cross-modal translation are fundamental yet challenging tasks in computer vision, with their performance directly impacting downstream applications. Existing approaches typically treat these tasks independently, developing specialized models that fail to exploit the intrinsic relationships between different modalities. This limitation not only restricts model generalizability but also hinders further performance improvements. In this paper, we propose a joint optimization framework for image generation and fusion. Specifically, we generalize multimodal image tasks as the fusion and transformation of cross-modal features, and design a hybrid task training strategy. At the data level, we introduce a self-supervised and mutual-supervised hybrid mechanism for content-style feature decoupling, which achieves superior feature separation through stepwise training on intra-modal and cross-modal data. At the model level, we construct a triple-branch decoupling head along with fusion and transformation modules to ensure synchronous and efficient execution of dual tasks. Our method not only breaks through the single task limitation of the model, but also innovatively introduces mixed supervision into multimodal processing. We conduct comprehensive experiments covering four modalities fusion tasks on seven popular datasets. Extensive experimental results demonstrate that our method achieves superior performance on two tasks as compared of the respective state-of-the-art methods, and show impressive cross-task generalization capability. Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002, Liyuan Pan |
Knowl. Based Syst. | 7 |
| 2025 | Unsupervised Degradation Representation Learning for Unpaired Restoration of Images and Point CloudsabstractRestoration tasks in low-level vision aim to restore high-quality (HQ) data from their low-quality (LQ) observations. To circumvents the difficulty of acquiring paired data in real scenarios, unpaired approaches that aim to restore HQ data solely on unpaired data are drawing increasing interest. Since restoration tasks are tightly coupled with the degradation model, unknown and highly diverse degradations in real scenarios make learning from unpaired data quite challenging. In this paper, we propose a degradation representation learning scheme to address this challenge. By learning to distinguish various degradations in the representation space, our degradation representations can extract implicit degradation information in an unsupervised manner. Moreover, to handle diverse degradations, we develop degradation-aware (DA) convolutions with flexible adaption to various degradations to fully exploit the degrdation information in the learned representations. Based on our degradation representations and DA convolutions, we introduce a generic framework for unpaired restoration tasks. Based on our framework, we propose UnIRnet and UnPRnet for unpaired image and point cloud restoration tasks, respectively. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on unpaired image and point cloud restoration tasks show that our UnIRnet and UnPRnet achieve state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Graph Laplacian regularization for fast infrared small target detection
Ting Liu 0017, Yongxian Liu, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 5 |
| 2025 | TO-LF: A Texture and Occlusion-Oriented Benchmark Dataset for Light Field Disparity EstimationabstractAccurate disparity estimation in light field (LF) imaging remains challenging due to the narrow baseline between adjacent sub-aperture images (SAIs) and the occlusion effect. Existing learning-based methods suffer from degraded performance in complex scenarios owing to the scarcity of high-quality and diverse training data. To address this limitation, we propose a Texture and Occlusion-oriented Light Field dataset (TO-LF) containing 78 carefully curated images. Unlike the widely used HCI 4D LF benchmark, TO-LF not only provides more training samples but also introduces a more challenging test set with complex occlusions and significant textureless regions. Furthermore, we present a viewpoint-selective sub-pixel cost volume construction method (VS-Sub), which extends disparity labels to the subpixel level for denser cost volumes, and employs dynamic dilated convolutions to differentiate between occluded and non-occluded viewpoints. Comprehensive experiments demonstrate that our framework achieves state-of-the-art (SOTA) performance in disparity estimation. Shubo Zhou, Yunlong Wang 0003, Yingqian Wang 0002, Fei Liu 0031, Xueqin Jiang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Fixed Relative Pose Prior for Camera Array Self-CalibrationabstractCamera arrays have unique advantages in various computer vision tasks, such as 3D scene reconstruction and depth estimation. For these tasks, precise calibration of sub-cameras is crucial. Since the baselines of sub-cameras are usually small, it is challenging to calibrate the camera array through a single recording of the scene. Consequently, the majority of existing calibration methods address this issue by recording a scene at different spatial locations. However, this approach neglects the prior that the relative pose of the sub-cameras remains unchanged across different locations, which leads to an increase in cumulative reprojection errors. In this letter, we propose to incorporate this fixed relative pose prior to precisely calibrate the camera array. Specifically, we first capture dual-array frames by recording a scene at two spatial locations. Then, we incorporate the fixed relative pose prior to the camera array calibration process by integrating the linear constraint into the organization of sub-aperture images (SAIs). Our method maintains the minimum necessary degrees of freedom for the calibration model, and reduces cumulative reprojection error. Moreover, we develop a real-world light field dataset for comprehensive performance evaluation. Experimental results demonstrate that our method can achieve higher calibration accuracy as compared to existing methods. Our code and dataset are available athttps://github.com/Zhangyaning-NUDT/Fixed-relative-pose-prior-for-camera-array-self-calibration. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Learning Rotation-Invariant Neighbor Consensus for Mismatch Removal
Fanzhi Cao, Wei An 0003, Tianxin Shi, Yingqian Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | CWIMamba: Cross-Scale Windowed Integration State Space Model for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) intends to detect potential anomalous targets hidden in the background of hyperspectral images (HSIs) and has garnered substantial attention in various remote sensing photography and surveying applications. Recent research advances in the HAD domain have highlighted the significance of deep convolutional networks (DCNs) and vision transformers (ViTs)-based formulas. However, DCNs are long-range dependency-limited networks, whereas ViTs bear the computational burden of quadratic complexity. Owing to their prominent nonlocal representations and linear complexity, Mamba-based approaches have drawn growing attention. Our study pioneers the integration of Mamba into HAD tasks, presenting CWIMamba, which introduces a novel cross-scale windowed integration state space model for considering the spatial distribution characteristics of the anomaly targets. Specifically, we devise a cross-scale windowed state space model (CSWSSM) to scan the spatial-spectral features based on the window-based bottleneck SSM with different scales. For better multiscale feature integration, a multiscale spatial-spectral feature adaptive integration (MS3FAI) method is explored to generate an intensified representation of multiscale feature interaction and fusion based on the elaborate adaptive spatial-spectral weighting scheme. Moreover, we also devised a Haar discrete wavelet transform convolution module (HDWTCM) to fully replenish the local informative representation and enhance the discriminative frequency characteristics between anomalies and background, introducing more inductive local features for accurate background reconstruction and anomaly suppression. Extensive experiments on five multifarious HAD datasets and seven indicators substantiate the state-of-the-art detection performance, demonstrating the effectiveness of CWIMamba. Wei An 0003, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Shilin Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Sparsity-Aware Global Channel Pruning for Infrared Small-Target Detection NetworksabstractFor infrared small-target detection, convolutional neural network (CNN)-based methods have demonstrated promising performance. However, due to the small size of the targets, existing infrared small-target detection methods necessitate intricate structures with intensive computation to maintain distinctive features of targets in deep layers, which poses great challenges for deployment on edge devices with constrained resources. The current pruning methods predominantly focus on the optimization of classification networks, with an emphasis on semantic information. Nevertheless, the spatial details crucial for infrared small-target detection are excessively pruned in the shallow layers, resulting in a significant degradation of detection performance. In this article, based on the sparse distribution of small targets in infrared images, we propose a sparsity-aware global channel pruning (SAGCP) framework to optimize infrared small-target detection networks. Specifically, sparse modeling is used to code the target region for the first time, and sparse priors can be induced into the feature map for the identification of redundant channels. Without the need for extra structures or intricate criteria to identify redundant channels, our pruning method can leverage the inherent properties of infrared small targets to extract more robust features and obtain more compact models. When SAGCP is applied to the existing infrared small-target detection methods, the pruned network is superior in model efficiency and detection performance. For example, when applying our method to DNA-Net, the pruned model can achieve a 72.34% reduction in parameters, and a 57.49% decrease in floating point operations (FLOPs), but a 2.02% increase in intersection over union (IoU). Shuanglin Wu, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Motion and Appearance Decoupling Representation for Event CamerasabstractEvent cameras, with high temporal resolution and high dynamic range, have shown great potential under extreme scenarios such as high-speed movement and low illumination. However, previous event representation methods typically aggregate event data into a single dense tensor, often overlooking the dynamic changes of events within a given time unit. This limitation can introduce historical artifacts and semantic inconsistencies, ultimately degrading model performance. Inspired by human visual prior, we propose a motion and appearance decoupling (MAD) event representation to disentangle the mixed spatial-temporal event tensor into two independent branches. This bio-inspired design helps the network extract discriminative temporal (i.e., motion) and spatial (i.e., appearance) information, thus reducing the network's learning burden toward complex high-level interpretation tasks. In our method, the event motion guided attention module (EMGA) is designed to achieve temporal and spatial feature interaction and fusion sequentially. Based on EMGA, three specially designed decoder heads are proposed for several representative event-based tasks (i.e., object detection, semantic segmentation, and human pose estimation). Experimental results demonstrate that our method achieves state-of-the-art performance on the above three tasks, which reveals that our method is an easy-to-implement replacement for currently event-based methods. Our code is available at: https://github.com/ChenYichen9527/MAD-representation. Boyang Li 0007, Yingqian Wang 0002, Xinyi Ying, Longguang Wang, Chushu Zhang, Yulan Guo, Wei An 0003 |
IEEE Trans. Image Process. | 3 |
| 2025 | Direction-Coded Temporal U-Shape Module for Multiframe Infrared Small Target DetectionabstractInfrared small target (IRST) detection aims at separating targets from cluttered background. Although many deep learning-based single-frame IRST (SIRST) detection methods have achieved promising detection performance, they cannot deal with extremely dim targets while suppressing the clutters since the targets are spatially indistinctive. Multiframe IRST (MIRST) detection can well handle this problem by fusing the temporal information of moving targets. However, the extraction of motion information is challenging since general convolution is insensitive to motion direction. In this article, we propose a simple yet effective direction-coded temporal U-shape module (DTUM) for MIRST detection. Specifically, we build a motion-to-data mapping to distinguish the motion of targets and clutters by indexing different directions. Based on the motion-to-data mapping, we further design a direction-coded convolution block (DCCB) to encode the motion direction into features and extract the motion information of targets. Our DTUM can be equipped with most single-frame networks to achieve MIRST detection. Moreover, in view of the lack of MIRST datasets, including dim targets, we build a multiframe infrared small and dim target dataset (namely, NUDT-MIRSDT) and propose several evaluation metrics. The experimental results on the NUDT-MIRSDT dataset demonstrate the effectiveness of our method. Our method achieves the state-of-the-art performance in detecting infrared small and dim targets and suppressing false alarms. Our codes will be available at https://github.com/TinaLRJ/Multi-frame-infrared-small-target-detection-DTUM. Ruojing Li, Wei An 0003, Boyang Li 0007, Yingqian Wang 0002, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Real-World Light Field Image Super-Resolution Via Degradation ModulationabstractRecent years have witnessed the great advances of deep neural networks (DNNs) in light field (LF) image super-resolution (SR). However, existing DNN-based LF image SR methods are developed on a single fixed degradation (e.g., bicubic downsampling), and thus cannot be applied to super-resolve real LF images with diverse degradation. In this article, we propose a simple yet effective method for real-world LF image SR. In our method, a practical LF degradation model is developed to formulate the degradation process of real LF images. Then, a convolutional neural network is designed to incorporate the degradation prior into the SR process. By training on LF images using our formulated degradation, our network can learn to modulate different degradation while incorporating both spatial and angular information in LF images. Extensive experiments on both synthetically degraded and real-world LF images demonstrate the effectiveness of our method. Compared with existing state-of-the-art single and LF image SR methods, our method achieves superior SR performance under a wide range of degradation, and generalizes better to real LF images. Codes and models are available at https://yingqianwang.github.io/LF-DMnet/. Yingqian Wang 0002, Zhengyu Liang, Longguang Wang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Learning Coupled Dictionaries from Unpaired Data for Image Super-ResolutionabstractThe difficulty of acquiring high-resolution (HR) and low-resolution (LR) image pairs in real scenarios limits the performance of existing learning-based image super-resolution (SR) methods in the real world. To conduct training on real-world unpaired data, current methods focus on synthesizing pseudo LR images to associate unpaired images. However, the realness and diversity of pseudo LR images are vulnerable due to the large image space. In this paper, we cir-cumvent the difficulty of image generation and propose an alternative to build the connection between unpaired images in a compact proxy space. Specifically, we first construct coupled HR and LR dictionaries, and then encode HR and LR images into a common latent code space using these dictionaries. In addition, we develop an autoencoder-based framework to couple these dictionaries during optimization by reconstructing input HR and LR images. The coupled dictionaries enable our method to employ a shal-low network architecture with only 18 layers to achieve efficient image SR. Extensive experiments show that our method (DictSR) can effectively model the LR-to-HR mapping in coupled dictionaries and produces state-of-the-art performance on benchmark datasets. Longguang Wang, Juncheng Li 0003, Yingqian Wang 0002, Qingyong Hu, Yulan Guo |
CVPR | 3 |
| 2024 | Learning Remote Sensing Object Detection With Single Point SupervisionabstractPointly Supervised Object Detection (PSOD) has attracted considerable interests due to its lower labeling cost as compared to box-level supervised object detection. However, the complex scenes, densely packed and dynamic-scale objects in Remote Sensing (RS) images hinder the development of PSOD methods in RS field. In this paper, we make the first attempt to achieve RS object detection with single point supervision, and propose a PSOD method tailored for RS images. Specifically, we design a point label upgrader (PLUG) to generate pseudo box labels from single point labels, and then use the pseudo boxes to supervise the optimization of existing detectors. Moreover, to handle the challenge of the densely packed objects in RS images, we propose a sparse feature guided semantic prediction module which can generate high-quality semantic maps by fully exploiting informative cues from sparse objects. Extensive ablation studies on the DOTA dataset have validated the effectiveness of our method. Our method can achieve significantly better performance as compared to state-of-the-art image-level and point-level supervised detection methods, and reduce the performance gap between PSOD and box-level supervised object detection. Code is available at https://github.com/heshitian/PLUG. Shitian He, Huanxin Zou, Yingqian Wang 0002, Boyang Li 0007, Ning Jing |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Mixed-Precision Network Quantization for Infrared Small Target SegmentationabstractNetwork quantization is leveraged to reduce the model size, memory footprint, and computational cost of deep neural networks. It is achieved by representing float weights and activations with lower bit counterparts, which is essential for model deployment on resource-limited devices. However, due to the extremely small size of infrared small targets in the feature map, low-bit quantization could lead to huge information loss of small targets and thus causes severe segmentation performance degradation. To achieve low-bit quantization while maintaining the segmentation performance, we first study the quantization sensitivity of small target segmentation network and observe the sensitivity heterogeneity of different layers in the network. Specifically, feature maps in shallow layers and encoder subnetwork are more vulnerable to information loss caused by quantization as compared to deep layers and decoder subnetwork. Based on these observations, we are motivated to assign a different bitwidth for each block according to their quantization sensitivity. A simple yet effective symmetrically progressive decreasing mixed-precision quantization (SPMix-Q) method is proposed to achieve high-performance segmentation under low-bit quantization (i.e., 2.42 bits for weights and 3.82 bits for activations). The experimental results show that our SPMix-Q achieves comparable accuracy with only 1/13 model size, 1/4.6 memory footprint, and 1/29 computational cost to the full-precision counterparts. Compared with the homogeneous low-bit quantization methods, our method achieves much better performance in terms of intersection of union (IoU) on the benchmark datasets. Our mobile-system-on-a-chip (SOC) (e.g., Kyrin 980, Snapdragon 660, and Dimensity 800U) deployable android application package (APK) is available at:https://github.com/YeRen123455/SIRST-Quantization-Deployment. Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Tianhao Wu 0014, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Label Assignment Matters: A Gaussian Assignment Strategy for Tiny Object DetectionabstractRecently, impressive improvements have been achieved in general object detection. However, tiny object detection remains a very challenging problem since tiny objects only occupy a few pixels. Consequently, the label assignment strategies used in general object detectors are not suitable for tiny object detection, because these algorithms tend to assign few or even no positive samples for tiny objects. In this article, we propose a simple yet effective Gaussian assignment (GA) strategy to solve this problem. Specifically, we first model the bounding boxes as 2-D Gaussian distributions and then encode training samples with a threshold. This strategy can assign more high-quality positive samples for tiny objects and adjust the weight of positive samples to balance the contribution from different-size objects. Extensive experiments on four tiny object detection datasets show that the proposed strategy significantly and consistently improves the performance of single-stage tiny object detectors. In particular, with our strategy, we bridge the performance gap between single-stage and state-of-the-art multistage detectors on the AI-TOD dataset (24.2% versus 24.8% in mAP) while maintaining the inference speed. The code is available athttps://github.com/zf020114/GaussianAssignment. Feng Zhang 0046, Shilin Zhou 0001, Yingqian Wang 0002, Xueying Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | DDAug: Differentiable Data Augmentation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation(WSSS) with image-level labels has witnessed promising advances with the help ofclass activation maps(CAM). However, CAM is always confined to small discriminative seed regions due to its simple classification loss guided training manner. To handle this problem, recent works introduced specifically designed regularizations and modules to expand the CAM seed regions, serving as the final segmentation masks. In this paper, we surprisingly find that the classification loss could suppress the gains from these regularization and modules in the late training phase, thereby limiting the further growth of CAM, which we call as theexplicit supervision disturb(ESD) issue. Interestingly, we find that specificdata augmentation(DA) operations (e.g., CutMix) can relieve such ESD issue, and the benefits introduced by different DA operations vary a lot. To maximize the benefits, we proposedifferentiable data augmentation(DDAug) to automatically search for the proper DA policy. Specifically, we design amulti-level search spaceto sequentially sample DA operations with different properties. Extensive experiments demonstrate that the proposed DDAug can alleviate the ESD issue and introduce consistent improvements to various popular WSSS methods, achieving the state-of-the-art performance on the MS COCO 2014 and PASCAL VOC 2012 datasets. Boyang Li 0007, Fei Zhang 0016, Longguang Wang, Yingqian Wang 0002, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Multim. | 4 |
| 2024 | Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T VideosabstractTracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on singlemodality scenes, and overlook the complementary characteristics of tiny objects captured by multiple remote sensors. To enhance tracking performance by integrating complementary information from multiple sources, we propose a novel framework called HGT-Track (Heterogeneous Graph Transformer based Multi-Tiny-Object Tracking). Specifically, we first employ a Transformer-based encoder to embed images from different modalities. Subsequently, we utilize Heterogeneous Graph Transformer to aggregate spatial and temporal information from multiple modalities to generate detection and tracking features. Additionally, we introduce a target re-detection module (ReDet) to ensure tracklet continuity by maintaining consistency across different modalities. Furthermore, this paper introduces the first benchmark VT-Tiny-MOT (Visible-Thermal Tiny MultiObject Tracking) for RGB-T fused multiple tiny object tracking. Extensive experiments are conducted on VT-Tiny-MOT, and the results have demonstrated the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of MOTA (Multiple-Object Tracking Accuracy) and ID-F1 score. The code and dataset will be made available at https://github.com/xuqingyu26/HGTMT Longguang Wang, Weidong Sheng, Yingqian Wang 0002, Chao Ma 0014, Wei An 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point SupervisionabstractTraining a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we make the first attempt to achieve infrared small target detection with point-level supervision. Interestingly, during the training phase supervised by point labels, we discover that CNNs first learn to segment a cluster of pixels near the targets, and then gradually converge to predict groundtruth point labels. Motivated by this “mapping degeneration” phenomenon, we propose a label evolution framework named label evolution with single point supervision (LESPS) to progressively expand the point label by leveraging the intermediate predictions of CNNs. In this way, the network predictions can finally approximate the updated pseudo labels, and a pixel-level target mask can be obtained to train CNNs in an end-to-end manner. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Experimental results show that CNNs equipped with LESPS can well recover the target masks from corresponding point labels, and can achieve over 70% and 95% of their fully supervised performance in terms of pixel-level intersection over union (IoU) and object-level probability of detection (Pd), respectively. Code is available at https://github.com/XinyiYing/LESPS. Xinyi Ying, Li Liu 0002, Yingqian Wang 0002, Ruojing Li, Zaiping Lin, Weidong Sheng, Shilin Zhou 0001 |
CVPR | 3 |
| 2023 | Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-level annotations. To reduce the annotation burden, we propose the first method to achieve SIRST detection with single-point supervision. The core idea of this work is to recover the per-pixel mask of each target from the given single point label by using clustering approaches, which looks simple but is indeed challenging since targets are always insalient and accompanied with background clutters. To handle this issue, we introduce randomness to the clustering process by adding noise to the input images, and then obtain much more reliable pseudo masks by averaging the clustered results. Thanks to this "Monte Carlo" clustering approach, our method can accurately recover pseudo masks and thus turn arbitrary fully supervised SIRST detection networks into weakly supervised ones with only single point annotation. Experiments on four datasets demonstrate that our method can be applied to existing SIRST detection networks to achieve comparable performance with their fully-supervised counterparts, which reveals that single-point supervision is strong enough for SIRST detection. Our code will be available at: https://github.com/YeRen123455/SIRST-Single-Point-Supervision. Boyang Li 0007, Yingqian Wang 0002, Longguang Wang, Fei Zhang 0016, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo |
ICCV | 2 |
| 2023 | Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-ResolutionabstractExploiting spatial-angular correlation is crucial to light field (LF) image super-resolution (SR), but is highly challenging due to its non-local property caused by the disparities among LF images. Although many deep neural networks (DNNs) have been developed for LF image SR and achieved continuously improved performance, existing methods cannot well leverage the long-range spatial-angular correlation and thus suffer a significant performance drop when handling scenes with large disparity variations. In this paper, we propose a simple yet effective method to learn the non-local spatial-angular correlation for LF image SR. In our method, we adopt the epipolar plane image (EPI) representation to project the 4D spatial-angular correlation onto multiple 2D EPI planes, and then develop a Transformer network with repetitive self-attention operations to learn the spatial-angular correlation by modeling the dependencies between each pair of EPI pixels. Our method can fully incorporate the information from all angular views while achieving a global receptive field along the epipolar line. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Comparative results on five public datasets show that our method not only achieves state-of-the-art SR performance but also performs robust to disparity variations. Code is publicly available at https://github.com/ZhengyuLiang24/EPIT. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001, Yulan Guo |
ICCV | 2 |
| 2023 | Disentangling Light Fields for Super-Resolution and Disparity EstimationabstractLight field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been proposed for various LF image processing tasks. However, it is challenging for CNNs to effectively process LF images since the spatial and angular information are highly inter-twined with varying disparities. In this paper, we propose a generic mechanism to disentangle these coupled information for LF image processing. Specifically, we first design a class of domain-specific convolutions to disentangle LFs from different dimensions, and then leverage these disentangled features by designing task-specific modules. Our disentangling mechanism can well incorporate the LF structure prior and effectively handle 4D LF data. Based on the proposed mechanism, we develop three networks (i.e., DistgSSR, DistgASR and DistgDisp) for spatial super-resolution, angular super-resolution and disparity estimation. Experimental results show that our networks achieve state-of-the-art performance on all these three tasks, which demonstrates the effectiveness, efficiency, and generality of our disentangling mechanism. Project page: https://yingqianwang.github.io/DistgLF/. Yingqian Wang 0002, Longguang Wang, Gaochang Wu, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Exploring Fine-Grained Sparsity in Convolutional Neural Networks for Efficient InferenceabstractNeural networks contain considerable redundant computation, which drags down the inference efficiency and hinders the deployment on resource-limited devices. In this paper, we study the sparsity in convolutional neural networks and propose a generic sparse mask mechanism to improve the inference efficiency of networks. Specifically, sparse masks are learned in both data and channel dimensions to dynamically localize and skip redundant computation at a fine-grained level. Based on our sparse mask mechanism, we develop SMPointSeg, SMSR, and SMStereo for point cloud semantic segmentation, single image super-resolution, and stereo matching tasks, respectively. It is demonstrated that our sparse masks are well compatible to different model components and network architectures to accurately localize redundant computation, with computational cost being significantly reduced for practical speedup. Extensive experiments show that our SMPointSeg, SMSR, and SMStereo achieve state-of-the-art performance on benchmark datasets in terms of both accuracy and efficiency. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Combining Deep Denoiser and Low-rank Priors for Infrared Small Target DetectionabstractMany existing low-rank methods have achieved good detection performance in uniform scenes, but they suffer from a high false alarm rate in complex noisy scenes. Therefore, it is important to improve the detection performance of low-rank models in noisy scenes. In this paper, we first formulate an implicit regularizer by plugging a denoising neural network (termed as deep denoiser), which can learn deep image priors from a large number of natural images. Then, we use the weighted sum of weighted tensor nuclear norm for more accurate background estimation. Finally, alternating direction multiplier method is used to solve the model under the plug-and-play framework. By integrating low-rank prior with deep denoiser prior, our model achieves higher accuracy. Experiments on different scenes demonstrate that our method achieves an improved performance in terms of visual effects and quantitative metrics. Specially, the overall accuracy of AUC value (AUCOA) achieved by the proposed method on Sequences 1-6 are 1.24%, 1.16%, 0.63%, 1.9%, 0.82%, 2.06% higher than those achieved by the second top performing methods, respectively. Ting Liu 0017, Jun-Gang Yang, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 4 |
| 2023 | Not All Patches Are Equal: Hierarchical Dataset Condensation for Single Image Super-ResolutionabstractAlthough the performance of single image super-resolution (SR) has been significantly improved with deep neural networks, existing methods commonly require millions of iterations for training, which not only limits their training efficiency, but also causes considerable energy consumption. In this paper, we comprehensively study the redundancy of existing training datasets and reveal that not all patches are equal for SR network training. We observe that a large percentage of patches with low textures or similar textures lead to high computation costs but make low contributions to SR performance. Then, we propose a dataset condensation method to remove these redundant patches hierarchically. Extensive experiments demonstrate that our dataset condensation method can effectively reduce the redundancy of SR datasets with a 90% condensation rate on DIV2K. With our condensed dataset, baseline networks can achieve significant improvement in terms of training efficiency while maintaining competitive accuracy. Codes are available athttps://github.com/QingtangDing/DCSR. Qingtang Ding, Zhengyu Liang, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang |
IEEE Signal Process. Lett. | 4 |
| 2023 | You Only Train Once: Learning a General Anomaly Enhancement Network With Random Masks for Hyperspectral Anomaly DetectionabstractIn this paper, we introduce a new approach to address the challenge of generalization in hyperspectral anomaly detection (AD). Our method eliminates the need for adjusting parameters or retraining on new test scenes as required by most existing methods. Employing an image-level training paradigm, we achieve a general anomaly enhancement network for hyperspectral AD that only needs to be trained once. Trained on a set of anomaly-free hyperspectral images with random masks, our network can learn the spatial context characteristics between anomalies and background in an unsupervised way. Additionally, a plug-and-play model selection module is proposed to search for a spatial-spectral transform domain that is more suitable for AD task than the original data. To establish a unified benchmark to comprehensive evaluate our method and existing methods, we develop a large-scale hyperspectral AD dataset (HAD100) that includes 100 real test scenes with diverse anomaly targets. In comparison experiments, we combine our network with a parameter-free detector, and achieve the optimal balance between detection accuracy and inference speed among state-of-the-art AD methods. Experimental results also show that our method still achieves competitive performance when the training and test set are captured by different sensor devices. Our code is available at https://github.com/ZhaoxuLi123/AETNet. Zhaoxu Li, Yingqian Wang 0002, Qiang Ling 0002, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Infrared Small Target Detection via Nonconvex Tensor Tucker Decomposition With Factor PriorabstractInfrared small target detection in complex scenes is an important but challenging research hotspot in infrared early warning fields. Previous studies have proved that low-rank Tucker decomposition (TD) achieves good detection performance in complex scenes. However, a key limitation of existing low-rank TD methods is that the rank needs to be set in advance, and an inaccurate predefined rank can lead to performance degradation. Inspired by the theorem that n-rank is upper bounded by the rank of each Tucker factor matrix, we propose a nonconvex tensor TD model with factor prior for infrared small target detection. In our method, we use a logdet-based function to constrain the latent factors of low-rank TD, which avoids empirical rank selection and sufficiently uses the latent data structure information in the factor matrix. Meanwhile, performing singular value decomposition (SVD) calculations on small factor matrices can reduce computational complexity. Then, group sparsity regularized total variation is used to better exploit the shared sparse pattern of difference images, which helps better remove background clutter and obtain better detection results. Finally, the proposed method is efficiently solved by the well-designed alternating direction method of multipliers (ADMM). Extensive experimental results demonstrate that our method is more effective and robust in complex scenes than other state-of-the-art methods. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Representative Coefficient Total Variation for Efficient Infrared Small Target DetectionabstractLow-rank and sparse decomposition based models are powerful and robust tools for infrared small target detection. However, due to the calculation of singular value decomposition (SVD) and the optimization of complex regularization terms, existing low-rank models often suffer from high computational complexity. To solve this problem, based on the theorem that representative coefficient matrix obtained by orthogonal transformation of data matrix can inherit the spatial structure of data matrix, we propose a representative coefficient total variation (RCTV) method for efficient infrared small target detection. In our method, we use total variational to constraint representative coefficient matrix instead of data matrix to describe local smooth prior, which helps remove noise and reduce computational complexity. Meanwhile, we control the number of columns in the representative coefficient matrix to maintain the low-rank characteristics of background, which avoids SVD calculation and improves detection efficiency. Therefore, the RCTV regularization can simultaneously describe local smooth prior and low-rank prior. Moreover, to better enhance the sparsity of targets and distinguish sparse non-target points, we use the log-sum function to adaptively assign weights to targets. It helps obtain more accurate detection performance. The proposed model is efficiently solved by the alternating direction multiplier method (ADMM). A large number of experiments show that the proposed method is superior to existing low-rank methods in both detection accuracy and efficiency. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | MTU-Net: Multilevel TransUNet for Space-Based Infrared Tiny Ship DetectionabstractSpace-based infrared tiny ship detection aims at separating tiny ships from the images captured by Earth-orbiting satellites. Due to the extremely large image coverage area (e.g., thousands of square kilometers), candidate targets in these images are much smaller, dimer, and more changeable than those targets observed by aerial- and land-based imaging devices. Existing short imaging distance-based infrared datasets and target detection methods cannot be well adopted to the space-based surveillance task. To address these problems, we develop a space-based infrared tiny ship detection dataset (namely, NUDT-SIRST-Sea) with 48 space-based infrared images and$17\,598$pixel-level tiny ship annotations. Each image covers about$10\,000$km2of area with$10 \ 000\,\, \times \ 10 \ 000$pixels. Considering the extreme characteristics (e.g., small, dim, and changeable) of those tiny ships in such challenging scenes, we propose a multilevel TransUNet (MTU-Net) in this article. Specifically, we design a vision Transformer (ViT) convolutional neural network (CNN) hybrid encoder to extract multilevel features. Local feature maps are first extracted by several convolution layers and then fed into the multilevel feature extraction module [multilevel ViT module (MVTM)] to capture long-distance dependency. We further propose a copy–rotate–resize–paste (CRRP) data augmentation approach to accelerate the training phase, which effectively alleviates the issue of sample imbalance between targets and background. Besides, we design a FocalIoU loss to achieve both target localization and shape description. Experimental results on the NUDT-SIRST-Sea dataset show that our MTU-Net outperforms traditional and existing deep learning-based single-frame infrared small target (SIRST) methods in terms of probability of detection, false alarm rate, and intersection over union. Our code is available athttps://github.com/TianhaoWu16/Multi-level-TransUNet-for-Space-based-Infrared-Tiny-ship-Detection Tianhao Wu 0014, Boyang Li 0007, Yihang Luo, Yingqian Wang 0002, Ting Liu 0017, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | RepISD-Net: Learning Efficient Infrared Small-Target Detection Network via Structural Re-ParameterizationabstractInfrared small target detection is a challenging task for deep learning-based methods because targets tend to disappear in the deep layers. To handle this problem, existing deep neural networks usually apply various dense and skip connections for feature maintenance. Although these well-designed networks have achieved good detection performance, the complex network structures reduce their efficiency. In this paper, we propose a simple yet efficient network (RepISD-Net) for infrared small target detection. The core of our RepISD-Net is to use different network architectures but equivalent model parameters for training and inference, respectively. Specifically, in the training phase, we design a parallel multi-branch edge compensation block (ECB) to enhance the local salient features and capture finer contour characteristic of infrared small targets. In the inference phase, the multi-branch topology structures are merged into a single branch with only cascaded 3×3 convolutions for fast inference. We conduct extensive experiments on several public datasets to validate the effectiveness of our method. Experimental results demonstrate that our RepISD-Net can achieve comparable or even better detection performance with significant acceleration in inference speed as compared to state-of-the-art infrared small target detection methods. Code is submitted for review and will be released upon acceptance. Shuanglin Wu, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Incorporating Deep Background Prior Into Model-Based Method for Unsupervised Moving Vehicle Detection in Satellite VideosabstractBackground reconstruction is a key step of moving object detection in satellite videos. Most existing model-based methods exploit low-rank prior to recover background, which have achieved good performance but suffered degradation under complex and dynamic scenes. In this paper, we introduce a deep background prior into model-based methods for moving vehicle detection in satellite videos. Our deep background prior is obtained by a background reconstruction network, which can learn to reconstruct background from consecutive frames. By applying our deep background prior into model-based methods, a closed-form solution can be obtained via alternating direction method of multipliers (ADMM) and then detection results can be acquired through iterative optimization. More importantly, our background reconstruction network can be trained in an unsupervised way by introducing specifically designed loss, thus relieving the dependence on large-scale labeled dataset. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed method. Ting Liu 0017, Xinyi Ying, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Dense Nested Attention Network for Infrared Small Target DetectionabstractSingle-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds. With the advances of deep learning, CNN-based methods have yielded promising results in generic object detection due to their powerful modeling capability. However, existing CNN-based methods cannot be directly applied to infrared small targets since pooling layers in their networks could lead to the loss of targets in deep layers. To handle this problem, we propose a dense nested attention network (DNA-Net) in this paper. Specifically, we design a dense nested interactive module (DNIM) to achieve progressive interaction among high-level and low-level features. With the repetitive interaction in DNIM, the information of infrared small targets in deep layers can be maintained. Based on DNIM, we further propose a cascaded channel and spatial attention module (CSAM) to adaptively enhance multi-level features. With our DNA-Net, contextual information of small targets can be well incorporated and fully exploited by repetitive fusion and enhancement. Moreover, we develop an infrared small target dataset (namely, NUDT-SIRST) and propose a set of evaluation metrics to conduct comprehensive performance evaluation. Experiments on both public and our self-developed datasets demonstrate the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of probability of detection (${P}_{d}$), false-alarm rate (${F}_{a}$), and intersection of union ($IoU$). Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 4 |
| 2022 | Occlusion-Aware Cost Constructor for Light Field Depth EstimationabstractMatching cost construction is a key step in light field (LF) depth estimation, but was rarely studied in the deep learning era. Recent deep learning-based LF depth estimation methods construct matching cost by sequentially shifting each sub-aperture image (SAI) with a series of pre-defined offsets, which is complex and time-consuming. In this paper, we propose a simple and fast cost constructor to construct matching cost for LF depth estimation. Our cost constructor is composed by a series of convolutions with specifically designed dilation rates. By applying our cost constructor to SAI arrays, pixels under predefined disparities can be integrated and matching cost can be constructed without using any shifting operation. More importantly, the proposed cost constructor is occlusion-aware and can handle occlusions by dynamically modulating pixels from different views. Based on the proposed cost constructor, we develop a deep network for LF depth estimation. Our network ranks first on the commonly used 4D LF benchmark in terms of the mean square error (MSE), and achieves a faster running time than other state-of-the-art methods. Yingqian Wang 0002, Longguang Wang, Zhengyu Liang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 1 |
| 2022 | Learnable Lookup Table for Neural Network QuantizationabstractNeural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(·) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use predefined functions (e.g., exponential function) with learnable parameters to build the quantizer for joint optimization. However, these complicated quantizers introduce considerable computational overhead during inference since activation quantization should be conducted online. In this paper, we formulate the quantization process as a simple lookup operation and propose to learn lookup tables as quantizers. Specifically, we develop differentiable lookup tables and introduce several training strategies for optimization. Our lookup tables can be trained with the network in an end-to-end manner to fit the distributions in different layers and have very small additional computational cost. Comparison with previous methods show that quantized networks using our lookup tables achieve state-of-the-art performance on image classification, image super-resolution, and point cloud classification tasks. Longguang Wang, Yingqian Wang 0002, Li Liu 0002, Wei An 0003, Yulan Guo |
CVPR | 3 |
| 2022 | RSMOT: Remote Sensing Multi-Object Tracking Network with Local Motion Prior for Objects in Satellite VideosabstractMulti-object tracking (MOT) in satellite videos is a new and challenging task. The difficulties stem from the extremely small objects and the low contrast between objects and background. To tackle the challenges of MOT in satellite videos, a multi-object tracking method is proposed in this paper to incorporate the local motion prior into the network. Specifically, we design a local cost volume construction module to obtain tracking offsets between adjacent frames. Based on the tracking offsets, features of previous frames can be propagated to the current frame to incorporate spatio-temporal information. We conduct extensive experiments on videos from Jilin-1 satellite, and the results demonstrate the effectiveness of the proposed method. Shuanglin Wu, Yingqian Wang 0002, Wei An 0003 |
IGARSS | 3 |
| 2022 | Enhancing Mid-Low-Resolution Ship Detection With High-Resolution Feature DistillationabstractTo enhance mid–low-resolution ship detection, existing methods generally use image super-resolution (SR) as a preprocessing step and feed the super-resolved images to the detectors. However, these methods only use high-resolution (HR) images as ground-truth labels to supervise the training of their SR module but overlook the rich HR information in the detection stage. Inspired by the recent advances in knowledge distillation, in this letter, we design a feature distillation framework to fully exploit the information in ground-truth HR images to handle mid–low-resolution ship detection. Our framework consists of a student network and a teacher network. The student network first super-resolves input images using an SR module and then feeds the super-resolved images to the detection module. The teacher network whose architecture is the same as the student detection module directly takes HR images as input to generate HR feature representation and then distills these HR features to the student network through a distillation loss. Using our feature distillation framework, HR images are not only used as ground-truth labels to train the SR module but also provide “ground-truth” features to train the detection module, which enhances the detection performance of the student network. We apply our framework to several popular detectors, includingFCOS,Faster-RCNN,Mask-RCNN, andCascase-RCNN, and conduct extensive ablation studies to validate its effectiveness and generality. Experimental results on the HRSC2016, DOTA, and NWPU VHR-10 datasets demonstrate that, when applying our framework toFaster-RCNN, our method can outperform several state-of-the-art detection methods in terms of mAP50 and mAP75. Shitian He, Huanxin Zou, Yingqian Wang 0002, Runlin Li |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | DARDet: A Dense Anchor-Free Rotated Object Detector in Aerial ImagesabstractRotated object detection in aerial images has received increasing attention for a wide range of applications. However, it is also a challenging task due to the huge variations of scale, rotation, aspect ratio, and densely arranged targets. Most existing methods heavily rely on a large number of predefined anchors with different scales, angles, and aspect ratios, and are optimized with a distance loss. Therefore, these methods are sensitive to anchor hyperparameters and easily suffer from performance degradation caused by boundary discontinuity. To handle this problem, in this letter, we propose a dense anchor-free rotated object detector (DARDet) for rotated object detection in aerial images. Our DARDet directly predicts five parameters of rotated boxes at each foreground pixel of feature maps. We design a new alignment convolution module (ACM) to extract aligned features and introduce a pixels-intersection over union (PIoU) loss for precise and stable regression. Our method achieves state-of-the-art performance on three commonly used aerial objects datasets (i.e., DOTA, HRSC2016, and UCAS-AOD) while keeping high efficiency. Code is available athttps://github.com/zf020114/DARDet. Feng Zhang 0046, Xueying Wang 0001, Shilin Zhou 0001, Yingqian Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Parallax Attention for Unsupervised Stereo Correspondence LearningabstractStereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Light Field Image Super-Resolution With TransformersabstractLight field (LF) image super-resolution (SR) aims at reconstructing high-resolution LF images from their low-resolution counterparts. Although CNN-based methods have achieved remarkable performance in LF image SR, these methods cannot fully model the non-local properties of the 4D LF data. In this paper, we propose a simple but effective Transformer-based method for LF image SR. In our method, an angular Transformer is designed to incorporate complementary information among different views, and a spatial Transformer is developed to capture both local and long-range dependencies within each sub-aperture image. With the proposed angular and spatial Transformers, the beneficial information in an LF can be fully exploited and the SR performance is boosted. We validate the effectiveness of our angular and spatial Transformers through extensive ablation studies, and compare our method to recent state-of-the-art methods on five public LF datasets. Our method achieves superior SR performance with a small model size and low computational cost. Code is available at.1 Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Dense Dual-Attention Network for Light Field Image Super-ResolutionabstractLight field (LF) images can be used to improve the performance of image super-resolution (SR) because both angular and spatial information is available. It is challenging to incorporate distinctive information from different views for LF image SR. Moreover, the long-term information from the previous layers can be weakened as the depth of network increases. In this paper, we propose a dense dual-attention network for LF image SR. Specifically, we design a view attention module to adaptively capture discriminative features across different views and a channel attention module to selectively focus on informative information across all channels. These two modules are fed to two branches and stacked separately in a chain structure for adaptive fusion of hierarchical features and distillation of valid information. Meanwhile, a dense connection is used to fully exploit multi-level information. Extensive experiments demonstrate that our dense dual-attention mechanism can capture informative information across views and channels to improve SR performance. Comparative results show the advantage of our method over state-of-the-art methods on public datasets. Yu Mo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Gated Recurrent Multiattention Network for VHR Remote Sensing Image ClassificationabstractWith the advances of deep learning, many recent CNN-based methods have yielded promising results for image classification. In very high-resolution (VHR) remote sensing images, the contributions of different regions to image classification can vary significantly, because informative areas are generally limited and scattered throughout the whole image. Therefore, how to pay more attention to these informative areas and better incorporate them over long distances are two main challenges to be addressed. In this article, we propose a gated recurrent multiattention neural network (GRMA-Net) to address these problems. Because informative features generally occur at multiple stages in a network (i.e., local texture features at shallow layers and global profile features at deep layers), we use multilevel attention modules to focus on informative regions to extract more discriminative features. Then, these features are arranged as spatial sequences and fed into a deep-gated recurrent unit (GRU) to capture long-range dependency and contextual relationship. We evaluate our method on the UC Merced (UCM), Aerial Image dataset (AID), NWPU-RESISC (NWPU), and Optimal-31 (Optimal) datasets. Experimental results have demonstrated the superior performance of our method as compared to other state-of-the-art methods. Boyang Li 0007, Yulan Guo, Jun-Gang Yang, Longguang Wang, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Spectral-Spatial Deep Support Vector Data Description for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to distinguish anomalies from background-by-background modeling. Deep learning has been applied to HAD and achieves promising detection results. However, there exist several issues that need to be addressed: 1) unrealistic Gaussian assumption on the latent representations may limit its application; 2) deep features are not well-suited to anomaly detection due to the separation between feature learning and anomaly detection; 3) lack of adequate exploitation of spectral-spatial features; 4) negative effect caused by spectral band redundancy. In this article, we propose an end-to-end trainable deep one-class classification network for HAD. Specifically, a minimal enclosing hypersphere is trained to involve the deep features of background samples. These background samples are selected by a density clustering-based method. In this way, feature learning and anomaly detection are incorporated into a unified framework. Meanwhile, there is no explicit Gaussian assumption on the background features. Moreover, due to the complementarity of spectral and spatial features, a novel feature fusion strategy is proposed to fuse spectral and spatial features extracted by a two-stream deep convolutional autoencoder network. Finally, a band attention module is used to automatically learn small weights for redundant bands and thus reduce the negative effect caused by redundant bands. Experimental results on five public datasets demonstrate the superiority of the proposed method compared to several state-of-the-art HAD methods in the detection performance. Kun Li 0029, Qiang Ling 0002, Yao Qin 0002, Yingqian Wang 0002, Yaoming Cai, Zaiping Lin, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Nonconvex Tensor Low-Rank Approximation for Infrared Small Target DetectionabstractInfrared small target detection is an important fundamental task in the infrared system. Therefore, many infrared small target detection methods have been proposed, in which the low-rank model has been used as a powerful tool. However, most low-rank-based methods assign the same weights for different singular values, which will lead to inaccurate background estimation. Considering that different singular values have different importance and should be treated discriminatively, in this article, we propose a nonconvex tensor low-rank approximation (NTLA) method for infrared small target detection. In our method, NTLA regularization adaptively assigns different weights to different singular values for accurate background estimation. Based on the proposed NTLA, we propose asymmetric spatial–temporal total variation (ASTTV) regularization to achieve more accurate background estimation in complex scenes. Compared with the traditional total variation approach, ASTTV exploits different smoothness intensities for spatial and temporal regularization. We design an efficient algorithm to find the optimal solution for our method. Compared with some state-of-the-art methods, the proposed method achieves an improvement in terms of various evaluation metrics. Extensive experimental results in various complex scenes demonstrate that our method has strong robustness and a low false-alarm rate. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yang Sun 0006, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A BenchmarkabstractSatellite video cameras can provide continuous observation for a large-scale area, which is important for many remote sensing applications. However, achieving moving object detection and tracking in satellite videos remains challenging due to the insufficient appearance information of objects and lack of high-quality datasets. In this article, we first build a large-scale satellite video dataset with rich annotations for the task of moving object detection and tracking. This dataset is collected by the Jilin-1 satellite constellation and composed of 47 high-quality videos with 1 646 038 instances of interest for object detection and 3711 trajectories for object tracking. We then introduce a motion modeling baseline to improve the detection rate and reduce false alarms based on accumulative multiframe differencing and robust matrix completion. Finally, we establish the first public benchmark for moving object detection and tracking in satellite videos and extensively evaluate the performance of several representative approaches on our dataset. Comprehensive experimental analyses and insightful conclusions are also provided. The dataset is available athttps://github.com/QingyongHu/VISO. Qingyong Hu, Hao Liu 0061, Feng Zhang 0046, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Arbitrary-Oriented Ship Detection Through Center-Head Point ExtractionabstractShip detection in remote sensing images plays a crucial role in various applications and has drawn increasing attention in recent years. However, existing arbitrary-oriented ship detection methods are generally developed on a set of predefined rotated anchor boxes. These predefined boxes not only lead to inaccurate angle predictions but also introduce extra hyperparameters and high computational cost. Moreover, the prior knowledge of ship size has not been fully exploited by existing methods, which hinders the improvement of their detection accuracy. Aiming at solving the above issues, in this article, we propose a center-head point extraction-based detector (CHPDet) to achieve arbitrary-oriented ship detection in remote sensing images. Our CHPDet formulates arbitrary-oriented ships as rotated boxes with head points that are used to determine the direction. Also, a rotated Gaussian kernel is used to map the annotations into target heatmaps. Keypoint estimation is performed to find the center of ships. Then, the size and head point of the ships are regressed. The orientation-invariant model (OIM) is also used to produce orientation-invariant feature maps. Finally, we use the target size as prior to fine-tune the results. Moreover, we introduce a new dataset for multiclass arbitrary-oriented ship detection in remote sensing images at a fixed ground sample distance (GSD) that is named FGSD2021. Experimental results on FGSD2021 and two other widely used datasets, i.e., HRSC2016 and UCAS-AOD, demonstrate that our CHPDet achieves the state-of-the-art performance and can well distinguish between bow and stern. Code and FGSD2021 dataset are available athttps://github.com/zf020114/CHPDet. Feng Zhang 0046, Xueying Wang 0001, Shilin Zhou 0001, Yingqian Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Exploring Sparsity in Image Super-Resolution for Efficient InferenceabstractCurrent CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for those flat regions. Therefore, existing CNN-based methods involve redundant computation in flat regions, which increases their computational cost and limits their applications on mobile devices. In this paper, we explore the sparsity in image SR to improve inference efficiency of SR networks. Specifically, we develop a Sparse Mask SR (SMSR) network to learn sparse masks to prune redundant computation. Within our SMSR, spatial masks learn to identify "important" regions while channel masks learn to mark redundant channels in those "unimportant" regions. Consequently, redundant computation can be accurately localized and skipped while maintaining comparable performance. It is demonstrated that our SMSR achieves state-of-the-art performance with 41%/33%/27% FLOPs being reduced for ×2/3/4 SR. Code is available at: https://github.com/LongguangWang/SMSR. Longguang Wang, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003, Yulan Guo |
CVPR | 3 |
| 2021 | Unsupervised Degradation Representation Learning for Blind Super-ResolutionabstractMost existing CNN-based super-resolution (SR) methods are developed based on an assumption that the degradation is fixed and known (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation is different from their assumption. To handle various unknown degradations in real-world applications, previous methods rely on degradation estimation to reconstruct the SR image. Nevertheless, degradation estimation methods are usually time-consuming and may lead to SR failure due to large estimation errors. In this paper, we propose an unsupervised degradation representation learning scheme for blind SR without explicit degradation estimation. Specifically, we learn abstract representations to distinguish various degradations in the representation space rather than explicit estimation in the pixel space. Moreover, we introduce a Degradation-Aware SR (DASR) network with flexible adaption to various degradations based on the learned representations. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on both synthetic and real images show that our network achieves state-of-the-art performance for the blind SR task. Code is available at: https://github.com/LongguangWang/DASR. Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 2 |
| 2021 | Learning A Single Network for Scale-Arbitrary Super-ResolutionabstractRecently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., ×2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to learn a scale-arbitrary image SR network from scale-specific networks. Specifically, we develop a plug-in module for existing SR networks to perform scale-arbitrary SR, which consists of multiple scale-aware feature adaption blocks and a scale-aware upsampling layer. Moreover, conditional convolution is used in our plug-in module to generate dynamic scale-aware filters, which enables our network to adapt to arbitrary scale factors. Our plug-in module can be easily adapted to existing networks to realize scale-arbitrary SR with a single model. These networks plugged with our module can produce promising results for non-integer and asymmetric SR while maintaining state-of-the-art performance for SR with integer scale factors. Besides, the additional computational and memory cost of our module is very small. Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
ICCV | 2 |
| 2021 | Shipsrdet: An End-to-End Remote Sensing Ship Detector Using Super-Resolved Feature RepresentationabstractHigh-resolution remote sensing images can provide abundant appearance information for ship detection. Although several existing methods use image super-resolution (SR) approaches to improve the detection performance, they consider image SR and ship detection as two separate processes and overlook the internal coherence between these two correlated tasks. In this paper, we explore the potential benefits introduced by image SR to ship detection, and propose an end-to-end network named ShipSRDet. In our method, we not only feed the super-resolved images to the detector but also integrate the intermediate features of the SR network with those of the detection network. In this way, the informative feature representation extracted by the SR network can be fully used for ship detection. Experimental results on the HRSC dataset validate the effectiveness of our method. Our ShipSRDet can recover the missing details from the input image and achieves promising ship detection performance. Shitian He, Huanxin Zou, Yingqian Wang 0002, Runlin Li |
IGARSS | 3 |
| 2021 | Deep Bilateral Learning for Stereo Image Super-ResolutionabstractBilateral filter has demonstrated its effectiveness in many traditional methods for image restoration tasks. In this letter, we incorporate the idea of bilateral grid processing in a CNN framework and propose a bilateral stereo super-resolution network (BSSRnet). Specifically, we use a parallax-attention module to incorporate information from left and right views to learn content-aware bilateral filters. Then, these bilateral filters are used to recover missing details at different spatial locations while preserving stereo consistency. Our network is fully differentiable and is robust to both content and disparity variations. Comparative results show that our BSSRnet achieves state-of-the-art performance on the Flickr1024, Middlebury, KITTI 2012 and KITTI 2015 datasets. Source code is available at. Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Xinpu Deng |
IEEE Signal Process. Lett. | 3 |
| 2021 | Light Field Image Super-Resolution Using Deformable ConvolutionabstractLight field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among LF images. In this paper, we propose a deformable convolution network (i.e., LF-DFnet) to handle the disparity problem for LF image SR. Specifically, we design an angular deformable alignment module (ADAM) for feature-level alignment. Based on ADAM, we further propose a collect-and-distribute approach to perform bidirectional alignment between the center-view feature and each side-view feature. Using our approach, angular information can be well incorporated and encoded into features of each view, which benefits the SR reconstruction of all LF images. Moreover, we develop a baseline-adjustable LF dataset to evaluate SR performance under different disparity variations. Experiments on both public and our self-developed datasets have demonstrated the superiority of our method. Our LF-DFnet can generate high-resolution images with more faithful details and achieve state-of-the-art reconstruction accuracy. Besides, our LF-DFnet is more robust to disparity variations, which has not been well addressed in literature. Yingqian Wang 0002, Jun-Gang Yang, Longguang Wang, Xinyi Ying, Tianhao Wu 0014, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 1 |
| 2021 | Spatial-Angular Attention Network for Light Field ReconstructionabstractTypical learning-based light field reconstruction methods demand in constructing a large receptive field by deepening their networks to capture correspondences between input views. In this paper, we propose a spatial-angular attention network to perceive non-local correspondences in the light field, and reconstruct high angular resolution light field in an end-to-end manner. Motivated by the non-local attention mechanism (Wang et al., 2018; Zhang et al., 2019), a spatial-angular attention module specifically for the high-dimensional light field data is introduced to compute the response of each query pixel from all the positions on the epipolar plane, and generate an attention map that captures correspondences along the angular dimension. Then a multi-scale reconstruction structure is proposed to efficiently implement the non-local attention in the low resolution feature space, while also preserving the high frequency components in the high-resolution feature space. Extensive experiments demonstrate the superior performance of the proposed spatial-angular attention network for reconstructing sparsely-sampled light fields with Non-Lambertian effects. Gaochang Wu, Yingqian Wang 0002, Yebin Liu, Lu Fang 0001, Tianyou Chai |
IEEE Trans. Image Process. | 2 |
| 2020 | Spatial-Angular Interaction for Light Field Image Super-Resolution
Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
ECCV (23) | 1 |
| 2020 | DeOccNet: Learning to See Through Foreground Occlusions in Light FieldsabstractBackground objects occluded in some views of a light field (LF) camera can be seen by other views. Consequently, occluded surfaces are possible to be reconstructed from LF images. In this paper, we handle the LF de-occlusion (LF-DeOcc) problem using a deep encoder-decoder network (namely, DeOccNet). In our method, sub-aperture images (SAIs) are first given to the encoder to incorporate both spatial and angular information. The encoded representations are then used by the decoder to render an occlusion-free center-view SAI. To the best of our knowledge, DeOccNet is the first deep learning-based LF-DeOcc method. To handle the insufficiency oftraining data, we propose an LF synthesis approach to embed selected occlusion masks into existing LF images. Besides, several synthetic and real-world LFs are developed for performance evaluation. Experimental results show that, after training on the generated data, our DeOccNet can effectively remove foreground occlusions and achieves superior performance as compared to other state-of-the-art methods. Source codes are available at: https://github.com/YingqianWang/DeOccNet. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Longguang Wang, Wei An 0003, Yulan Guo |
WACV | 1 |
| 2020 | High-precision refocusing method with one interpolation for camera array imagesabstractCamera array image refocusing can change the in‐focus region so that objects lying on a specified plane are in focus, whereas objects lying off this plane are blurred. Existing refocusing methods for camera array or light field images usually contain two interpolations. Since interpolation brings distortion, especially on the sharp edge in the images, existing methods are not sufficiently precise. In order to improve the quality of the refocusing result, the authors propose a high‐precision method to refocus camera array images. They first back‐project the pixel coordinates to a corresponding location on the focal plane in the world coordinate. Then they reproject the world coordinates to the pixel coordinates by using the parameters of each camera in the array. After that, they align the images with the focal plane by employing interpolation according to the acquired pixel coordinates. Finally, they get the synthetic image refocused on the focal plane by averaging the resulted images. In the proposed method, only one interpolation is used. So that it alleviates the quality degradation of the refocused image compared to the existing methods. Experiments on real‐world scenes (captured by their self‐developed light field devices) demonstrate that their method can yield better results than the existing methods. Jun-Gang Yang, Yingqian Wang 0002, Chengjin An, Wei An 0003 |
IET Image Process. | 3 |
| 2020 | A Stereo Attention Module for Stereo Image Super-ResolutionabstractIn stereo image super-resolution (SR), exploiting both intra-view and cross-view information is significant but challenging. As existing single image SR (SISR) methods are powerful in intra-view information exploitation, in this letter, we propose a generic stereo attention module (SAM) to extend arbitrary SISR networks for stereo image SR. Specifically, we apply two identical pretrained SISR networks to stereo images. The extracted stereo features at different stages are fed to SAMs to interact cross-view information. Finally, the intra-view and cross-view information is incorporated by SISR networks for stereo image SR. Experiments on the KITTI2012, KITTI2015 and Middlebury datasets have demonstrated the effectiveness of our scheme. Using SAM, we can exploit cross-view information while maintaining the superiority of intra-view information exploitation, resulting in notable performance gain to SISR networks. Moreover, SRResNet equipped with our SAM outperforms the state-of-the-art stereo SR methods. Source code is available at https://github.com/XinyiYing/SAM. Xinyi Ying, Yingqian Wang 0002, Longguang Wang, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 2 |
| 2020 | Deformable 3D Convolution for Video Super-ResolutionabstractThe spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal motion compensation are usually performed sequentially. In this paper, we propose a deformable 3D convolution network (D3Dnet) to incorporate spatio-temporal information from both spatial and temporal dimensions for video SR. Specifically, we introduce deformable 3D convolution (D3D) to integrate deformable convolution with 3D convolution, obtaining both superior spatio-temporal modeling capability and motion-aware modeling flexibility. Extensive experiments have demonstrated the effectiveness of D3D in exploiting spatio-temporal information. Comparative results show that our network achieves state-of-the-art SR performance. Code is available at: https://github.com/XinyiYing/D3Dnet. Xinyi Ying, Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 3 |
| 2019 | Learning Parallax Attention for Stereo Image Super-ResolutionabstractStereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets. Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 2 |
| 2019 | Selective Light Field Refocusing for Camera Arrays Using Bokeh Rendering and SuperresolutionabstractCamera arrays provide spatial and angular information within a single snapshot. With refocusing methods, focal planes can be altered after exposure. In this letter, we propose a light field refocusing method to improve the imaging quality of camera arrays. In our method, the disparity is first estimated. Then, the unfocused region (bokeh) is rendered by using a depth-based anisotropic filter. Finally, the refocused image is produced by a reconstruction-based superresolution approach where the bokeh image is used as a regularization term. Our method can selectively refocus images with focused region being superresolved and bokeh being esthetically rendered. Our method also enables postadjustment of depth of field. We conduct experiments on both public and self-developed datasets. Our method achieves superior visual performance with acceptable computational cost as compared to the other state-of-the-art methods. Yingqian Wang 0002, Jun-Gang Yang, Yulan Guo, Wei An 0003 |
IEEE Signal Process. Lett. | 1 |