EDBT 2026 Demo / reviewers in the wild / expert
Jun-Gang Yang
dblp:208/0174 · also Jungang Yang 0001
· DBLP profile ↗
42ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0002-3127-8705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 12 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diving Into Epipolar Transformers for Light Field Super-Resolution and Disparity EstimationabstractLight field (LF) cameras capture the light rays of a 3D scene from multiple views simultaneously, and thus provide a more immersive experience of the real world as compared to traditional cameras. Although significant progress has been made in various LF image processing tasks, it remains challenging to effectively model the non-local spatial-angular correlations inherent in LF images, particularly when dealing with complex disparity variations. In this paper, we focus on orthogonal epipolar geometry of LF images and propose a generic Epipolar Transformer mechanism that incorporates geometrically meaningful correlations along the epipolar lines. Our Epipolar Transformer mechanism enjoys the following benefits: learning effective and diverse LF feature representations, delivering satisfactory results without redundant architectural designs, and enabling flexible extension to various LF-related tasks with simple adaptations. For LF spatial and angular super-resolution, our methods not only achieve state-of-the-art performance on benchmark datasets, but also demonstrate superior and robust performance on large disparity variations. For disparity estimation, we explore the use of geometry information encoded in our Epipolar Transformer to directly regress the disparity results, effectively avoiding the limitation of a fixed maximum disparity. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Yulan Guo, Li Liu 0002, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Revisiting Subspace Disentangling for Light Field Spatial Super-Resolution
Yingqian Wang 0002, Xueying Wang 0001, Zhengyu Liang, Longguang Wang, Lvli Tian, Jun-Gang Yang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Unsupervised Degradation Representation Learning for Unpaired Restoration of Images and Point CloudsabstractRestoration tasks in low-level vision aim to restore high-quality (HQ) data from their low-quality (LQ) observations. To circumvents the difficulty of acquiring paired data in real scenarios, unpaired approaches that aim to restore HQ data solely on unpaired data are drawing increasing interest. Since restoration tasks are tightly coupled with the degradation model, unknown and highly diverse degradations in real scenarios make learning from unpaired data quite challenging. In this paper, we propose a degradation representation learning scheme to address this challenge. By learning to distinguish various degradations in the representation space, our degradation representations can extract implicit degradation information in an unsupervised manner. Moreover, to handle diverse degradations, we develop degradation-aware (DA) convolutions with flexible adaption to various degradations to fully exploit the degrdation information in the learned representations. Based on our degradation representations and DA convolutions, we introduce a generic framework for unpaired restoration tasks. Based on our framework, we propose UnIRnet and UnPRnet for unpaired image and point cloud restoration tasks, respectively. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on unpaired image and point cloud restoration tasks show that our UnIRnet and UnPRnet achieve state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Graph Laplacian regularization for fast infrared small target detection
Ting Liu 0017, Yongxian Liu, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 3 |
| 2025 | Fixed Relative Pose Prior for Camera Array Self-CalibrationabstractCamera arrays have unique advantages in various computer vision tasks, such as 3D scene reconstruction and depth estimation. For these tasks, precise calibration of sub-cameras is crucial. Since the baselines of sub-cameras are usually small, it is challenging to calibrate the camera array through a single recording of the scene. Consequently, the majority of existing calibration methods address this issue by recording a scene at different spatial locations. However, this approach neglects the prior that the relative pose of the sub-cameras remains unchanged across different locations, which leads to an increase in cumulative reprojection errors. In this letter, we propose to incorporate this fixed relative pose prior to precisely calibrate the camera array. Specifically, we first capture dual-array frames by recording a scene at two spatial locations. Then, we incorporate the fixed relative pose prior to the camera array calibration process by integrating the linear constraint into the organization of sub-aperture images (SAIs). Our method maintains the minimum necessary degrees of freedom for the calibration model, and reduces cumulative reprojection error. Moreover, we develop a real-world light field dataset for comprehensive performance evaluation. Experimental results demonstrate that our method can achieve higher calibration accuracy as compared to existing methods. Our code and dataset are available athttps://github.com/Zhangyaning-NUDT/Fixed-relative-pose-prior-for-camera-array-self-calibration. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Sparsity-Aware Global Channel Pruning for Infrared Small-Target Detection NetworksabstractFor infrared small-target detection, convolutional neural network (CNN)-based methods have demonstrated promising performance. However, due to the small size of the targets, existing infrared small-target detection methods necessitate intricate structures with intensive computation to maintain distinctive features of targets in deep layers, which poses great challenges for deployment on edge devices with constrained resources. The current pruning methods predominantly focus on the optimization of classification networks, with an emphasis on semantic information. Nevertheless, the spatial details crucial for infrared small-target detection are excessively pruned in the shallow layers, resulting in a significant degradation of detection performance. In this article, based on the sparse distribution of small targets in infrared images, we propose a sparsity-aware global channel pruning (SAGCP) framework to optimize infrared small-target detection networks. Specifically, sparse modeling is used to code the target region for the first time, and sparse priors can be induced into the feature map for the identification of redundant channels. Without the need for extra structures or intricate criteria to identify redundant channels, our pruning method can leverage the inherent properties of infrared small targets to extract more robust features and obtain more compact models. When SAGCP is applied to the existing infrared small-target detection methods, the pruned network is superior in model efficiency and detection performance. For example, when applying our method to DNA-Net, the pruned model can achieve a 72.34% reduction in parameters, and a 57.49% decrease in floating point operations (FLOPs), but a 2.02% increase in intersection over union (IoU). Shuanglin Wu, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Real-World Light Field Image Super-Resolution Via Degradation ModulationabstractRecent years have witnessed the great advances of deep neural networks (DNNs) in light field (LF) image super-resolution (SR). However, existing DNN-based LF image SR methods are developed on a single fixed degradation (e.g., bicubic downsampling), and thus cannot be applied to super-resolve real LF images with diverse degradation. In this article, we propose a simple yet effective method for real-world LF image SR. In our method, a practical LF degradation model is developed to formulate the degradation process of real LF images. Then, a convolutional neural network is designed to incorporate the degradation prior into the SR process. By training on LF images using our formulated degradation, our network can learn to modulate different degradation while incorporating both spatial and angular information in LF images. Extensive experiments on both synthetically degraded and real-world LF images demonstrate the effectiveness of our method. Compared with existing state-of-the-art single and LF image SR methods, our method achieves superior SR performance under a wide range of degradation, and generalizes better to real LF images. Codes and models are available at https://yingqianwang.github.io/LF-DMnet/. Yingqian Wang 0002, Zhengyu Liang, Longguang Wang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-ResolutionabstractExploiting spatial-angular correlation is crucial to light field (LF) image super-resolution (SR), but is highly challenging due to its non-local property caused by the disparities among LF images. Although many deep neural networks (DNNs) have been developed for LF image SR and achieved continuously improved performance, existing methods cannot well leverage the long-range spatial-angular correlation and thus suffer a significant performance drop when handling scenes with large disparity variations. In this paper, we propose a simple yet effective method to learn the non-local spatial-angular correlation for LF image SR. In our method, we adopt the epipolar plane image (EPI) representation to project the 4D spatial-angular correlation onto multiple 2D EPI planes, and then develop a Transformer network with repetitive self-attention operations to learn the spatial-angular correlation by modeling the dependencies between each pair of EPI pixels. Our method can fully incorporate the information from all angular views while achieving a global receptive field along the epipolar line. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Comparative results on five public datasets show that our method not only achieves state-of-the-art SR performance but also performs robust to disparity variations. Code is publicly available at https://github.com/ZhengyuLiang24/EPIT. Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001, Yulan Guo |
ICCV | 4 |
| 2023 | Disentangling Light Fields for Super-Resolution and Disparity EstimationabstractLight field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been proposed for various LF image processing tasks. However, it is challenging for CNNs to effectively process LF images since the spatial and angular information are highly inter-twined with varying disparities. In this paper, we propose a generic mechanism to disentangle these coupled information for LF image processing. Specifically, we first design a class of domain-specific convolutions to disentangle LFs from different dimensions, and then leverage these disentangled features by designing task-specific modules. Our disentangling mechanism can well incorporate the LF structure prior and effectively handle 4D LF data. Based on the proposed mechanism, we develop three networks (i.e., DistgSSR, DistgASR and DistgDisp) for spatial super-resolution, angular super-resolution and disparity estimation. Experimental results show that our networks achieve state-of-the-art performance on all these three tasks, which demonstrates the effectiveness, efficiency, and generality of our disentangling mechanism. Project page: https://yingqianwang.github.io/DistgLF/. Yingqian Wang 0002, Longguang Wang, Gaochang Wu, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Combining Deep Denoiser and Low-rank Priors for Infrared Small Target DetectionabstractMany existing low-rank methods have achieved good detection performance in uniform scenes, but they suffer from a high false alarm rate in complex noisy scenes. Therefore, it is important to improve the detection performance of low-rank models in noisy scenes. In this paper, we first formulate an implicit regularizer by plugging a denoising neural network (termed as deep denoiser), which can learn deep image priors from a large number of natural images. Then, we use the weighted sum of weighted tensor nuclear norm for more accurate background estimation. Finally, alternating direction multiplier method is used to solve the model under the plug-and-play framework. By integrating low-rank prior with deep denoiser prior, our model achieves higher accuracy. Experiments on different scenes demonstrate that our method achieves an improved performance in terms of visual effects and quantitative metrics. Specially, the overall accuracy of AUC value (AUCOA) achieved by the proposed method on Sequences 1-6 are 1.24%, 1.16%, 0.63%, 1.9%, 0.82%, 2.06% higher than those achieved by the second top performing methods, respectively. Ting Liu 0017, Jun-Gang Yang, Yingqian Wang 0002, Wei An 0003 |
Pattern Recognit. | 3 |
| 2023 | Learning scalable dynamic filter in convolutional networks
Shuanglin Wu, Xinyi Ying, Longguang Wang, Jun-Gang Yang, Wei An 0003 |
Pattern Recognit. Lett. | 5 |
| 2023 | Not All Patches Are Equal: Hierarchical Dataset Condensation for Single Image Super-ResolutionabstractAlthough the performance of single image super-resolution (SR) has been significantly improved with deep neural networks, existing methods commonly require millions of iterations for training, which not only limits their training efficiency, but also causes considerable energy consumption. In this paper, we comprehensively study the redundancy of existing training datasets and reveal that not all patches are equal for SR network training. We observe that a large percentage of patches with low textures or similar textures lead to high computation costs but make low contributions to SR performance. Then, we propose a dataset condensation method to remove these redundant patches hierarchically. Extensive experiments demonstrate that our dataset condensation method can effectively reduce the redundancy of SR datasets with a 90% condensation rate on DIV2K. With our condensed dataset, baseline networks can achieve significant improvement in terms of training efficiency while maintaining competitive accuracy. Codes are available athttps://github.com/QingtangDing/DCSR. Qingtang Ding, Zhengyu Liang, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang |
IEEE Signal Process. Lett. | 5 |
| 2023 | Infrared Small Target Detection via Nonconvex Tensor Tucker Decomposition With Factor PriorabstractInfrared small target detection in complex scenes is an important but challenging research hotspot in infrared early warning fields. Previous studies have proved that low-rank Tucker decomposition (TD) achieves good detection performance in complex scenes. However, a key limitation of existing low-rank TD methods is that the rank needs to be set in advance, and an inaccurate predefined rank can lead to performance degradation. Inspired by the theorem that n-rank is upper bounded by the rank of each Tucker factor matrix, we propose a nonconvex tensor TD model with factor prior for infrared small target detection. In our method, we use a logdet-based function to constrain the latent factors of low-rank TD, which avoids empirical rank selection and sufficiently uses the latent data structure information in the factor matrix. Meanwhile, performing singular value decomposition (SVD) calculations on small factor matrices can reduce computational complexity. Then, group sparsity regularized total variation is used to better exploit the shared sparse pattern of difference images, which helps better remove background clutter and obtain better detection results. Finally, the proposed method is efficiently solved by the well-designed alternating direction method of multipliers (ADMM). Extensive experimental results demonstrate that our method is more effective and robust in complex scenes than other state-of-the-art methods. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Representative Coefficient Total Variation for Efficient Infrared Small Target DetectionabstractLow-rank and sparse decomposition based models are powerful and robust tools for infrared small target detection. However, due to the calculation of singular value decomposition (SVD) and the optimization of complex regularization terms, existing low-rank models often suffer from high computational complexity. To solve this problem, based on the theorem that representative coefficient matrix obtained by orthogonal transformation of data matrix can inherit the spatial structure of data matrix, we propose a representative coefficient total variation (RCTV) method for efficient infrared small target detection. In our method, we use total variational to constraint representative coefficient matrix instead of data matrix to describe local smooth prior, which helps remove noise and reduce computational complexity. Meanwhile, we control the number of columns in the representative coefficient matrix to maintain the low-rank characteristics of background, which avoids SVD calculation and improves detection efficiency. Therefore, the RCTV regularization can simultaneously describe local smooth prior and low-rank prior. Moreover, to better enhance the sparsity of targets and distinguish sparse non-target points, we use the log-sum function to adaptively assign weights to targets. It helps obtain more accurate detection performance. The proposed model is efficiently solved by the alternating direction multiplier method (ADMM). A large number of experiments show that the proposed method is superior to existing low-rank methods in both detection accuracy and efficiency. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MTU-Net: Multilevel TransUNet for Space-Based Infrared Tiny Ship DetectionabstractSpace-based infrared tiny ship detection aims at separating tiny ships from the images captured by Earth-orbiting satellites. Due to the extremely large image coverage area (e.g., thousands of square kilometers), candidate targets in these images are much smaller, dimer, and more changeable than those targets observed by aerial- and land-based imaging devices. Existing short imaging distance-based infrared datasets and target detection methods cannot be well adopted to the space-based surveillance task. To address these problems, we develop a space-based infrared tiny ship detection dataset (namely, NUDT-SIRST-Sea) with 48 space-based infrared images and$17\,598$pixel-level tiny ship annotations. Each image covers about$10\,000$km2of area with$10 \ 000\,\, \times \ 10 \ 000$pixels. Considering the extreme characteristics (e.g., small, dim, and changeable) of those tiny ships in such challenging scenes, we propose a multilevel TransUNet (MTU-Net) in this article. Specifically, we design a vision Transformer (ViT) convolutional neural network (CNN) hybrid encoder to extract multilevel features. Local feature maps are first extracted by several convolution layers and then fed into the multilevel feature extraction module [multilevel ViT module (MVTM)] to capture long-distance dependency. We further propose a copy–rotate–resize–paste (CRRP) data augmentation approach to accelerate the training phase, which effectively alleviates the issue of sample imbalance between targets and background. Besides, we design a FocalIoU loss to achieve both target localization and shape description. Experimental results on the NUDT-SIRST-Sea dataset show that our MTU-Net outperforms traditional and existing deep learning-based single-frame infrared small target (SIRST) methods in terms of probability of detection, false alarm rate, and intersection over union. Our code is available athttps://github.com/TianhaoWu16/Multi-level-TransUNet-for-Space-based-Infrared-Tiny-ship-Detection Tianhao Wu 0014, Boyang Li 0007, Yihang Luo, Yingqian Wang 0002, Ting Liu 0017, Jun-Gang Yang, Wei An 0003, Yulan Guo |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | RepISD-Net: Learning Efficient Infrared Small-Target Detection Network via Structural Re-ParameterizationabstractInfrared small target detection is a challenging task for deep learning-based methods because targets tend to disappear in the deep layers. To handle this problem, existing deep neural networks usually apply various dense and skip connections for feature maintenance. Although these well-designed networks have achieved good detection performance, the complex network structures reduce their efficiency. In this paper, we propose a simple yet efficient network (RepISD-Net) for infrared small target detection. The core of our RepISD-Net is to use different network architectures but equivalent model parameters for training and inference, respectively. Specifically, in the training phase, we design a parallel multi-branch edge compensation block (ECB) to enhance the local salient features and capture finer contour characteristic of infrared small targets. In the inference phase, the multi-branch topology structures are merged into a single branch with only cascaded 3×3 convolutions for fast inference. We conduct extensive experiments on several public datasets to validate the effectiveness of our method. Experimental results demonstrate that our RepISD-Net can achieve comparable or even better detection performance with significant acceleration in inference speed as compared to state-of-the-art infrared small target detection methods. Code is submitted for review and will be released upon acceptance. Shuanglin Wu, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Occlusion-Aware Cost Constructor for Light Field Depth EstimationabstractMatching cost construction is a key step in light field (LF) depth estimation, but was rarely studied in the deep learning era. Recent deep learning-based LF depth estimation methods construct matching cost by sequentially shifting each sub-aperture image (SAI) with a series of pre-defined offsets, which is complex and time-consuming. In this paper, we propose a simple and fast cost constructor to construct matching cost for LF depth estimation. Our cost constructor is composed by a series of convolutions with specifically designed dilation rates. By applying our cost constructor to SAI arrays, pixels under predefined disparities can be integrated and matching cost can be constructed without using any shifting operation. More importantly, the proposed cost constructor is occlusion-aware and can handle occlusions by dynamically modulating pixels from different views. Based on the proposed cost constructor, we develop a deep network for LF depth estimation. Our network ranks first on the commonly used 4D LF benchmark in terms of the mean square error (MSE), and achieves a faster running time than other state-of-the-art methods. Yingqian Wang 0002, Longguang Wang, Zhengyu Liang, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 4 |
| 2022 | Parallax Attention for Unsupervised Stereo Correspondence LearningabstractStereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Light Field Image Super-Resolution With TransformersabstractLight field (LF) image super-resolution (SR) aims at reconstructing high-resolution LF images from their low-resolution counterparts. Although CNN-based methods have achieved remarkable performance in LF image SR, these methods cannot fully model the non-local properties of the 4D LF data. In this paper, we propose a simple but effective Transformer-based method for LF image SR. In our method, an angular Transformer is designed to incorporate complementary information among different views, and a spatial Transformer is developed to capture both local and long-range dependencies within each sub-aperture image. With the proposed angular and spatial Transformers, the beneficial information in an LF can be fully exploited and the SR performance is boosted. We validate the effectiveness of our angular and spatial Transformers through extensive ablation studies, and compare our method to recent state-of-the-art methods on five public LF datasets. Our method achieves superior SR performance with a small model size and low computational cost. Code is available at.1 Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Dense Dual-Attention Network for Light Field Image Super-ResolutionabstractLight field (LF) images can be used to improve the performance of image super-resolution (SR) because both angular and spatial information is available. It is challenging to incorporate distinctive information from different views for LF image SR. Moreover, the long-term information from the previous layers can be weakened as the depth of network increases. In this paper, we propose a dense dual-attention network for LF image SR. Specifically, we design a view attention module to adaptively capture discriminative features across different views and a channel attention module to selectively focus on informative information across all channels. These two modules are fed to two branches and stacked separately in a chain structure for adaptive fusion of hierarchical features and distillation of valid information. Meanwhile, a dense connection is used to fully exploit multi-level information. Extensive experiments demonstrate that our dense dual-attention mechanism can capture informative information across views and channels to improve SR performance. Comparative results show the advantage of our method over state-of-the-art methods on public datasets. Yu Mo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Gated Recurrent Multiattention Network for VHR Remote Sensing Image ClassificationabstractWith the advances of deep learning, many recent CNN-based methods have yielded promising results for image classification. In very high-resolution (VHR) remote sensing images, the contributions of different regions to image classification can vary significantly, because informative areas are generally limited and scattered throughout the whole image. Therefore, how to pay more attention to these informative areas and better incorporate them over long distances are two main challenges to be addressed. In this article, we propose a gated recurrent multiattention neural network (GRMA-Net) to address these problems. Because informative features generally occur at multiple stages in a network (i.e., local texture features at shallow layers and global profile features at deep layers), we use multilevel attention modules to focus on informative regions to extract more discriminative features. Then, these features are arranged as spatial sequences and fed into a deep-gated recurrent unit (GRU) to capture long-range dependency and contextual relationship. We evaluate our method on the UC Merced (UCM), Aerial Image dataset (AID), NWPU-RESISC (NWPU), and Optimal-31 (Optimal) datasets. Experimental results have demonstrated the superior performance of our method as compared to other state-of-the-art methods. Boyang Li 0007, Yulan Guo, Jun-Gang Yang, Longguang Wang, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Nonconvex Tensor Low-Rank Approximation for Infrared Small Target DetectionabstractInfrared small target detection is an important fundamental task in the infrared system. Therefore, many infrared small target detection methods have been proposed, in which the low-rank model has been used as a powerful tool. However, most low-rank-based methods assign the same weights for different singular values, which will lead to inaccurate background estimation. Considering that different singular values have different importance and should be treated discriminatively, in this article, we propose a nonconvex tensor low-rank approximation (NTLA) method for infrared small target detection. In our method, NTLA regularization adaptively assigns different weights to different singular values for accurate background estimation. Based on the proposed NTLA, we propose asymmetric spatial–temporal total variation (ASTTV) regularization to achieve more accurate background estimation in complex scenes. Compared with the traditional total variation approach, ASTTV exploits different smoothness intensities for spatial and temporal regularization. We design an efficient algorithm to find the optimal solution for our method. Compared with some state-of-the-art methods, the proposed method achieves an improvement in terms of various evaluation metrics. Extensive experimental results in various complex scenes demonstrate that our method has strong robustness and a low false-alarm rate. Ting Liu 0017, Jun-Gang Yang, Boyang Li 0007, Yang Sun 0006, Yingqian Wang 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Unsupervised Degradation Representation Learning for Blind Super-ResolutionabstractMost existing CNN-based super-resolution (SR) methods are developed based on an assumption that the degradation is fixed and known (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation is different from their assumption. To handle various unknown degradations in real-world applications, previous methods rely on degradation estimation to reconstruct the SR image. Nevertheless, degradation estimation methods are usually time-consuming and may lead to SR failure due to large estimation errors. In this paper, we propose an unsupervised degradation representation learning scheme for blind SR without explicit degradation estimation. Specifically, we learn abstract representations to distinguish various degradations in the representation space rather than explicit estimation in the pixel space. Moreover, we introduce a Degradation-Aware SR (DASR) network with flexible adaption to various degradations based on the learned representations. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on both synthetic and real images show that our network achieves state-of-the-art performance for the blind SR task. Code is available at: https://github.com/LongguangWang/DASR. Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 5 |
| 2021 | Cgan-Net: Class-Guided Asymmetric Non-Local Network for Real-Time Semantic SegmentationabstractBy introducing various non-local blocks to capture the long-range dependencies, remarkable progress has been achieved in semantic segmentation recently. However, the improvement in segmentation accuracy usually comes at the price of significant reductions in network efficiency, as non-local block usually requires expensive computation and memory cost for dense pixel-to-pixel correlation. In this paper, we introduce a Class-Guided Asymmetric Non-local Network (CGAN-Net) to enhance the class-discriminability in learned feature map, while maintaining real-time efficiency. The key to our approach is to calculate the dense similarity matrix in coarse semantic prediction maps, instead of the high-dimensional latent feature map. This is not only computationally and memory efficient, but helps to learn query-dependent global context. Experiments conducted on Cityscape and CamVid demonstrate the compelling performance of our CGAN-Net. In particular, our network achieves 76.8% mean IoU on the Cityscapes test set with a speed of 38 FPS for 1024×2048 images on a single Tesla V100 GPU. Qingyong Hu, Jun-Gang Yang, Yulan Guo |
ICASSP | 3 |
| 2021 | Learning A Single Network for Scale-Arbitrary Super-ResolutionabstractRecently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., ×2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to learn a scale-arbitrary image SR network from scale-specific networks. Specifically, we develop a plug-in module for existing SR networks to perform scale-arbitrary SR, which consists of multiple scale-aware feature adaption blocks and a scale-aware upsampling layer. Moreover, conditional convolution is used in our plug-in module to generate dynamic scale-aware filters, which enables our network to adapt to arbitrary scale factors. Our plug-in module can be easily adapted to existing networks to realize scale-arbitrary SR with a single model. These networks plugged with our module can produce promising results for non-integer and asymmetric SR while maintaining state-of-the-art performance for SR with integer scale factors. Besides, the additional computational and memory cost of our module is very small. Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
ICCV | 4 |
| 2021 | Infrared Dim and Small Target Detection via Multiple Subspace Learning and Spatial-Temporal Patch-Tensor ModelabstractRobust detection of infrared small and dim targets with highly heterogeneous backgrounds plays an indispensable role in infrared search and tracking (IRST) system, which is still a challenging problem. In this article, a novel method based on multisubspace learning and spatial-temporal tensor data structure is aimed to solve this problem. First, a tensor data structure is constructed to use inner correlation in spatial and temporal domain of the infrared image sequence. Second, in consideration of the complex and heterogeneous backgrounds in infrared images, the proposed method promotes the multisubspace property to tensor domain to separate the target and the background more accurately. Finally, an efficient and effective optimization algorithm based on alternating direction method of multipliers (ADMM) is designed to solve this problem. Experimental results on various and real scenes demonstrate the superiority of the proposed method compared to other five baseline methods. Yang Sun 0006, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Light Field Image Super-Resolution Using Deformable ConvolutionabstractLight field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among LF images. In this paper, we propose a deformable convolution network (i.e., LF-DFnet) to handle the disparity problem for LF image SR. Specifically, we design an angular deformable alignment module (ADAM) for feature-level alignment. Based on ADAM, we further propose a collect-and-distribute approach to perform bidirectional alignment between the center-view feature and each side-view feature. Using our approach, angular information can be well incorporated and encoded into features of each view, which benefits the SR reconstruction of all LF images. Moreover, we develop a baseline-adjustable LF dataset to evaluate SR performance under different disparity variations. Experiments on both public and our self-developed datasets have demonstrated the superiority of our method. Our LF-DFnet can generate high-resolution images with more faithful details and achieve state-of-the-art reconstruction accuracy. Besides, our LF-DFnet is more robust to disparity variations, which has not been well addressed in literature. Yingqian Wang 0002, Jun-Gang Yang, Longguang Wang, Xinyi Ying, Tianhao Wu 0014, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 2 |
| 2020 | Spatial-Angular Interaction for Light Field Image Super-Resolution
Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
ECCV (23) | 3 |
| 2020 | DeOccNet: Learning to See Through Foreground Occlusions in Light FieldsabstractBackground objects occluded in some views of a light field (LF) camera can be seen by other views. Consequently, occluded surfaces are possible to be reconstructed from LF images. In this paper, we handle the LF de-occlusion (LF-DeOcc) problem using a deep encoder-decoder network (namely, DeOccNet). In our method, sub-aperture images (SAIs) are first given to the encoder to incorporate both spatial and angular information. The encoded representations are then used by the decoder to render an occlusion-free center-view SAI. To the best of our knowledge, DeOccNet is the first deep learning-based LF-DeOcc method. To handle the insufficiency oftraining data, we propose an LF synthesis approach to embed selected occlusion masks into existing LF images. Besides, several synthetic and real-world LFs are developed for performance evaluation. Experimental results show that, after training on the generated data, our DeOccNet can effectively remove foreground occlusions and achieves superior performance as compared to other state-of-the-art methods. Source codes are available at: https://github.com/YingqianWang/DeOccNet. Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Longguang Wang, Wei An 0003, Yulan Guo |
WACV | 3 |
| 2020 | High-precision refocusing method with one interpolation for camera array imagesabstractCamera array image refocusing can change the in‐focus region so that objects lying on a specified plane are in focus, whereas objects lying off this plane are blurred. Existing refocusing methods for camera array or light field images usually contain two interpolations. Since interpolation brings distortion, especially on the sharp edge in the images, existing methods are not sufficiently precise. In order to improve the quality of the refocusing result, the authors propose a high‐precision method to refocus camera array images. They first back‐project the pixel coordinates to a corresponding location on the focal plane in the world coordinate. Then they reproject the world coordinates to the pixel coordinates by using the parameters of each camera in the array. After that, they align the images with the focal plane by employing interpolation according to the acquired pixel coordinates. Finally, they get the synthetic image refocused on the focal plane by averaging the resulted images. In the proposed method, only one interpolation is used. So that it alleviates the quality degradation of the refocused image compared to the existing methods. Experiments on real‐world scenes (captured by their self‐developed light field devices) demonstrate that their method can yield better results than the existing methods. Jun-Gang Yang, Yingqian Wang 0002, Chengjin An, Wei An 0003 |
IET Image Process. | 1 |
| 2019 | Learning Parallax Attention for Stereo Image Super-ResolutionabstractStereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets. Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo |
CVPR | 5 |
| 2019 | Selective Light Field Refocusing for Camera Arrays Using Bokeh Rendering and SuperresolutionabstractCamera arrays provide spatial and angular information within a single snapshot. With refocusing methods, focal planes can be altered after exposure. In this letter, we propose a light field refocusing method to improve the imaging quality of camera arrays. In our method, the disparity is first estimated. Then, the unfocused region (bokeh) is rendered by using a depth-based anisotropic filter. Finally, the refocused image is produced by a reconstruction-based superresolution approach where the bokeh image is used as a regularization term. Our method can selectively refocus images with focused region being superresolved and bokeh being esthetically rendered. Our method also enables postadjustment of depth of field. We conduct experiments on both public and self-developed datasets. Our method achieves superior visual performance with acceptable computational cost as compared to the other state-of-the-art methods. Yingqian Wang 0002, Jun-Gang Yang, Yulan Guo, Wei An 0003 |
IEEE Signal Process. Lett. | 2 |
| 2019 | Learning Multi-View Representation With LSTM for 3-D Shape Recognition and RetrievalabstractShape representation for 3-D models is an important topic in computer vision, multimedia analysis, and computer graphics. Recent multiview-based methods demonstrate promising performance for 3-D shape recognition and retrieval. However, most multiview-based methods ignore the correlations of multiple views or suffer from high computional cost. In this paper, we propose a novel multiview-based network architecture for 3-D shape recognition and retrieval. Our network combines convolutional neural networks (CNNs) with long short-term memory (LSTM) to exploit the correlative information from multiple views. Well-pretrained CNNs with residual connections are first used to extract a low-level feature of each view image rendered from a 3-D shape. Then, a LSTM and a sequence voting layer are employed to aggregate these features into a shape descriptor. The highway network and a three-step training strategy are also adopted to boost the optimization of the deep network. Experimental results on two public datasets demonstrate that the proposed method achieves promising performance for 3-D shape recognition and the state-of-the-art performance for the 3-D shape retrieval. Chao Ma 0014, Yulan Guo, Jun-Gang Yang, Wei An 0003 |
IEEE Trans. Multim. | 3 |
| 2017 | An Attitude Jitter Correction Method for Multispectral Parallax Imagery Based on Compressive SensingabstractAttitude jitter is a common problem for high-resolution earth-observation satellites and can diminish the geo-positioning and mapping performance of observed images. It is especially necessary to address this problem when high-performance attitude measurements are unavailable. Therefore, an attitude jitter correction method for multispectral parallax imagery that utilizes the compressive-sensing technology is proposed in this letter. In the proposed method, the attitude jitter is estimated from the parallax disparities of different band images, and then the image displacement caused by attitude jitter can be corrected. Using the normalized cross correlation method and compressive-sensing technology, the proposed method can deal with the condition of texture-feature deficiency in the partial image. The multispectral images of the Terra and ZY-3 satellites are used as experimental data to evaluate the proposed method. The registration errors of different bands are greatly reduced in both the cross- and along-track directions, and the experiment results indicate that the proposed method is effective for correcting the attitude jitter of both satellites. Jun Chen 0007, Jun-Gang Yang, Wei An 0003, Zhi-Jie Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Compressed Sensing Radar Imaging With Compensation of Observation Position ErrorabstractCompressed sensing (CS) based radar imaging requires the use of a mathematical model of the observation process. Inaccuracies in the observation model may cause defocusing in the reconstructed images. In the observation process, the observation positions are usually not known perfectly. Imperfect knowledge of the observation positions is a major source of model errors in imaging. In this paper, a method is proposed to compensate the observation position errors in CS-based radar imaging. Instead of treating the observation-position-induced model errors as phase errors in the data, the proposed method can determine the observation position errors as part of the imaging process. It uses an iterative algorithm, which cycles through steps of target reconstruction and observation position error estimation and compensation. The proposed method can estimate the observation position errors accurately, and the reconstruction quality of the target images can be improved significantly. Simulation results and experimental results from rail-mounted radar and airborne synthetic aperture radar are presented to show the effectiveness of the proposed method. Jun-Gang Yang, Xiaotao Huang 0001, John S. Thompson, Tian Jin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Sparse MIMO Array Forward-Looking GPR Imaging Based on Compressed Sensing in Clutter EnvironmentabstractThis paper presents a sparse multiple-input and multiple-output (MIMO) array and sparse frequency ground-penetrating radar (GPR) imaging scheme based on compressed sensing (CS). Since the targets of interest for GPR are usually sparse, the number of the MIMO array elements and frequencies can be reduced using CS theory. Thus, the system complexity and data acquisition time can be reduced accordingly. Considering the serious clutter in forward-looking GPR, we propose two methods for the CS reconstruction in clutter environment. The first one is a clutter suppression preprocessing method, which can effectively suppress the azimuth clutter and short range clutter outside the reconstruction region and significantly improve the reconstruction result. The second one is to determine the regularization parameter for the CS reconstruction in clutter environment. We refer to this reconstruction process as basis pursuit declutter. The proposed imaging scheme can produce pointlike and less cluttered images of sparse targets using fewer array elements and frequencies. Results from simulated data, trihedral reflector, and real buried land mine experimental data are presented to show the validity of the proposed methods. The experimental data are acquired by the vehicle-mounted stepped-frequency forward-looking ground-penetrating virtual aperture radar, which is designed and developed by the National University of Defense Technology. Jun-Gang Yang, Tian Jin 0001, Xiaotao Huang 0001, John S. Thompson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Random-Frequency SAR Imaging Based on Compressed SensingabstractStepped-frequency waveforms can achieve an ultrawide bandwidth by using a sequence of single-frequency pulses. The advantages of stepped-frequency waveforms are low hardware requirements and high resolution. However, the stepped-frequency waveform requires a long time period to transmit the signals, which limits its application in synthetic aperture radar (SAR). The available imaging range width is usually very narrow, unless the range and azimuth resolutions are both decreased. In this paper, a random-frequency SAR imaging scheme based on compressed sensing is proposed. If the targets are sparse or compressible, it is sufficient to transmit only a small number of random frequencies to reconstruct the image of the targets. This means that the limitations of the stepped-frequency technique for SAR can be overcome. The available imaging range width can be enlarged significantly, while the range and azimuth resolutions are both maintained. Random undersampling is very easy to implement for both range and azimuth dimensions, and no new hardware components are needed. Simulation and experimental results are presented to demonstrate the validity of the proposed method. Jun-Gang Yang, John S. Thompson, Xiaotao Huang 0001, Tian Jin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Segmented Reconstruction for Compressed Sensing SAR ImagingabstractThe compressed sensing (CS) synthetic aperture radar (SAR) imaging scheme can use random undersampled data to reconstruct images of sparse or compressible targets. However, compared to Nyquist sampling, the cost of the CS imaging scheme is the long reconstruction time, particularly for the conventional reconstruction strategy, which reconstructs the whole scene in one process. It also needs a large memory to access the sensing matrix used for reconstruction. In this paper, a segmented reconstruction strategy for the CS SAR imaging scheme is proposed. The whole scene is split into a set of small subscenes, so that the reconstruction time can be reduced significantly. The proposed method also needs much less memory for computation than the conventional method. In this proposed method, the range profiles are reconstructed first, and then, the range profiles can be split into subpatches. Subscenes can be reconstructed by using the subpatch data, and the whole scene can be obtained by combining the reconstructed subscenes. Simulation and experimental results are shown to demonstrate the validity of the proposed method. Jun-Gang Yang, John S. Thompson, Xiaotao Huang 0001, Tian Jin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | FMCW radar near field three-dimensional imagingabstractA system with 3-D imaging capability can be implemented by using a frequency modulated continuous wave (FMCW) radar which synthesizes a two-dimensional (2-D) planar aperture. A millimeter-wave FMCW three-dimensional (3-D) imaging system can be used for the detection of concealed weapons and contrabands at airports or other security checkpoints, since millimeter-wave can readily penetrate common clothing material. A 3-D image can be formed by coherently integrating the backscatter data over the measured frequency bandwidth and the two spatial coordinates of the 2-D synthetic aperture. This paper presents a 3-D imaging algorithm for near field FMCW radar. This algorithm is an extension of the 2-D range migration algorithm (RMA). We derive the formulation in detail by using the principle of stationary phase (POSP). A 3-D version Stolt interpolation is used in this algorithm. Accurate image reconstruction and high computational efficiency of this algorithm are demonstrated through simulation results. Jun-Gang Yang, John S. Thompson, Xiaotao Huang 0001, Tian Jin 0001 |
ICC | 1 |
| 2012 | Synthetic Aperture Radar Imaging Using Stepped Frequency WaveformabstractThis paper presents a synthetic aperture radar (SAR) imaging system using a stepped frequency waveform. The main problem in stepped frequency SAR is the range difference in one sequence of pulses. This paper analyzes the influence of range difference in one sequence of pulses in detail and proposes a method to compensate this influence in the Doppler domain. A Stolt interpolation for the stepped frequency signal is used to focus the data accurately. The parameters for the stepped frequency SAR are analyzed, and some criteria are proposed for system design. Finally, simulation and experimental results are presented to demonstrate the validity of the proposed method. The experimental data are collected by using a stepped frequency radar mounted on a rail. Jun-Gang Yang, Xiaotao Huang 0001, Tian Jin 0001, John S. Thompson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | New Approach for SAR Imaging of Ground Moving Targets Based on a Keystone TransformabstractWe propose here a new approach for synthetic aperture radar (SAR) imaging of ground moving targets. The unique characteristic of this approach is that range curvature (i.e., quadratic range migration) can be corrected by a simple processing step. A keystone transform is used to correct range walk (i.e., linear range migration) for all targets without knowing their velocities, and the range curvature is corrected in the range-Doppler domain. The advantage of this approach is that it is simple to implement and can correct range curvature for all targets in one processing step, so that it is computationally efficient. Simulation and experimental SAR data processing results are presented to demonstrate the validity of the proposed approach. Jun-Gang Yang, Xiaotao Huang 0001, Tian Jin 0001, John S. Thompson |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2011 | An Interpolated Phase Adjustment by Contrast Enhancement Algorithm for SARabstractPhase adjustment by contrast enhancement (PACE) is an autofocus algorithm that is capable of performance that is unattainable by conventional techniques. It is a nonparametric method that requires no constraints on the type of phase error to be measured. The algorithm does not require special data culling techniques or the presence of isolated scatterers. However, the drawback of PACE algorithm is that the number of estimated variables is very large; it leads to a long computational time. The azimuth sampling frequency is commonly much bigger than the bandwidth of phase error in SAR image so that we can estimate part of the phase error variables and then obtain the whole variables by interpolation; this induces the interpolated phase adjustment by contrast enhancement (IPACE) algorithm. The IPACE algorithm can remarkably reduce the computational time while maintaining the accuracy. This letter has derived the detailed processing of IPACE, and the results of the experiments using real SAR data are presented to show the validity of the proposed algorithm. Jun-Gang Yang, Xiaotao Huang 0001, Tian Jin 0001, Guoyi Xue |
IEEE Geosci. Remote. Sens. Lett. | 1 |