Yongqiang Zhao 0001

dblp:180/6275 · also Yong-Qiang Zhao 0001 · DBLP profile ↗
← Back
79ranked-venue papers
11as first author
46since 2021 · last 2027
0000-0002-6974-7327ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2027 Real-world underwater polarization image restoration via joint internal and external prior learning
Yuanyang Bu, Yongqiang Zhao 0001
Signal Process.3
2025 Transductive gradient injection for improved hyperspectral image denoising
Yuanyang Bu, Yongqiang Zhao 0001, Jize Xue, Seong G. Kong, Jiaxin Yao, Jonathan Cheung-Wai Chan, Pan Liu 0001
Eng. Appl. Artif. Intell.2
2025 Progressive self-supervised framework for anomaly detection in hyperspectral images
Pan Liu 0001, Yuanyang Bu, Yongqiang Zhao 0001, Seong G. Kong
Eng. Appl. Artif. Intell.3
2025 Color-aware fusion of nighttime infrared and visible images
abstract
Pixel-level fusion of visible and infrared images has demonstrated promise in enhancing information representation. However, nighttime image fusion remains challenging due to low and uneven lighting. Existing fusion methods neglect the preservation of color-related information at night, resulting in unsatisfactory outcomes with insufficient brightness. This paper presents a novel color image fusion framework to prevent color distortion, thus generating results more aligned with human perception. Firstly, we design an image fusion network to retain color information from visible images under low-light conditions. Secondly, we incorporate mature low-light enhancement technology into the network as a flexible component to produce fusion results under normal illumination. The training process is carefully designed to address potential issues of overexposure or noise amplification. Finally, we utilize knowledge distillation to create a lightweight end-to-end network that directly generates fusion results under normal lighting conditions from pairs of low-light images. Experimental results demonstrate that our proposed framework outperforms existing methods in nighttime scenarios.
Jiaxin Yao, Yongqiang Zhao 0001, Yuanyang Bu, Seong G. Kong
Eng. Appl. Artif. Intell.2
2025 Enhancing Visual Data Completion With Pseudo Side Information Regularization
abstract
Unsupervised image restoration methods relying on a single data source often face challenges in achieving high-quality visual data completion due to the absence of additional supplementary information. This paper presents a novel optimization framework to address this limitation and further enhance the performance of image restoration. The framework generates pseudo side information (PSI) and utilizes it to guide the process of visual data completion. We introduce a pseudo side information regularizer (PSIR) tailored specifically for visual data completion tasks. The PSIR comprises two components: the PSI generator and updater, responsible for generating and refining the PSI, and the neural self-expressive prior (NSEP), which identifies a prior matching the desired result and PSI during optimization. Notably, our method achieves comprehensive visual data completion across various data types without the need for additional reference side information or training data. Extensive experimental evaluations conducted on spectral data (including color images, multispectral images, and hyperspectral images), video data (including gray video, color video, and hyperspectral video), magnetic resonance image, and real cloud data demonstrate the superiority of our approach over other state-of-the-art completion methods under different missing rate scenarios.
Pan Liu 0001, Yuanyang Bu, Yongqiang Zhao 0001, Seong G. Kong
IEEE Trans. Circuits Syst. Video Technol.3
2025 Collaborative Multimodal Fusion Network for Multiagent Perception
abstract
With the increasing popularity of autonomous driving systems and their applications in complex transportation scenarios, collaborative perception among multiple intelligent agents has become an important research direction. Existing single-agent multimodal fusion approaches are limited by their inability to leverage additional sensory data from nearby agents. In this article, we present the collaborative multimodal fusion network (CMMFNet) for distributed perception in multiagent systems. CMMFNet first extracts modality-specific features from LiDAR point clouds and camera images for each agent using dual-stream neural networks. To overcome the ambiguity in-depth prediction, we introduce a collaborative depth supervision module that projects dense fused point clouds onto image planes to generate more accurate depth ground truths. We then present modality-aware fusion strategies to aggregate homogeneous features across agents while preserving their distinctive properties. To align heterogeneous LiDAR and camera features, we introduce a modality consistency learning method. Finally, a transformer-based fusion module dynamically captures cross-modal correlations to produce a unified representation. Comprehensive evaluations on two extensive multiagent perception datasets, OPV2V and V2XSet, affirm the superiority of CMMFNet in detection performance, establishing a new benchmark in the field.
Lei Zhang 0166, Binglu Wang, Yongqiang Zhao 0001, Yuan Yuan 0006, Tianfei Zhou, Zhijun Li 0001
IEEE Trans. Cybern.3
2025 Guided Feature Fusion: Zero-Shot Framework for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising algorithms often face challenges with non-uniform, band-dependent noise and fail to fully utilize high signal-to-noise ratio (SNR) spectral bands. To address these issues, this paper presents a zero-shot framework that leverages high-SNR bands as scene-specific references to guide the restoration of degraded bands. The framework features a dual-branch network architecture with the guide branch extracting high-fidelity spatial structures from high-SNR bands and the denoising branch restoring degraded bands through dynamic feature fusion. Fusion modules, enhanced with spectral-spatial attention, enable adaptive and spatially aware refinement between branches. Unlike conventional methods relying on external training datasets or fixed priors, the proposed approach adapts to the scene-specific quality disparities within each HSI. Extensive experiments on synthetic and real-world datasets demonstrate its superior ability in preserving spatial-spectral fidelity and handling complex noise scenarios, outperforming state-of-the-art methods.
Ziqin Zhang, Yuanyang Bu, Pan Liu 0001, Jiaxin Yao, Yongqiang Zhao 0001, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.6
2024 Inheriting Bayer's Legacy: Joint Remosaicing and Denoising for Quad Bayer Image Sensor
Haijin Zeng, Jiezhang Cao, Shaoguang Huang, Yongqiang Zhao 0001, Hiêp Quang Luong, Jan Aelterman, Wilfried Philips
Int. J. Comput. Vis.5
2024 Mixed norm regularized models for low-rank tensor completion
Yuanyang Bu, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
Inf. Sci.2
2024 Transferable Multiple Subspace Learning for Hyperspectral Image Super-Resolution
abstract
In real hyperspectral scenes, heterogeneous spatial details and noises make a single subspace assumptions unrealistic. In this letter, a novel transferable multiple tensor subspace learning scheme is proposed for super-resolution enhancement of hyperspectral image (HSI). The intrinsic assumption is that the nonlocal patch tensors extracted from HSIs are derived from multiple tensor low-rank subspaces, which is compatible with practical data distribution and may better characterize the complex structures underlying HSIs. The transferable subspace structures are embedded into both nonblind and semi-blind HSI super-resolution. The alternating direction method of multipliers (ADMMs) algorithm is derived for model learning. The superiority of our method is demonstrated by comprehensive experiments on both synthetic and real datasets.
Yuanyang Bu, Yongqiang Zhao 0001, Jize Xue, Jiaxin Yao, Jonathan Cheung-Wai Chan
IEEE Geosci. Remote. Sens. Lett.2
2024 Temporal Action Localization in the Deep Learning Era: A Survey
abstract
The temporal action localization research aims to discover action instances from untrimmed videos, representing a fundamental step in the field of intelligent video understanding. With the advent of deep learning, backbone networks have been instrumental in providing representative spatiotemporal features, while the end-to-end learning paradigm has enabled the development of high-quality models through data-driven training. Both supervised and weakly supervised learning approaches have contributed to the rapid progress of temporal action localization, resulting in a multitude of methods and a large body of literature, making a comprehensive survey a pressing necessity. This paper presents a thorough analysis of existing action localization works, offering a well-organized taxonomy that highlights the strengths and weaknesses of each strategy. In the realm of supervised learning, in addition to the anchor mechanism, we introduce a novel classification mechanism to categorize and summarize existing works. Similarly, for weakly supervised learning, we extend the traditional pre-classification and post-classification mechanisms by providing a fresh perspective on enhancement strategies. Furthermore, we shed light on the bottleneck of confidence estimation, a critical yet overlooked aspect of current works. By conducting detailed analyses, this survey serves as a valuable resource for researchers, providing beneficial guidance to newcomers and inspiring seasoned researchers alike.
Binglu Wang, Yongqiang Zhao 0001, Le Yang 0008, Teng Long 0001, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Navigating Uncertainty: Semantic-Powered Image Enhancement and Fusion
abstract
The fusion of infrared and visible imagery plays a crucial role in environmental monitoring. Existing approaches aim to achieve high-quality perceptual results for human observers and robust outcomes for machine-based high-level tasks by adopting a joint design of fusion and segmentation. However, design constraints imposed by predetermined fusion rules limit the precision of high-level tasks, and drawbacks in the image domains to be fused are often overlooked. This paper presents a novel semantic-powered infrared and visible image fusion framework to address these issues. The key feature of our approach is the utilization of trainable gains and weights in enhancement and fusion processes, influenced solely by segmentation and serving as uncertainty parameters. We propose a two-stage training strategy: initially, training a combined enhancement and fusion network with random uncertainty parameters, followed by the estimation of semantic-driven uncertainty parameters. The enhancement and fusion process is optimized within the Laplacian Pyramid framework to ensure efficient computation. Experimental results highlight the significance of modeling the fusion process with uncertainty for achieving satisfactory fusion and segmentation outcomes.
Jiaxin Yao, Yongqiang Zhao 0001, Seong G. Kong
IEEE Signal Process. Lett.2
2024 Physics-Driven Multispectral Filter Array Pattern Optimization and Hyperspectral Image Reconstruction
abstract
This paper presents a hyperspectral image (HSI) reconstruction technique based on physics-driven optimization of multispectral filter array (MSFA) patterns. The encoding of HSIs using an MSFA and their decoding through deep learning has gained increasing attention. However, previous studies have seldom explored pattern optimization from a physical perspective during the encoding process. In this paper, we apply a spectral sensitivity function (SSF) response model to generate the MSFA, and the goal of encoder optimization extends from SSF to physical structural parameters. To fully utilize spatial and spectral information in the decoding process, we design an end-to-end dual-branch spatial-spectral fusion network (DSFNet). By jointly optimizing the MSFA with the SSF response model and DSFNet, the proposed method significantly improves the reconstruction accuracy of HSI. When compared with existing HSI reconstruction methods, our proposed approach achieves state-of-the-art performance in both metric and visual quality.
Pan Liu 0001, Yongqiang Zhao 0001, Seong G. Kong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Tensor Convolution-Like Low-Rank Dictionary for High-Dimensional Image Representation
abstract
High-dimensional image representation is a challenging task since data has the intrinsic low-dimensional and shift-invariant characteristics. Currently, popular methods, such as tensor-Singular Value Decomposition (t-SVD), have limited ability in expressing shift-invariant subspace knowledge underlying data. To these problem, we propose a high-dimensional image representation framework based on Tensor Convolution-like Low-Rank Dictionary (TCLRD), which considers the shift-invariant low-dimensional structure of a tensor-valued data by convolution-like low-rank dictionary learning and coefficient coding, to promote the high-dimensional image representation ability. To be specific, we first define the TCLRD framework with low-rank constraint for dictionary and coefficient, in which tensor factorization and tensor-tensor product over frequency domain can be understood as convolution-like operation when describing shift-invariant. Then, the tensor Schatten-p norm is introduced to verify that TCLRD has rational mathematical interpretation. We study the TCLRD minimization problem in tensor completion with the ADMM-based optimization algorithm. The efficient solving scheme with TCLRD is extendable to various low-rank models like tensor robust principal component analysis and subspace clustering, and prove their theoretical guarantees based on generalization error. Extensive experimental results demonstrate the proposed TCLRD methods are beyond state-of-the-arts in typical tasks, including image denoising, HSI completion and image clustering.
Jize Xue, Yongqiang Zhao 0001, Tongle Wu, Jonathan Cheung-Wai Chan
IEEE Trans. Circuits Syst. Video Technol.2
2024 U²PNet: An Unsupervised Underwater Image-Restoration Network Using Polarization
abstract
This article presents U 2PNet, a novel unsupervised underwater image restoration network using polarization for improving signal-to-noise ratio and image quality in underwater imaging environments. Traditional methods for underwater image restoration using polarization require specific cues or pairs of underwater polarization datasets, which limit their practical applications. Our proposed method requires only one mosaicked polarized image of the scene and does not require datasets for pretraining or specific cues. We design two subnetworks (T-net and B textsubscript ∞ -net) to accurately estimate the transmission map and background light, and unique nonreference loss functions to ensure effective restoration. Our experiments are based on an indoor polarization simulated dataset and a real polarization image dataset constructed from our underwater robotic platform equipped with polarization cameras. Experiment results demonstrate that our proposed method achieves state-of-the-art performance on both simulated and real underwater polarization images. The code and datasets will be available at https://github.com/polwork/U-2Pnet.
Linghao Shen, Haisheng Xia, Yongqiang Zhao 0001, Ning Li 0038, Seong G. Kong, Binglu Wang, Zhijun Li 0001
IEEE Trans. Cybern.4
2024 Scale-Aware Backprojection Transformer for Single Remote Sensing Image Super-Resolution
abstract
Backprojection networks have achieved promising super-resolution performance for nature images but not well be explored in the remote sensing image super-resolution (RSISR) field due to the high computation costs. In this article, we propose a scale-aware backprojection Transformer termed SPT for RSISR. SPT incorporates the backprojection learning strategy into a Transformer framework. It consists of scale-aware backprojection-based self-attention layers (SPALs) for scale-aware low-resolution feature learning and scale-aware backprojection-based Transformer blocks (SPTBs) for hierarchical feature learning. A backprojection-based reconstruction module (PRM) is also introduced to enhance the hierarchical features for image reconstruction. SPT stands out by efficiently learning low-resolution features without excessive modules for high-resolution processing, resulting in lower computational resources. Experimental results on UCMerced and AID datasets demonstrate that SPT obtains state-of-the-art results compared to other leading RSISR methods.
Jinglei Hao, Wukai Li, Yongqiang Zhao 0001, Shunzhou Wang, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.5
2024 Polarization-Driven Solution for Mitigating Scattering and Uneven Illumination in Underwater Imagery
abstract
This paper introduces a Polarization-Driven Solution (PDS) to enhance the contrast of underwater imagery degraded by light scattering and uneven illumination. Images taken in underwater environments suffer from reduced contrast due to the combined effects of light scattering and non-uniform illumination. We present an Underwater Joint Degradation Model (UJDM) that effectively describes the compounded impacts of scattering and uneven illumination. By exploiting the polarization distinctions between objects and scattered light, we mitigate the deleterious effects of light scattering. Additionally, we leverage the polarization information to persist across changes in illumination, enhancing contrast and detail, especially in darker image regions. Building upon this foundation, our proposed underwater image restoration technique synergistically combines de-scattering with low-light image enhancement. We establish a benchmark database comprising both simulated and authentic underwater polarization images. Experiment results demonstrate that our proposed technique outperforms state-of-the-art underwater image contrast enhancement algorithms, validated through both subjective assessments and objective evaluation metrics. The source code and the datasets are available at https://github.com/polwork/PDS.
Linghao Shen, Mohamed Reda, Yongqiang Zhao 0001, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.4
2024 Rapid Hyperspectral Anomaly Detection Using Discriminative Band Selection
abstract
Hyperspectral image (HSI) exhibits high-quality spectral signals that convey subtle differences, enabling the discrimination of similar materials and providing a unique advantage for anomaly detection (AD). Fine spectral of anomalies can be effectively identified amidst heterogeneous background pixels. Given the similarity of materials in spatial and spectral dimensions, joint utilization of spatial and spectral information enhances detection performance. However, many existing AD approaches for HSIs usually achieve high accuracy at the expense of high computational complexity. In response to the requirements of practical detection scenarios-efficiency, robustness, and accuracy-this article introduces a rapid and robust AD algorithm through discriminative band selection for HSIs. We propose a spatial-spectral feature extraction strategy to ensure detection accuracy. Initially, to effectively mine context information across a broad spectral range, the HSI cube in space is partitioned into several groups using a coarse-to-fine strategy. Subsequently, we identify the most relevant and informative bands based on spatial local density and spectral information entropy, forming the coarse HSI bands subset. Following this, we design a multiband target-background ratio (MBTBR) to capture strongly discriminative bands, resulting in the fine HSI bands subset. Finally, we present an adaptively spatial—spectral feature extraction strategy to detect anomalous targets. Extensive experimental results on real hyperspectral datasets demonstrate that the proposed method achieves satisfactory performance compared to the state-of-the-art algorithms, validating its strong robustness and low computational complexity simultaneously.
Hao-Fang Yan, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.2
2024 Few-Sample Anomaly Detection in Industrial Images With Edge Enhancement and Cascade Residual Feature Refinement
abstract
In industrial inspection scenarios, the scarcity of data and the varying appearances of anomalies pose significant challenges for existing methods in accurately localizing anomaly edges and reducing detection errors. To address these issues, we propose a few-sample anomaly detection method based on edge enhancement and cascade optimization of residual features. Our approach includes a distribution transformation-based augmentation method to generate a variety of augmented images that closely resemble the distribution of real anomaly images. We introduce an anomaly detection method combined with an edge-guided feature enhancement module and a complementary feature attention module, which accentuates features at the edges of anomalies and emphasizes anomalous regions in a cascaded structure to achieve refined anomaly localization. Extensive experiments on three widely used datasets demonstrate that the proposed approach outperforms state-of-the-art methods in detection and localization accuracy.
Naifu Yao, Yongqiang Zhao 0001, Seong G. Kong
IEEE Trans. Ind. Informatics2
2024 Unsupervised Spectral Demosaicing With Lightweight Spectral Attention Networks
abstract
This paper presents a deep learning-based spectral demosaicing technique trained in an unsupervised manner. Many existing deep learning-based techniques relying on supervised learning with synthetic images, often underperform on real-world images, especially as the number of spectral bands increases. This paper presents a comprehensive unsupervised spectral demosaicing (USD) framework based on the characteristics of spectral mosaic images. This framework encompasses a training method, model structure, transformation strategy, and a well-fitted model selection strategy. To enable the network to dynamically model spectral correlation while maintaining a compact parameter space, we reduce the complexity and parameters of the spectral attention module. This is achieved by dividing the spectral attention tensor into spectral attention matrices in the spatial dimension and spectral attention vector in the channel dimension. This paper also presents Mosaic 25 , a real 25-band hyperspectral mosaic image dataset featuring various objects, illuminations, and materials for benchmarking purposes. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed method outperforms conventional unsupervised methods in terms of spatial distortion suppression, spectral fidelity, robustness, and computational cost. Our code and dataset are publicly available at https://github.com/polwork/Unsupervised-Spectral-Demosaicing.
Haijin Zeng, Yongqiang Zhao 0001, Seong G. Kong, Yuanyang Bu
IEEE Trans. Image Process.3
2024 BEVRefiner: Improving 3D Object Detection in Bird's-Eye-View via Dual Refinement
abstract
Many multi-view camera-based 3D object detection models transform the image features into Bird’s-Eye-View (BEV) via the Lift-Splat-Shoot (LSS) mechanism, which “lifts” 2D camera-view features to the 3D voxel space based on the predicted depth distribution and then “splats” 3D features into a BEV plane for subsequent 3D object detection. However, the BEV feature in such a one-stage view transformation scheme heavily relies on the quality of the predicted depth distribution and 2D camera-view features, which further determines the final detection performance. In this paper, we propose a BEVRefiner model which performs dual refinement for both depth prediction and 2D camera-view features. On the one hand, we perform light-weight depth refinement in the depth distribution frustum space by incorporating 3D context and depth distribution prior. On the other hand, we reproject the BEV feature back to each camera view to enhance 2D image features. In this way, the original camera-view features can be enhanced by implicitly incorporating 3D contexts and multi-view contexts, which cannot be achieved in the original 2D camera view. We also propose to use dominant depth bins only for the reprojection to save computational burden. Finally, we generate the refined BEV feature using the refined depth distribution and camera-view features for more accurate 3D object detection. Our BEVRefiner can be plugged into LSS-based BEV detectors and we perform extensive experiments on the representative model BEVDet, which strongly verified the efficiency of our proposed approach under several settings.
Binglu Wang, Lei Zhang 0166, Nian Liu 0002, Rao Muhammad Anwer, Hisham Cholakkal, Yongqiang Zhao 0001, Zhijun Li 0001
IEEE Trans. Intell. Transp. Syst.7
2024 Coarse-to-Fine Nutrition Prediction
abstract
Healthy dietary intake has a broad influence on the quality of life, and nutrition prediction plays a great role in the auxiliary decision-making of diet. Given a food image, existing nutrition prediction methods directly regress the nutrition content. However, due to the complex variations in food images, such as differences in viewpoint and lighting conditions, directly regressing the nutrition content faces significant challenges. The complexity of the food image data results in a high-dimensional and feature-rich input space, which poses difficulties for traditional regression models to efficiently navigate and optimize. Consequently, the direct regression paradigm usually generates inaccurate nutrition predictions. To alleviate the ambiguity challenge in the prediction progress, we propose to narrow the searchable space for the model's predictions by decomposing the direct regression into two steps: first coarsely selecting the nutrition scope and then finely refining the prediction value, forming a coarse-to-fine nutrition prediction paradigm. Although the process of coarse prediction which selects a bin from a series of scope bins can be formulated as a standard classification problem, it exhibits a distinguishable characteristic, i.e. the closer to the ground truth bin, the less punishment in the training phase. However, most of the current methods have ignored this phenomenon, thus, we specially design the linearly smoothed label in the nutrition prediction task to reveal the relative distance to the ground truth bin, leading to extraordinary improvements. Furthermore, we conduct a pair-wise comparison among all bins by extending the 1D label into 2D space and propose the structure loss to guide the bin selection process effectively. Due to the narrowed decision space, the nutrition prediction problem can be effectively optimized, and the proposed method achieves promising results on three benchmarks ECUSTFD, VFD and Nutrition5K, demonstrating the efficiency of the coarse-to-fine paradigm equipped with the linear-smoothed structure loss.
Binglu Wang, Tianci Bu, Zaiyi Hu, Le Yang 0009, Yongqiang Zhao 0001, Xuelong Li 0001
IEEE Trans. Multim.5
2024 Unsupervised Deep Tensor Network for Hyperspectral-Multispectral Image Fusion
abstract
Fusing low-resolution (LR) hyperspectral images (HSIs) with high-resolution (HR) multispectral images (MSIs) is a significant technology to enhance the resolution of HSIs. Despite the encouraging results from deep learning (DL) in HSI-MSI fusion, there are still some issues. First, the HSI is a multidimensional signal, and the representability of current DL networks for multidimensional features has not been thoroughly investigated. Second, most DL HSI-MSI fusion networks need HR HSI ground truth for training, but it is often unavailable in reality. In this study, we integrate tensor theory with DL and propose an unsupervised deep tensor network (UDTN) for HSI-MSI fusion. We first propose a tensor filtering layer prototype and further build a coupled tensor filtering module. It jointly represents the LR HSI and HR MSI as several features revealing the principal components of spectral and spatial modes and a sharing code tensor describing the interaction among different modes. Specifically, the features on different modes are represented by the learnable filters of tensor filtering layers, the sharing code tensor is learned by a projection module, in which a co-attention is proposed to encode the LR HSI and HR MSI and then project them onto the sharing code tensor. The coupled tensor filtering module and projection module are jointly trained from the LR HSI and HR MSI in an unsupervised and end-to-end way. The latent HR HSI is inferred with the sharing code tensor, the features on spatial modes of HR MSIs, and the spectral mode of LR HSIs. Experiments on simulated and real remote-sensing datasets demonstrate the effectiveness of the proposed method.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Neural Networks Learn. Syst.3
2024 LCH: fast RGB-D salient object detection on CPU via lightweight convolutional network with hybrid knowledge distillation
Binglu Wang, Fan Zhang 0070, Yongqiang Zhao 0001
Vis. Comput.3
2023 Core: Cooperative Reconstruction for Multi-Agent Perception
abstract
This paper presents Core, a conceptually simple, effective and communication-efficient model for multi-agent cooperative perception. It addresses the task from a novel perspective of cooperative reconstruction, based on two key insights: 1) cooperating agents together provide a more holistic observation of the environment, and 2) the holistic observation can serve as valuable supervision to explicitly guide the model learning how to reconstruct the ideal observation based on collaboration. Core instantiates the idea with three major components: a compressor for each agent to create more compact feature representation for efficient broadcasting, a lightweight attentive collaboration component for cross-agent message aggregation, and a reconstruction module to reconstruct the observation based on aggregated feature representations. This learning-to-reconstruct idea is task-agnostic, and offers clear and reasonable supervision to inspire more effective collaboration, eventually promoting perception tasks. We validate Core on two large-scale multi-agent percetion dataset, OPV2V and V2X-Sim, in two tasks, i.e., 3D object detection and semantic segmentation. Results demonstrate that Core achieves state-of-the-art performance, and is more communication-efficient.
Binglu Wang, Lei Zhang 0006, Zhaozhong Wang, Yongqiang Zhao 0001, Tianfei Zhou
ICCV4
2023 Transformed Structured Sparsity With Smoothness for Hyperspectral Image Deblurring
abstract
Due to the influence of imaging equipment or environment, a hyperspectral image (HSI) is often unavoidably blurred in the acquisition process, which results in the spatial and spectral information loss of the HSI. The existing HSI deblurring methods can address the problem, however, they neglect the intrinsic structured sparsity and thus reduce the deblurring performance. Aiming at this issue, we propose a new HSI deblurring method based on transformed structured sparsity with smoothness (TSSS). We first use the local piecewise smoothness to obtain the spatial and spectral sparsity of an HSI in the gradient domain. Then, to capture the refined sparsity, we exploit the transform sparsity learning framework to encode the structured sparsity self-adaptively in transform space, where the sparse structures of transformed operators can be depicted by Laplacian scale mixture (LSM), i.e., the sparsity can be expressed as the product of a hidden positive scalar multiplier and a Laplacian vector. The visual and quantitative comparisons of experimental results on three HSI datasets indicate that our method outperforms state-of-the-arts.
Jinglei Hao, Jize Xue, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Geosci. Remote. Sens. Lett.3
2023 Learning pixel-adaptive weights for portrait photo retouching
Binglu Wang, Chengzhe Lu, Yongqiang Zhao 0001, Ning Li 0038, Xuelong Li 0001
Pattern Recognit.4
2023 Laplacian Pyramid Fusion Network With Hierarchical Guidance for Infrared and Visible Image Fusion
abstract
The fusion of infrared and visible images combines the information from two complementary imaging modalities for various computer vision tasks. Many existing techniques, however, fail to maintain a uniform overall style and keep salient details of individual modalities simultaneously. This paper presents an end-to-end Laplacian Pyramid Fusion Network with hierarchical guidance (HG-LPFN) that takes advantage of pixel-level saliency reservation of Laplacian Pyramid and global optimization capability of deep learning. The proposed scheme generates hierarchical saliency maps through Laplacian Pyramid decomposition and modal difference calculation. In the pyramid fusion mode, all sub-networks are connected in a bottom-up manner. The sub-network for low-frequency fusion focuses on extracting universal features to produce an opposite style while sub-networks for high-frequency fusion determine how much the details of each modality will be retained. Taking the style, details, and background into consideration, we design a set of novel loss functions to supervise both low-frequency images and full-resolution images under the guidance of saliency maps. Experimental results on public datasets demonstrate that the proposed HG-LPFN outperforms the state-of-the-art image fusion techniques.
Jiaxin Yao, Yongqiang Zhao 0001, Yuanyang Bu, Seong G. Kong, Jonathan Cheung-Wai Chan
IEEE Trans. Circuits Syst. Video Technol.2
2023 Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion-Based Transformer Network for Remote Sensing Image Super-Resolution
abstract
Remote sensing image super-resolution (RSISR) plays a vital role in enhancing spatial detials and improving the quality of satellite imagery. Recently, Transformer-based models have shown competitive performance in RSISR. To mitigate the quadratic computational complexity resulting from global self-attention, various methods constrain attention to a local window, enhancing its efficiency. Consequently, the receptive fields in a single attention layer are inadequate, leading to insufficient context modeling. Furthermore, while most transform-based approaches reuse shallow features through skip connections, relying solely on these connections treats shallow and deep features equally, impeding the model’s ability to characterize them. To address these issues, we propose a novel transformer architecture called Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion Based Transformer Network (SPIFFNet) for RSISR. Our proposed model effectively enhances context cognition and understanding of the entire image, facilitating efficient integration of features cross-stages. The model incorporates Cross-Spatial Pixel Integration Attention (CSPIA) to introduce contextual information into a local window, while Cross-Stage Feature Fusion Attention (CSFFA) adaptively fuses features from the previous stage to improve feature expression in line with the requirements of the current stage. We conducted comprehensive experiments on multiple benchmark datasets, demonstrating the superior performance of our proposed SPIFFNet in terms of both quantitative metrics and visual quality when compared to state-of-the-art methods. Our code is available at https://github.com/Dr-Lyt/SPIFFNet.
Lingtong Min, Binglu Wang, Le Zheng, Yongqiang Zhao 0001, Le Yang 0008, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Hybrid Attention-Based U-Shaped Network for Remote Sensing Image Super-Resolution
abstract
Recently, remote sensing image super-resolution (RSISR) has drawn considerable attention and made great breakthroughs based on convolutional neural networks (CNNs). Due to the scale and richness of texture and structural information frequently recurring inside the same remote sensing images (RSIs) but varying greatly with different RSIs, state-of-the-art CNN-based methods have begun to explore the multiscale global features in RSIs by using attention mechanisms. However, they are still insufficient to explore significant content attention clues in RSIs. In this article, we present a new hybrid attention-based U-shaped network (HAUNet) for RSISR to effectively explore the multiscale features and enhance the global feature representation by hybrid convolution-based attention. It contains two kinds of convolutional attention-based single-scale feature extraction modules (SEM) to explore the global spatial context information and abstract content information, and a cross-scale interaction module (CIM) as the skip connection between different scale feature outputs of encoders to bridge the semantic and resolution gaps between them. Considering the existence of equipment with poor hardware facilities, we further design a lighter HAUNet-S with about 596K parameters. Experimental attribution analysis method LAM results demonstrate that our HAUNet is a more efficient way to capture meaningful content information and quantitative results can show that our HAUNet can significantly improve the performance of RSISR on four remote sensing test datasets. Meanwhile, HAUNET-S also maintains competitive performance. Our code is available athttps://github.com/likakakaka/HAUNet_RSISR.
Binglu Wang, Yongqiang Zhao 0001, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Spectral Super-Resolution Based on Dictionary Optimization Learning via Spectral Library
abstract
Extensive works have been reported in hyperspectral images (HSIs) and multispectral images (MSIs) fusion to raise the spatial resolution of HSIs. However, the limited acquisition of HSIs has been an obstacle to such approaches. Spectral super-resolution (SSR) of MSI is a challenging and less investigated topic, which can also provide high-resolution synthetic HSIs. To deal with this high ill-posedness problem, we perform super-resolution enhancement of MSIs in the spectral domain by incorporating a spectral library as a priori. First, an aligned spectral library, which maps the open-source spectral library to a specific spectral library created for the reconstructed HR HSI, is represented. An intermediate latent HSI is obtained by fusing the spatial information from MSI and the hyperspectral information from a specific spectral library. Then, we use low-rank attribute embedding to transfer latent HSI into a robust subspace. Finally, a low-rank HSI dictionary representing the hyperspectral information is learned from the latent HSI. The adaptive sparse coefficient of MSI is obtained with a nonnegative constraint. By fusing these two terms, we get the final HR HSI. The proposed SSR model does not require any pretraining stages. We confirm the validity and superiority of our proposed SSR algorithm by comparing it with several benchmark state-of-the-art approaches on different datasets.
Hao-Fang Yan, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.2
2023 Joint Denoising-Demosaicking Network for Long-Wave Infrared Division-of-Focal-Plane Polarization Images With Mixed Noise Level Estimation
abstract
Denoising and demosaicking long-wave infrared (LWIR) division-of-focal-plane (DoFP) polarization images are crucial for various vision applications. However, existing methods rely on the sequential application of individual denoising and demosaicking processes, which may result in the accumulation of errors produced by each process. To address this issue, we propose a joint denoising and demosaicking method for LWIR DoFP images based on a three-stage progressive deep convolutional neural network. To ensure the generalization ability of this network, it is essential to have adequate training data that closely resembles real data. Therefore, we model the complex noise sources that affect LWIR DoFP images as mixed Poisson-Additive-Stripe noise and construct a least-squares problem based on the polarization measurement redundancy error to estimate the parameters of this model on real images. Subsequently, the estimated noise parameters are used to generate training data that enables the network to learn accurate polarization image statistics and improve its generalization ability. The experimental results demonstrate the effectiveness of the proposed method in enhancing the image restoration performance on real LWIR DoFP polarization data.
Ning Li 0038, Binglu Wang, François Goudail, Yongqiang Zhao 0001, Quan Pan 0001
IEEE Trans. Image Process.4
2023 Prototype-Based Intent Perception
abstract
Intent perception is a novel task that aims to understand the intention of images, regular classification methods usually perform unsatisfactorily on intent perception due to the semantic ambiguity problem,i.e. the intra-class variety problem in which images of the same intent class may contain objects of different semantic categories and the inter-class confusion problem in which images of different intent classes may contain objects of similar semantic categories. To address this problem, this paper introduces prototype learning into the intent perception and proposes a unified framework named PIP-Net to reduce the influence of semantic ambiguity. Specifically, for each intent class, we first filter semantic ambiguity samples which are far away from the cluster center. Then we use features of the filtered samples to generate prototypes via clustering algorithm. Besides, we enhance the diversity between prototypes of different classes to better handle the inter-class confusion problem. To update the prototypes in the training process, we introduce a global matching algorithm to holistically match each feature with class prototypes, and use the momentum update strategy to stably update prototypes. Experimental results on the Intentonomy dataset demonstrate that our method can consistently outperform the traditional classification paradigm in multiple baseline models, and verify the effectiveness of our proposed prototype learning paradigm in addressing the intent perception problem. Our proposed PIP-Net achieves a new state-of-the-art performance on Intentonomy, including Macro F1 score of 31.57% and averaging F1 score of 41.85%. Code is available athttps://github.com/CodeMonsterPHD/PIP-Net.
Binglu Wang, Yongqiang Zhao 0001, Teng Long 0001, Xuelong Li 0001
IEEE Trans. Multim.3
2022 Exploring Sub-Action Granularity for Weakly Supervised Temporal Action Localization
abstract
Modeling cross-video relationship is an important issue for the weakly supervised temporal action localization task. To this end, traditional methods operate at the action level and rely on complicated strategies to prepare triplet samples, which only mines the cross-video relationships among three videos from two categories. In this work, we observe that action instances from different categories could exhibit similar motion patterns, i.e. subaction, and propose to operate at the sub-action granularity to elaborately explore cross-video relationships. However, only given video-level category labels, the sub-actions are undefined and not annotated. To tackle this challenge, we represent video features via a group of sub-actions,i.e.the sub-action family. Specifically, the sub-action family contains multiple feature vectors, where each vector is in charge of representing a specific sub-action. The sub-action family is shared among all videos in the dataset, while all videos contribute to the learning of the sub-action family. Consequently, we can not only get rid of the complicated sampling strategy but also thoroughly mine cross-video relationships from all available videos in the dataset. To learn feature vectors within the sub-action family, we employ a bottom-up temporal action localization paradigm and introduce an extra top-down branch. The sub-action family is introduced into the top-down branch, and it learns feature vectors via representing raw video features. Moreover, we propose a consistency loss to guide the learning process and a diversity loss to mine distinct sub-actions. Extensive experiments are carried out on three benchmark datasets,i.e.THUMOS14, ActivityNet v1.2 and ActivityNet v1.3, and the proposed method builds new high performance.
Binglu Wang, Yongqiang Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 When Laplacian Scale Mixture Meets Three-Layer Transform: A Parametric Tensor Sparsity for Tensor Completion
abstract
Recently, tensor sparsity modeling has achieved great success in the tensor completion (TC) problem. In real applications, the sparsity of a tensor can be rationally measured by low-rank tensor decomposition. However, existing methods either suffer from limited modeling power in estimating accurate rank or have difficulty in depicting hierarchical structure underlying such data ensembles. To address these issues, we propose a parametric tensor sparsity measure model, which encodes the sparsity for a general tensor by Laplacian scale mixture (LSM) modeling based on three-layer transform (TLT) for factor subspace prior with Tucker decomposition. Specifically, the sparsity of a tensor is first transformed into factor subspace, and then factor sparsity in the gradient domain is used to express the local similarity in within-mode. To further refine the sparsity, we adopt LSM by the transform learning scheme to self-adaptively depict deeper layer structured sparsity, in which the transformed sparse matrices in the sense of a statistical model can be modeled as the product of a Laplacian vector and a hidden positive scalar multiplier. We call the method as parametric tensor sparsity delivered by LSM-TLT. By a progressive transformation operator, we formulate the LSM-TLT model and use it to address the TC problem, and then the alternating direction method of multipliers-based optimization algorithm is designed to solve the problem. The experimental results on RGB images, hyperspectral images (HSIs), and videos demonstrate the proposed method outperforms state of the arts.
Jize Xue, Yongqiang Zhao 0001, Yuanyang Bu, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Cybern.2
2022 Multiple Instance Graph Learning for Weakly Supervised Remote Sensing Object Detection
abstract
Weakly supervised object detection (WSOD) has recently attracted much attention in the field of remote sensing, where only image-level labels that distinguish the existence of an object in images are required. However, existing methods frequently treat the most discriminative area of an object as the optimal solution and, meanwhile, ignore the fact that more than one instance may exist in a certain class in remote sensing images (RSIs). To address the issue, we propose a unique multiple instance graph (MIG) learning framework for WSOD in RSIs. The motivation of this work is twofold: 1) a spatial graph-based vote (SGV) mechanism is proposed to find high-quality objects by collecting the top-ranking votes with highly spatial overlap and 2) an appearance graph-based instance mining (AGIM) model is further constructed to exploit all possible instances with the same class by propagating the label information according to the apparent similarity. It is noted that the formulated MIG framework that collaborates SGV and AGIM is independent of extra hyperparameters or annotations. Experimental results reported for two well-known benchmarks, i.e., NWPU VHR-10.v2 and DIOR, testify to the superiority of the proposed framework by 55.9% and 25.11% mAPs.
Binglu Wang, Yongqiang Zhao 0001, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Variational Regularization Network With Attentive Deep Prior for Hyperspectral-Multispectral Image Fusion
abstract
Hyperspectral–multispectral image (HSI-MSI) fusion relies on a robust degradation model and data prior, where the former describes the degeneration of HSI in the spectral and spatial domains, and the latter reveals the latent statistics of the expected high-resolution (HR) HSI. In practice, the degradation model is often unknown, and the data prior is usually too complicated to be expressed analytically. In this study, we propose a variational network for HSI-MSI fusion (VaFuNet), in which the degradation model and data prior are implicitly represented by a deep learning network and jointly learned from the training data. A variational fusion model regularized by deep prior is first proposed, and then, it is optimized via a half-quadratic splitting and unfolded into a deep network. The deep prior is implicitly represented by a proximity operator. Due to the structural self-similarity, HSI possesses structural recurrences across different scales. To exploit such nonlocal prior and enhance the representability of network, we also propose a multiscale nonlocal attention and embed it into the deep prior proximity. The degradation model and deep prior proximity are jointly learned via end-to-end training. Experimental results on simulated and real-life HSI datasets demonstrate the effectiveness of the proposed VaFuNet HSI-MSI fusion method.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.3
2022 Multilayer Sparsity-Based Tensor Decomposition for Low-Rank Tensor Completion
abstract
Existing methods for tensor completion (TC) have limited ability for characterizing low-rank (LR) structures. To depict the complex hierarchical knowledge with implicit sparsity attributes hidden in a tensor, we propose a new multilayer sparsity-based tensor decomposition (MLSTD) for the low-rank tensor completion (LRTC). The method encodes the structured sparsity of a tensor by the multiple-layer representation. Specifically, we use the CANDECOMP/PARAFAC (CP) model to decompose a tensor into an ensemble of the sum of rank-1 tensors, and the number of rank-1 components is easily interpreted as the first-layer sparsity measure. Presumably, the factor matrices are smooth since local piecewise property exists in within-mode correlation. In subspace, the local smoothness can be regarded as the second-layer sparsity. To describe the refined structures of factor/subspace sparsity, we introduce a new sparsity insight of subspace smoothness: a self-adaptive low-rank matrix factorization (LRMF) scheme, called the third-layer sparsity. By the progressive description of the sparsity structure, we formulate an MLSTD model and embed it into the LRTC problem. Then, an effective alternating direction method of multipliers (ADMM) algorithm is designed for the MLSTD minimization problem. Various experiments in RGB images, hyperspectral images (HSIs), and videos substantiate that the proposed LRTC methods are superior to state-of-the-art methods.
Jize Xue, Yongqiang Zhao 0001, Shaoguang Huang, Wenzi Liao, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Neural Networks Learn. Syst.2
2021 Smooth Coupled Tucker Decomposition for Hyperspectral Image Super-Resolution
Yuanyang Bu, Yongqiang Zhao 0001, Jize Xue, Jonathan Cheung-Wai Chan
PRCV (3)2
2021 PFWNet: Pretraining neural network via feature jigsaw puzzle for weakly-supervised temporal action localization
Binglu Wang, Yongqiang Zhao 0001
Neurocomputing2
2021 I2Net: Mining intra-video and inter-video attention for temporal action localization
abstract
This paper focuses on two challenges for temporal action localization community, i.e., lack of long-term relationship and action pattern uncertainty. The former prevents the cooperation among multiple action instances within a video, while the latter may cause incomplete localizations or false positives. The lack of long-term relationship challenge results from the limited receptive field. Instead of stacking multiple layers or using large convolution kernels, we propose the intra-video attention mechanism to bring global receptive field to each temporal point. As for the action pattern uncertainty challenge, although it is hard to precisely depict the desired action pattern, paired videos that share the same action category can provide complementary information about action pattern. Consequently, we propose an inter-video attention mechanism to assist learning accurate action patterns. Based on the intra-video attention and inter-video attention, we propose a unified framework, namely I2Net, to tackle the challenging temporal action localization task. Given two videos containing sharing action categories, I2Net adopts the widely used one-stage action localization paradigm to dispose of them in parallel. As for two neighboring layers within the same video, the intra-video attention brings global information to each temporal point and helps to learn representative features. As for two parallel layers between two videos, the inter-video attention introduces complementary information to each video and helps to learn accurate action patterns. With the cooperation of intra-video and inter-video attention mechanisms, I2Net shows obvious performance gains over the baseline and builds new state-of-the-art on two widely-used benchmarks, i.e., THUMOS14 and ActivityNet v1.3.
Binglu Wang, Songhui Ma, Yongqiang Zhao 0001
Neurocomputing5
2021 Hybrid Local and Nonlocal 3-D Attentive CNN for Hyperspectral Image Super-Resolution
abstract
A deep convolutional neural network (CNN) has shown its great potential in hyperspectral image (HSI) super-resolution (SR). Integrating CNN with attention mechanism is expected to boost the SR performance. However, how to learn attention along the spectral, spatial, and channel dimensions of HSI is still an open issue, and the current attention mechanism is not efficient in capturing long-range interdependency in HSI. In this letter, we first design a local 3-D attention module to learn the spectral-spatial-channel attention by exploiting local contextual information in HSI. Then, we propose a nonlocal 3-D attention module, in which the long-range interdependency in HSI can be exploited for attention learning. By jointly embedding the local and nonlocal attention in a residual 3-D CNN, a hybrid local and nonlocal 3-D attentive CNN can be built for HSI SR. The experimental results show that local and nonlocal attention formulation leads to competitive SR performance.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Geosci. Remote. Sens. Lett.3
2021 POLO: Learning Explicit Cross-Modality Fusion for Temporal Action Localization
abstract
Temporal action localization aims at discovering action instances in untrimmed videos, where RGB and flow are two widely used feature modalities. Specifically, RGB chiefly reveals appearance and flow mainly depicts motion. Given RGB and flow features, previous methods employ the early fusion or late fusion paradigm to mine the complementarity between them. By concatenating raw RGB and flow features, the early fusion implicitly achieved complementarity by the network, but it partly discards the particularity of each modality. The late fusion independently maintains two branches to explore the particularity of each modality, but it only fuses the localization results, which is insufficient to mine the complementarity. In this work, we propose explicit cross-modality fusion (POLO) to effectively utilize the complementarity between two modalities and thoroughly explore the particularity of each modality. POLO performs cross-modality fusion via estimating the attention weight from RGB modality and employing it to flow modality (vice versa). This bridges the complementarity of one modality to supply the other. Assisted with the attention weight, POLO independently learns from RGB and flow features and explores the particularity of each modality. Extensive experiments on two benchmarks demonstrate the preferable performance of POLO.
Binglu Wang, Le Yang 0008, Yongqiang Zhao 0001
IEEE Signal Process. Lett.3
2021 Hyperspectral and Multispectral Image Fusion via Graph Laplacian-Guided Coupled Tensor Decomposition
abstract
We propose a novel graph Laplacian-guided coupled tensor decomposition (gLGCTD) model for fusion of hyperspectral image (HSI) and multispectral image (MSI) for spatial and spectral resolution enhancements. The coupled Tucker decomposition is employed to capture the global interdependencies across the different modes to fully exploit the intrinsic global spatial-spectral information. To preserve local characteristics, the complementary submanifold structures embedded in high-resolution (HR)-HSI are encoded by the graph Laplacian regularizations. The global spatial-spectral information captured by the coupled Tucker decomposition and the local submanifold structures are incorporated into a unified framework. The gLGCTD fusion framework is solved by a hybrid framework between the proximal alternating optimization (PAO) and the alternating direction method of multipliers (ADMM). Experimental results on both synthetic and real data sets demonstrate that the gLGCTD fusion method is superior to state-of-the-art fusion methods with a more accurate reconstruction of the HR-HSI.
Yuanyang Bu, Yongqiang Zhao 0001, Jize Xue, Jonathan Cheung-Wai Chan, Seong G. Kong, Jinhuan Wen, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.2
2021 No-Reference Physics-Based Quality Assessment of Polarization Images and Its Application to Demosaicking
abstract
Assessing the quality of polarization images is of significance for recovering reliable polarization information. Widely used quality assessment methods including peak signal-to-noise ratio and structural similarity index require reference data that is usually not available in practice. We introduce a simple and effective physics-based quality assessment method for polarization images that does not require any reference. This metric, based on the self-consistency of redundant linear polarization measurements, can thus be used to evaluate the quality of polarization images degraded by noise, misalignment, or demosaicking errors even in the absence of ground-truth. Based on this new metric, we propose a novel processing algorithm that significantly improves demosaicking of division-of-focal-plane polarization images by enabling efficient fusion between demosaicking algorithms and edge-preserving image filtering. Experimental results obtained on public databases and homemade polarization images show the effectiveness of the proposed method.
Ning Li 0038, Benjamin Le Teurnier, Matthieu Boffety, François Goudail, Yongqiang Zhao 0001, Quan Pan 0001
IEEE Trans. Image Process.5
2021 Spatial-Spectral Structured Sparse Low-Rank Representation for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution by fusing high-resolution multispectral image (HR-MSI) and low-resolution hyperspectral image (LR-HSI) aims at reconstructing high resolution spatial-spectral information of the scene. Existing methods mostly based on spectral unmixing and sparse representation are often developed from a low-level vision task perspective, they cannot sufficiently make use of the spatial and spectral priors available from higher-level analysis. To this issue, this paper proposes a novel HSI super-resolution method that fully considers the spatial/spectral subspace low-rank relationships between available HR-MSI/LR-HSI and latent HSI. Specifically, it relies on a new subspace clustering method named "structured sparse low-rank representation" (SSLRR), to represent the data samples as linear combinations of the bases in a given dictionary, where the sparse structure is induced by low-rank factorization for the affinity matrix. Then we exploit the proposed SSLRR model to learn the SSLRR along spatial/spectral domain from the MSI/HSI inputs. By using the learned spatial and spectral low-rank structures, we formulate the proposed HSI super-resolution model as a variational optimization problem, which can be readily solved by the ADMM algorithm. Compared with state-of-the-art hyperspectral super-resolution methods, the proposed method shows better performance on three benchmark datasets in terms of both visual and quantitative evaluation.
Jize Xue, Yongqiang Zhao 0001, Yuanyang Bu, Wenzi Liao, Jonathan Cheung-Wai Chan, Wilfried Philips
IEEE Trans. Image Process.2
2020 Full-Time Monocular Road Detection Using Zero-Distribution Prior of Angle of Polarization
Ning Li 0038, Yongqiang Zhao 0001, Quan Pan 0001, Seong G. Kong, Jonathan Cheung-Wai Chan
ECCV (25)2
2020 Hyperspectral Image Super-Resolution via Self-projected Smooth Prior
Yuanyang Bu, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
PRCV (1)2
2020 Enhanced Sparsity Prior Model for Low-Rank Tensor Completion
abstract
Conventional tensor completion (TC) methods generally assume that the sparsity of tensor-valued data lies in the global subspace. The so-called global sparsity prior is measured by the tensor nuclear norm. Such assumption is not reliable in recovering low-rank (LR) tensor data, especially when considerable elements of data are missing. To mitigate this weakness, this article presents an enhanced sparsity prior model for LRTC using both local and global sparsity information in a latent LR tensor. In specific, we adopt a doubly weighted strategy for nuclear norm along each mode to characterize global sparsity prior of tensor. Different from traditional tensor-based local sparsity description, the proposed factor gradient sparsity prior in the Tucker decomposition model describes the underlying subspace local smoothness in real-world tensor objects, which simultaneously characterizes local piecewise structure over all dimensions. Moreover, there is no need to minimize the rank of a tensor for the proposed local sparsity prior. Extensive experiments on synthetic data, real-world hyperspectral images, and face modeling data demonstrate that the proposed model outperforms state-of-the-art techniques in terms of prediction capability and efficiency.
Jize Xue, Yongqiang Zhao 0001, Wenzi Liao, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Neural Networks Learn. Syst.2
2019 Hyperspectral Image Super-Resolution Based on Multi-Scale Wavelet 3D Convolutional Neural Network
abstract
Super-resolution (SR) of hyperspectral image (HSI) is of significance for its applications. Wavelet decomposition can be used to capture textures and structures in the HSI. In this study, we propose a multi-scale wavelet 3D convolutional neural network (MW-3D-CNN) for HSI SR. Instead of reconstructing the high resolution (HR) HSI directly, we predict the wavelet coefficients of HR HSI with the proposed network, which is composed of an embedding subnet and a predicting subnet. Both of them are built with 3D convolutional layers. The embedding subnet extracts deep spatial-spectral features from the low resolution (LR) HSI and represents the LR HSI as a set of feature cubes. The feature cubes are then fed to the predicting subnet. There are multiple output branches in the predicting subnet, each of which corresponds to a wavelet sub-band and predicts the wavelet coefficients of HR HSI. By applying inverse wavelet transform to the predicted wavelet coefficients, the HR HSI can be obtained. In the training stage, we propose to train MW-3D-CNN with L1 norm loss, which is more suitable than the conventional L2 norm loss for penalizing the errors in different wavelet sub-bands. In the experiment, the performance is tested on several HSI datasets.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS2
2019 Spectral Super-Resolution for Multispectral Image Based on Spectral and Spatial Strategies
abstract
A spectral super-resolution method is proposed in this paper to recover a high spectral resolution hyperspectral (HS) image from multispectral (MS) images. The proposed method involves spectral improvement strategy and spatial preservation strategy. For spectral improvement strategy, auxiliary MS/HS image pairs of different landscapes are exploited to estimate spectral response relationship so that a HS image is obtained as an intermediate result. Then, spectral dictionary learning is exploited to recover more accurate spectral reconstruction result. Spatial preservation strategy is used as spatial constraint to ensure spatial consistency. Additionally, low rank property of HS image is also introduced to make use of global spectral coherence among HS bands. Experiments are conducted on real MS/HS data (ALI and Hyperion) captured by EO-1 satellite. Experiment results demonstrate the superiority of our proposed method to other state-of-art methods.
Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS2
2019 Image Enhancement of Shadow Region Based on Polarization Imaging
Mohamed Reda, Linghao Shen, Yongqiang Zhao 0001
PRCV (2)3
2019 Hyper-Laplacian regularized nonlocal low-rank matrix recovery for hyperspectral image compressive sensing reconstruction
Jize Xue, Yongqiang Zhao 0001, Wenzi Liao, Jonathan Cheung-Wai Chan
Inf. Sci.2
2019 Nonconvex tensor rank minimization and its applications to tensor recovery
Jize Xue, Yongqiang Zhao 0001, Wenzi Liao, Jonathan Cheung-Wai Chan
Inf. Sci.2
2019 Nonlocal Low-Rank Regularized Tensor Decomposition for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) enjoys great advantages over more traditional image types for various applications due to the extra knowledge available. For the nonideal optical and electronic devices, HSI is always corrupted by various noises, such as Gaussian noise, deadlines, and stripings. The global correlation across spectrum (GCS) and nonlocal self-similarity (NSS) over space are two important characteristics for HSI. In this paper, a nonlocal low-rank regularized CANDECOMP/PARAFAC (CP) tensor decomposition (NLR-CPTD) is proposed to fully utilize these two intrinsic priors. To make the rank estimation more accurate, a new manner of rank determination for the NLR-CPTD model is proposed. The intrinsic GCS and NSS priors can be efficiently explored under the low-rank regularized CPTD to avoid tensor rank estimation bias for denoising performance. Then, the proposed HSI denoising model is performed on tensors formed by nonlocal similar patches within an HSI. The alternating direction method of multipliers-based optimization technique is designed to solve the minimum problem. Compared with state-of-the-art methods, the proposed algorithm can greatly promote the denoising performance of an HSI in various quality assessments.
Jize Xue, Yongqiang Zhao 0001, Wenzi Liao, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.2
2019 Spectral Super-Resolution for Multispectral Image Based on Spectral Improvement Strategy and Spatial Preservation Strategy
abstract
While hyperspectral (HS) images play a significant role in many applications, they often suffer from issues such as low spatial resolution, low temporal resolution, and some of the acquired spectral bands are either with low signal-to-noise ratio (SNR) or invalid because of the very high-noise level. To address this issue, a spectral super-resolution method is proposed in this paper to recover a high-spectral-resolution HS image from multispectral (MS) images. The reconstructed HS image will have the same spatial resolution and coverage as the input MS image. The proposed method involves spectral improvement strategy and spatial preservation strategy. For spectral improvement strategy, auxiliary MS/HS image pairs of different landscapes are exploited to estimate spectral response relationship so that an HS image is obtained as an intermediate result. Then, spectral dictionary learning is exploited to recover a more accurate spectral reconstruction result. Spatial preservation strategy is used as a spatial constraint to ensure spatial consistency. In addition, the low-rank property of HS image is also introduced to make the use of global spectral coherence among HS bands. Experiments are conducted on both simulated and real datasets including spectral enhancement of RGB image and the MS image generated by AVIRIS data and real MS/HS data (ALI and Hyperion) captured by Earth Observing-1 (EO-1) satellite. Experiment results demonstrate the superiority of our proposed method to other state-of-the-art methods.
Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.2
2019 An Iterative Image Dehazing Method With Polarization
abstract
This paper presents a joint dehazing and denoising scheme for an image taken in hazy conditions. Conventional image dehazing methods may amplify the noise depending on the distance and density of the haze. To suppress the noise and improve the dehazing performance, an imaging model is modified by adding the process of amplifying the noise in hazy conditions. This model offers depth-chromaticity compensation regularization for the transmission map and chromaticity-depth compensation regularization for dehazing the image. The proposed iterative image dehazing method with polarization uses these two joint regularization schemes and the relationship between the transmission map and dehazed image. The transmission map and irradiance image are used to promote each other. To verify the effectiveness of the algorithm, polarizing images of different scenes in different days are collected. Different algorithms are applied to the original images. Experimental results demonstrate that the proposed scheme increases visibility in extreme weather conditions without amplifying the noise.
Linghao Shen, Yongqiang Zhao 0001, Qunnie Peng, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Multim.2
2018 Hyperspectral Imagery Denoising Using Multi-Linear Weighted Nuclear Norm Minimization
abstract
Classical matrix-based denoising methods for hyperspectral imagery (HSI) may cause spatial and spectral distortion. To improve denoising performance, a multi-linear weighted nuclear norm minimization was proposed for HSI denoising. By considering spectral continuity and inter-dependency of three unfolding modes, a multi-linear rank was proposed to model the spatial and spectral nonlocal similarity. To make the proposed method more tractable, a variable splitting based technique was used to solve the optimization problem. Experiment results reveal that the proposed method outperforms state-of-the-art methods both visually and quantitatively.
Xiangyang Kong, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS2
2018 Banana Disease Detection by Fusion of Close Range Hyperspectral Image and High-Resolution Rgb Image
abstract
Early detection of banana disease can limit the spread of disease, as well as reduce the treatment costs. Current methods focus on either manually interpretation or calculation of spectral indices (e.g., the normalized difference vegetation index). In this paper, we exploit the fusion of close range hyperspectral (HS) image and high-resolution (HR) visible RGB image for potential disease detection in banana leaves. Our approach applies the joint bilateral filter to transfer the textural structures of HR RGB image to low-resolution HS image and obtain an enhanced HS image. Initial experimental results on Musa acuminata (banana) leaf images demonstrate the efficiency of our fusion approach, with significant improvements over either single data source or some conventional methods.
Wenzi Liao, Daniel Ochoa 0001, Yongqiang Zhao 0001, Gladys Villegas, Wilfried Philips
IGARSS3
2018 Deformable Dictionary Learning for SAR Image Change Detection
abstract
This paper proposes a novel method based on deformable dictionary learning for detecting the regions of change between multitemporal image pairs. We build on our previous work, which constructed a pair of dictionaries. The main shortcoming of this method was its dependence on a large amount of training data. In practice, there is often a shortage of ground-truthed training images, which limits the expression capability of the resulting dictionaries. This paper overcomes this challenge by incorporating the concept of deformation, wherein each atom of a dictionary is no longer a simple image patch, but instead is a flexible image deformation function. This enables the creation of more expressive dictionaries, capable of generalizing to a far greater variety of image patterns, while using a far smaller amount of ground-truthed images for supervised dictionary training. Deformation similarity is employed for patch matching to find the best set of atoms in the difference image (DI) dictionary for reconstructing image patches for a new input DI. Each such atom can be deformed to achieve a better match, thus extending generality while reducing the number of atoms needed in the dictionary. Multiple deformed atoms are weighted and combined to best reconstruct the input DI patch. Then, the same set of deformations and weights is projected to the corresponding atoms in the CD dictionary to obtain the output change-detection map. Experiments in six realistic synthetic aperture radar data sets demonstrate the robustness and efficiency of the proposed method in comparison with five other state-of-the-art methods from the literature.
Lin Li 0016, Yongqiang Zhao 0001, Jinjun Sun, Rustam Stolkin, Quan Pan 0001, Jonathan Cheung-Wai Chan, Seong G. Kong, Zhunga Liu
IEEE Trans. Geosci. Remote. Sens.2
2018 Joint Spatial and Spectral Low-Rank Regularization for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) noise reduction is an active research topic in HSI processing due to its significance in improving the performance for object detection and classification. In this paper, we propose a joint spectral and spatial low-rank (LR) regularized method for HSI denoising, based on the assumption that the free-noise component in an observed signal can exist in latent low-dimensional structure while the noise component does not have this property. The proposed HSI denoising method not only considers the traditional LR property across the spectral domain but also leverages nonlocal LR property over the spatial domain. The main contribution of this paper is the incorporation of the low-rankness-based nonlocal similarity into sparse representation to characterize the spatial structure. Specially, the similar patches in each cluster usually contain similar sharp structure such as edges and textures; LR performed on cluster entitles to achieve a lower rank than that on the global spectral correlation. To make the proposed method more tractable and robust, we develop a variable splitting-based technique to solve the optimization problem. Experiment results on both simulated and real hyperspectral data sets demonstrate that the proposed method outperforms state-of-the-art methods with significant improvements both visually and quantitatively.
Jize Xue, Yongqiang Zhao 0001, Wenzi Liao, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.2
2018 Hyperspectral Image Super-Resolution Based on Spatial and Spectral Correlation Fusion
abstract
Super-resolution image reconstruction has been utilized to overcome the problem of spatial resolution limitation in hyperspectral (HS) imaging. To improve the spatial resolution of HS image, this paper proposes an HS-multispectral (MS) fusion method, which exploits spatial and spectral correlations and proper regularization. High spatial correlation between MS image and the desired high-resolution HS image is conserved via an over-completed dictionary, and the spectral degradation between them projected onto the space of sparsity is applied as the spectral constraint. The high spectral correlation between high-spatial- and low-spatial-resolution HS image is preserved through linear spectral unmixing. The idea of an interactive feedback proposed in our previous work is also used when dealing with spatial reconstruction and unmixing. Low-rank property is introduced in this paper to regularize the sparse coefficients of the HS patch matrix, which is utilized as the spatial constraint. Experiments on both simulated and real data sets demonstrate that the proposed fusion algorithm achieves lower spectral distortions and the super-resolution results are superior to those of other state-of-the-art methods.
Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.2
2017 Tensor non-local low-rank regularization for recovering compressed hyperspectral images
abstract
Sparsity-based methods have been widely used in hyperspectral imagery compression recovery (HSI-CR). However, most of the available HSI-CR methods work on vector space by vectorizing hyperspectral cubes in spatial and spectral domain, which will destroy spatial and spectral correlation and result in spatial and spectral information distortion in the recovery. At the same time, vectorization also make HSI's intrinsic structure sparsity cannot be utilized adequately. In this paper, a tensor non-local low-rank regularization (TNLR) approach is proposed to exploit essential structured sparsity and explore its advantages for CR of hyperspectral imagery. Specifically, a tensor nuclear norm penalty function is utilized as tensor low-rank regularization term to describe the spatial-and-spectral correlation hidden in HSI. To further improve the computational efficiency of the proposed algorithm, a fast implementation algorithm is developed by using the alternative direction multiplier method (ADMM) technique. Experimental results are shown that the proposed TNLR-CR algorithm can significantly outperform existing state-of-the-art CR techniques for hyperspectral image recovery.
Yongqiang Zhao 0001, Jize Xue, Jinglei Hao
ICIP1
2017 Learning and Transferring Deep Joint Spectral-Spatial Features for Hyperspectral Classification
abstract
Feature extraction is of significance for hyperspectral image (HSI) classification. Compared with conventional hand-crafted feature extraction, deep learning can automatically learn features with discriminative information. However, two issues exist in applying deep learning to HSIs. One issue is how to jointly extract spectral features and spatial features, and the other one is how to train the deep model when training samples are scarce. In this paper, a deep convolutional neural network with two-branch architecture is proposed to extract the joint spectral-spatial features from HSIs. The two branches of the proposed network are devoted to features from the spectral domain as well as the spatial domain. The learned spectral features and spatial features are then concatenated and fed to fully connected layers to extract the joint spectral-spatial features for classification. When the training samples are limited, we investigate the transfer learning to improve the performance. Low and mid-layers of the network are pretrained and transferred from other data sources; only top layers are trained with limited training samples extracted from the target scene. Experiments on Airborne Visible/Infrared Imaging Spectrometer and Reflective Optics System Imaging Spectrometer data demonstrate that the learned deep joint spectral-spatial features are discriminative, and competitive classification results can be achieved when compared with state-of-the-art methods. The experiments also reveal that the transferred features boost the classification performance.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.2
2017 Joint Hyperspectral Superresolution and Unmixing With Interactive Feedback
abstract
This paper presents an interactive feedback scheme of spatial resolution enhancement and spectral unmixing in hyperspectral imaging. Traditionally spatial resolution enhancement and spectral unmixing operations have been carried out separately, often in series. In such sequential processing, spatially enhanced hyperspectral images (HSIs) may introduce distortion in spectral fidelity making spectral unmixing results unreliable, or vice versa. Since both high- and low-resolution HSIs have the same endmembers, the deviation in spectral unmixing between targets and estimated high-resolution HSIs can be used as feedback to control spatial resolution enhancement. The spatial difference before and after unmixing can also be used as feedback to enhance spectral unmixing. Therefore, spectral unmixing is utilized as a constraint to spatial resolution enhancement, while spatial resolution enhancement helps improve spectral unmixing results. The performance of spatial resolution enhancement and spectral unmixing can be improved since one behaves like a prior to the other. Experimental results on both simulated and real HSI data sets demonstrate that the proposed interactive feedback scheme simultaneously achieved spatial resolution enhancement and spectral unmixing fidelity. This paper is an extended version of the previous work.
Yongqiang Zhao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.2
2016 Hyperspectral image classification using two-channel deep convolutional neural network
abstract
Performance of hyperspectral image classification depends on feature extraction. Compared with conventional hand-crafted feature extraction, deep learning can learn feature with more discriminative information. In this paper, a two-channel deep convolutional neural network (Two-CNN) is proposed to learn jointly spectral-spatial feature from hyperspectral image. The proposed model is composed of two channels of CNN, each of which learns feature from spectral domain and spatial domain respectively. The learned spectral feature and spatial feature are then concatenated and fed to fully connected layer to extract joint spectral-spatial feature for classification. When number of training samples is limited, we propose to train the deep model using transfer learning to improve the performance. Low-layer and mid-layer features of the deep model are learned and transferred from other scenes, only top-layer feature is learned using the limited training samples of the current scene. Experiment results on real data demonstrate the effectiveness of the proposed method.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS2
2016 A hyperspectral spatial-spectral enhancement algorithm
abstract
Low spatial and spectral resolution hyperspectral image will always degrade the performance of the subsequent applications, such as classification and object detection. The desired hyperspectral image is assumed to be reconstructed based on both high spatial and spectral features, which are always represented using endmembers and their abundances. In this paper, we propose a hyperspectral spatial and spectral resolution enhancement algorithm based on spectral unmixing and spatial constraints to simultaneously obtain high spatial-spectral resolution result. An intermediate high spatial but low spectral resolution HSI is introduced to establish mapping scheme of abundances and endmembers between low resolution input and desired high spatial-spectral resolution result. Experiments on the Sandigo dataset have illustrated that the proposed method is comparable or superior to other state-of-art methods.
Yongqiang Zhao 0001, Jingxiang Yang
IGARSS2
2016 Orthogonal Nonnegative Matrix Factorization Combining Multiple Features for Spectral-Spatial Dimensionality Reduction of Hyperspectral Imagery
abstract
Nonnegative matrix factorization (NMF), which can lead to nonsubtractive parts-based representation, has been demonstrated to be effective for dimensionality reduction of hyperspectral imagery (HSI). However, existing NMF methods applied to HSI use only a single spectral feature and do not take into consideration spatial information, such as texture or morphological features, while it has been widely acknowledged that exploiting multiple features can improve performance. Consequently, a variant of orthogonal NMF, which can not only achieve a nonnegative factorization but also exploit the complementary information that arises among heterogeneous features, is proposed for hyperspectral dimensionality reduction. The proposed method, which couples orthogonal NMF with a previous multiple-features-combining algorithm, yields a discriminative low-dimensional feature representation that matches the intuition that parts should sum to produce a whole. An efficient multiplicative updating procedure is derived, and its local convergence is guaranteed theoretically. Experimental results on two hyperspectral data sets demonstrate the effectiveness of the proposed method.
Jinhuan Wen, James E. Fowler, Mingyi He, Yongqiang Zhao 0001, Chengzhi Deng, Vineetha Menon
IEEE Trans. Geosci. Remote. Sens.4
2016 Coupled Sparse Denoising and Unmixing With Low-Rank Constraint for Hyperspectral Image
abstract
Hyperspectral image (HSI) denoising is significant for correct interpretation. In this paper, a sparse representation framework that unifies denoising and spectral unmixing in a closed-loop manner is proposed. While conventional approaches treat denoising and unmixing separately, the proposed scheme utilizes spectral information from unmixing as feedback to correct spectral distortion. Both denoising and spectral unmixing act as constraints to the others and are solved iteratively. Noise is suppressed via sparse coding, and fractional abundance in spectral unmixing is estimated using the sparsity prior of endmembers from a spectral library. The abundance of endmembers is used as a spectral regularizer for denoising based on the hypothesis that spectral signatures obtained from a denoising process result are close to those of unmixing. Unmixing restrains spectral distortion and results in better denoising, which reciprocally leads to further improvements in unmixing. The strength of our proposed method is illustrated by simulated and real HSIs with performance competitive to the state-of-the-art denoising and unmixing methods.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.2
2015 Specular reflection removal using local structural similarity and chromaticity consistency
abstract
This paper presents a technique that removes specular reflection in an image without sacrificing texture and chromaticity fidelity in diffuse regions of the image. This method is based on the observation that there exist structural similarity both in diffuse and specular components with original highlight image as well as strong local correlation in chromaticity in diffuse component. Local correlation in chromaticity regularizes chromaticity consistency. Structural similarity is measured using a gradient magnitude similarity map of a diffuse region. A diffuse component is estimated by solving local structure and chromaticity joint compensation problem. Experiment results show that the proposed method outperforms state-of-the-art specular reflection removal methods in terms of visual evaluation and the degree of information loss.
Yongqiang Zhao 0001, Qunnie Peng, Jize Xue, Seong G. Kong
ICIP1
2015 Hyperspectral Image Denoising via Sparse Representation and Low-Rank Constraint
abstract
Hyperspectral image (HSI) denoising is an essential preprocess step to improve the performance of subsequent applications. For HSI, there is much global and local redundancy and correlation (RAC) in spatial/spectral dimensions. In addition, denoising performance can be improved greatly if RAC is utilized efficiently in the denoising process. In this paper, an HSI denoising method is proposed by jointly utilizing the global and local RAC in spatial/spectral domains. First, sparse coding is exploited to model the global RAC in the spatial domain and local RAC in the spectral domain. Noise can be removed by sparse approximated data with learned dictionary. At this stage, only local RAC in the spectral domain is employed. It will cause spectral distortion. To compensate the shortcoming of local spectral RAC, low-rank constraint is used to deal with the global RAC in the spectral domain. Different hyperspectral data sets are used to test the performance of the proposed method. The denoising results by the proposed method are superior to results obtained by other state-of-the-art hyperspectral denoising methods.
Yongqiang Zhao 0001, Jingxiang Yang
IEEE Trans. Geosci. Remote. Sens.1
2014 Coupled hyperspectral super-resolution and unmixing
abstract
The acquired hyperspectral data are always in low resolution in both spatial and spectral domains, which will result in lots of mixed pixels and degrade the detection and recognition performance in civil and military applications. So many super resolution techniques are applied to overcome this limit. In this paper, we propose a coupled hyperspectral spatial super-resolution and spectral unmixing method based on sparse representation. Combing spatial super-resolution and spectral unmixing can precisely conserve both spatial information and spectral correlation among different bands. Spectral unmixing is taken as a regularization term in spatial super-resolution to test spectral consistency and avoid spectral distortion, while spatial super-resolution is used to enhance the resolution of abundance map after spectral unmixing.
Yongqiang Zhao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan
IGARSS1
2013 Hyperspectral image denoising via sparsity and low rank
abstract
Hyperspectral noise is unavoidable in capture and transmission process, and it will degrade the detection and classification performance greatly. Noise free signal can be approximated using few atom or basis, while noisy signal is not. There are lots of similar spatial-spectral structures in noise free hyperspectral image. On the other hand, hyperspectral image of different bands are highly correlated, the rank of hyperspectral data should be low. Based on these ideas, in this paper, we propose a hyperspectral denoising method in sparse representation framework with low rank and nonlocal regulation. Numerical experiment demonstrates that proposed denoising result is competitive with the state of art algorithm.
Yongqiang Zhao 0001, Jinxiang Yang
IGARSS1
2013 Automated classification of touching or overlapping M-FISH chromosomes by region fusion and homolog pairing
Yongqiang Zhao 0001, Seong G. Kong
Pattern Anal. Appl.1
2013 Joint segmentation and pairing of multispectral chromosome images
Yongqiang Zhao 0001, Xiaolin Wu 0001, Seong G. Kong, Lei Zhang 0006
Pattern Anal. Appl.1
2012 Hyperspectral imagery super-resolution by image fusion and compressed sensing
abstract
Low spatial resolution is the mainly drawback of hyperspectral imaging. Image super-resolution techniques can be applied to overcome the limits. This paper presents a new framework for improving the spatial resolution of hyperspectral images base by combing high-resolution spectral information and high-resolution spatial information by image fusion and compressed sensing. Based on the compressed sensing theory, small patches of hyperspectral observations from different wavelengths can be represented as weighted linear combinations of a small number of atoms in dictionary which is trained by using panchromatic images. Then hyperspectral image super-resolution is treated as a special image fusion problem with sparse constraints. To make the super-resolution reconstruction more accurate, local manifold projection is used as a regulation term. Extensive experiments on image super-resolution validate that proposed method achieves much better results.
Yongqiang Zhao 0001, Yaozhong Yang, Qingyong Zhang, Jinxiang Yang
IGARSS1
2011 Unsupervised Classification of Spectropolarimetric Data by Region-Based Evidence Fusion
abstract
Imaging spectropolarimetry is a new sensing method that can acquire the spectral, polarimetric, and spatial information of an interesting scene. They give the incomplete representations of a scene, respectively, and it is expected that combination of them will improve confidence in target identification and quality of the scene description. In this letter, a divide-and-conquer-based unsupervised spectropolarimetric data classification method is proposed to utilize the spatial, spectral, and polarimetric information jointly. First, a spectropolarimetric projection scheme is proposed to divide the whole data set into two parts: spatial-spectral and spatial-polarimetric domains. Then, a nonparametric technique is used to extract the homogeneous regions in these two domains. Each homogeneous region offers a reference spectrum and polarization, based on which a pseudosupervised spectropolarimetric classification scheme is developed by using evidence theory to fuse the information provided by the spectrum and polarization. The experimental results on real spectropolarimetric data demonstrate that the proposed divide-and-conquer-based classification scheme can achieve higher accuracy than the fuzzyc-means clustering method with spatial information constraints, which takes into account the spatial information during spectral and polarimetric clustering. Moreover, the experimental results also show the potential of spectropolarimetric classification.
Yongqiang Zhao 0001, Feiran Jie, Shibo Gao, Quan Pan 0001
IEEE Geosci. Remote. Sens. Lett.1
2011 Band-Subset-Based Clustering and Fusion for Hyperspectral Imagery Classification
abstract
This paper proposes a band-subset-based clustering and fusion technique to improve the classification performance in hyperspectral imagery. The proposed method can account for the varying data qualities and discrimination capabilities across spectral bands, and utilize the spectral and spatial information simultaneously. First, the hyperspectral data cube is partitioned into several nearly uncorrelated subsets, and an eigenvalue-based approach is proposed to evaluate the confidence of each subset. Then, a nonparametric technique is used to extract the arbitrarily-shaped clusters in spatial-spectral domain. Each cluster offers a reference spectral, based on which a pseudosupervised hyperspectral classification scheme is developed by using evidence theory to fuse the information provided by each subset. The experimental results on real Hyperspectral Digital Imagery Collection Experiment (HYDICE) demonstrate that the proposed pseudosupervised classification scheme can achieve higher accuracy than the spatially constrained fuzzy c-means clustering method. It can achieve nearly the same accuracy as the supervised K-Nearest Neighbor (KNN) classifier but is more robust to noise.
Yongqiang Zhao 0001, Lei Zhang 0006, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.1
2008 Object Detection by Spectropolarimeteric Imagery Fusion
abstract
In the past few years, imaging spectroscopy has been widely used. However, it only acquires intensity information in a narrow electromagnetic band, ignoring the polarimetric information of the electromagnetic wave and resulting in inaccurate object detection. According to electromagnetic theory, the reflected spectral signature depends on the elemental composition of objects residing within the scene, and the radiation's polarization characteristic is sensitive to surface features, such as relative smoothness and conductance. Independently, spectral and polarimetric features give incomplete representations of an object of interest. These representations are complementary, and it is expected that the combination of complementary information will reduce false alarms, improve confidence in target identification, and improve the quality of the scene description. Imaging spectropolarimetric technology as a new sensing method can acquire polarimetric information at narrow electromagnetic bands, but there are a few results showing how to combine these complementary features to detect objects in clutter. In this paper, a spectropolarimetric projection scheme is proposed to divide the spectropolarimetric data set into two parts: (1) a polarimetric spectrum data set and (2) a polarimetric data cube. Then, the polarimetric spectrum anomaly feature extraction method is used to deal with the polarimetric spectrum data set, and the adaptive polarimetric information fusion method is proposed to extract the feature from the polarimetric data cube. Finally, these features are combined by the Choquet integral to achieve better detection performance. The algorithm is applied to one complex scene, and detailed detection performance is evaluated.
Yongqiang Zhao 0001, Quan Pan 0001
IEEE Trans. Geosci. Remote. Sens.1