VLDB 2026 Research / reviewers in the wild / expert
Lizhi Wang 0001
dblp:37/5780-1
· DBLP profile ↗
57ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0002-1953-3339ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 7 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Physics-Informed Noise Models from Dark Frames for Low-Light Raw Image DenoisingabstractRecently, the mainstream practice for training low-light raw image denoising methods has shifted towards employing synthetic data. Noise modeling, which focuses on characterizing the noise distribution of real-world sensors, profoundly influences the effectiveness and practicality of synthetic data. Currently, physics-based noise modeling struggles to characterize the entire real noise distribution, while learning-based noise modeling impractically depends on paired real data. In this paper, we propose a novel strategy: learning the noise model from dark frames instead of paired real data, to break down the data dependency. Based on this strategy, we introduce an efficient physics-informed noise neural proxy (PNNP) to approximate the real-world sensor noise model. Specifically, we integrate physical priors into neural proxies and introduce three efficient techniques: physics-guided noise decoupling (PND), physics-aware proxy model (PPM), and differentiable distribution loss (DDL). PND decouples the dark frame into different components and handles different levels of noise flexibly, which reduces the complexity of noise modeling. PPM incorporates physical priors to constrain the synthetic noise, which promotes the accuracy of noise modeling. DDL provides explicit and reliable supervision for noise distribution, which promotes the precision of noise modeling. PNNP exhibits powerful potential in characterizing the real noise distribution. Extensive experiments on public datasets demonstrate superior performance in practical low-light raw image denoising. The source code will be publicly available at the https://fenghansen.github.io/publication/PNNP. Hansen Feng, Lizhi Wang 0001, Yiqi Huang, Yuzhi Wang, Lin Zhu 0012, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | DSNeRF: Dynamic View Synthesis for Ultra-Fast Scenes From Continuous Spike StreamsabstractSpike cameras generate binary spikes in response to light intensity changes, enabling high-speed visual perception with unprecedented temporal resolution. However, the unique characteristics of spike stream present significant challenges for reconstructing dense 3D scene representations, particularly in dynamic environments and under non-ideal lighting conditions. In this paper, we introduce DSNeRF, the first method to derive a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. We propose a novel mapping from pixel rays to the spike domain, integrating the spike generation process directly into NeRF training. Specifically, DSNeRF introduces an integrate-and-fire neuron layer that models non-idealities to capture intrinsic camera noise, including both random and fixed-pattern spike noise, thereby enhancing scene fidelity. Additionally, we propose a motion-guided spiking neuron layer and a long-term rendering photometric loss to better align dynamic spike streams, ensuring accurate scene geometry. Our method optimizes neural radiance fields to render photorealistic novel views from continuous spike streams, demonstrating advantages over other vision sensors in certain scenes. Empirical evaluations on both real and simulated sequences validate the effectiveness of our approach. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Toward Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge MiningabstractExisting object detection methods struggle to generalize across increasingly data domains while simultaneously adapting to the emergence of novel categories. To tackle this challenge, adaptive open-set object detection (AOOD) has been introduced, which employs supervised training on base categories within the source domain while enabling unsupervised adaptation to both base and novel categories in the target domain. However, existing AOOD approaches are still hindered by several limitations, including insufficient cross-domain feature representation, inter-category ambiguity in novel classes, and inherent feature bias toward the source domain. To overcome these issues, this paper proposes a category-level collaboration knowledge mining strategy designed to comprehensively exploit both inter-class and intra-class feature relationships across domains. Specifically, a clustering-based memory bank (CMB) is initially constructed to aggregate class prototype features, class auxiliary features, and intra-class disparity features, thereby embedding rich category-level knowledge into a unified memory structure. The CMB is iteratively updated through unsupervised clustering, which facilitates the modeling of intra-category relationships and enhances its capacity for cross-domain knowledge representation. Subsequently, a base-to-novel selection metric (BNSM) is designed to identify features corresponding to novel categories within the source domain by regulating the relationships between the novel categories and each base category. The selected features are then leveraged to initialize the object detector for the classification of novel categories. Finally, an adaptive feature assignment (AFA) strategy is introduced to transfer the learned category-level knowledge to the target domain, enabling the assignment of category labels to features. The memory bank is updated asynchronously with these assigned features to mitigate source domain bias. Extensive experiments conducted on diverse domain datasets demonstrate that the proposed method consistently outperforms state-of-the-art AOOD approaches, achieving performance gains of 1.1 to 5.5 mAP. Code is available at https://github.com/Jandsome/CCKM. Yuqi Ji, Junjie Ke, Lihuo He, Lizhi Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image DenoisingabstractImage denoising enhances image quality, serving as a foundational technique across various computational photography applications. The obstacle to clean image acquisition in real scenarios necessitates the development of self-supervised image denoising methods only depending on noisy images, especially a single noisy image. Existing self-supervised image denoising paradigms (Noise2Noise and Noise2Void) rely heavily on information-lossy operations, such as downsampling and masking, culminating in low-quality denoising performance. In this paper, we propose a novel self-supervised single image denoising paradigm, Positive2Negative, to break the information-lossy barrier. Our paradigm involves two key steps: Renoised Data Construction (RDC) and Denoised Consistency Supervision (DCS). RDC renoises the predicted denoised image by the predicted noise to construct multiple noisy images, preserving all the information of the original image. DCS ensures consistency across the multiple denoised images, supervising the network to learn robust denoising. Our Positive2Negative paradigm achieves state-of-the-art performance in self-supervised single image denoising with significant speed improvements. The code is released to the public at https://github.com/Li-Tong-621/P2N. Tong Li 0016, Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 2 |
| 2025 | Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image DenoisingabstractExisting single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still struggle to balance the contributions of two fields in the process of image fusion. In response to this, in this paper, we develop a cross-field Frequency Correlation Exploiting Network (FCENet) for NIR-assisted image denoising. We first propose the frequency correlation prior based on an in-depth statistical frequency analysis of NIR-RGB image pairs. The prior reveals the complementary correlation of NIR and RGB images in the frequency domain. Leveraging frequency correlation prior, we then establish a frequency learning framework composed of Frequency Dynamic Selection Mechanism (FDSM) and Frequency Exhaustive Fusion Mechanism (FEFM). FDSM dynamically selects complementary information from NIR and RGB images in the frequency domain, and FEFM strengthens the control of common and differential features during the fusion process of NIR and RGB features. Extensive experiments on simulated and real data validate that the proposed method outperforms other state-of-the-art methods. The code will be released at https://github.com/yuchenwang815/FCENet. Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 3 |
| 2025 | HDR Image Generation via Gain Map Decomposed Diffusion
Yuanshen Guan, Ruikang Xu, Yinuo Liao, Mingde Yao, Lizhi Wang 0001, Zhiwei Xiong |
ICCV | 5 |
| 2025 | Noise-Modeled Diffusion Models for Low-Light Spike Image Restoration
Lin Zhu 0012, Xijie Xiang, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 4 |
| 2025 | EMatch: A Unified Framework for Event-Based Optical Flow and Stereo MatchingabstractEvent cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus event data within the confines of task-specific domains, overlooking how tasks across the temporal and spatial domains can reinforce each other. In this paper, we reformulate event-based flow estimation and stereo matching as a unified dense correspondence matching problem, enabling us to solve both tasks within a single model by directly matching features in a shared representation space. Specifically, our method utilizes a Temporal Recurrent Network to aggregate event features across temporal or spatial domains, and a Spatial Contextual Attention to enhance knowledge transfer across event flows via temporal or spatial interactions. By utilizing a shared feature similarities module that integrates knowledge from event streams via temporal or spatial interactions, our network performs optical flow estimation from temporal event segment inputs and stereo matching from spatial event segment inputs simultaneously. We demonstrate that our unified model inherently supports multi-task fusion and cross-task transfer. Without the need for retraining for specific task, our model can effectively handle both optical flow and stereo estimation, achieving state-of-the-art performance on both tasks. Pengjie Zhang, Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 4 |
| 2025 | EvFocus: Learning to Reconstruct Sharp Images from Out-of-Focus Event StreamsabstractEvent cameras are innovative sensors that capture brightness changes as asynchronous events rather than traditional intensity frames. These cameras offer substantial advantages over conventional cameras, including high temporal resolution, high dynamic range, and the elimination of motion blur. However, defocus blur, a common image quality degradation resulting from out-of-focus lenses, complicates the challenge of event-based imaging. Due to the unique imaging mechanism of event cameras, existing focusing algorithms struggle to operate efficiently on sparse event data. In this work, we propose EvFocus, a novel architecture designed to reconstruct sharp images from defocus event streams for the first time. Our work includes the development of an event-based out-of-focus camera model and a simulator to generate realistic defocus event streams for robust training and testing. EvDefous integrates a temporal information encoder, a blur-aware two-branch decoder, and a reconstruction and re-defocus module to effectively learn and correct defocus blur. Extensive experiments on both simulated and real-world datasets demonstrate that EvFocus outperforms existing methods across varying lighting conditions and blur sizes, proving its robustness and practical applicability in event-based defocus imaging. Lin Zhu 0012, Xiantao Ma, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ICML | 4 |
| 2025 | Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse Events
Lin Zhu 0012, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2025 | APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechabstractAutomatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in predicting mean opinion scores (MOS) to assess synthetic speech, the neglect of fundamental auditory perception mechanisms limits consistency with human judgments. To address this issue, we propose an auditory perception guided-MOS prediction model (APG-MOS) that synergistically integrates auditory modeling with semantic analysis to enhance consistency with human judgments. Specifically, we first design a perceptual module, grounded in biological auditory mechanisms, to simulate cochlear functions, which encodes acoustic signals into biologically aligned electrochemical representations. Secondly, we propose a residual vector quantization (RVQ)-based semantic distortion modeling method to quantify the degradation of speech quality at the semantic level. Finally, we design a residual cross-attention module, coupled with a progressive learning strategy, to enable multimodal fusion of encoded electrochemical signals and semantic representations. Experiments demonstrate that APG-MOS achieves superior performance on two primary benchmarks. The implementation code is available at https://github.com/BNU-ERC-ITEA/APG-MOS. Zhicheng Lian, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 2 |
| 2025 | Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionabstractEvent cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms. Code: https://github.com/BIT-Vision/SATE. Lin Zhu 0012, Tengyu Long, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
NeurIPS | 4 |
| 2025 | Towards Ultra High-Speed Hyperspectral Imaging by Integrating Compressive and Neuromorphic Sampling
Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Learning Refractive-Diffractive Optics with Unidirectional Transformer for Large Field-of-View Imaging
Xiangtian Ma, Lizhi Wang 0001, Qilin Sun 0001, Lin Zhu 0012, Hua Huang 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Continuous Spatial-Spectral Reconstruction via Implicit Neural Representation
Ruikang Xu, Mingde Yao, Chang Chen 0004, Lizhi Wang 0001, Zhiwei Xiong |
Int. J. Comput. Vis. | 4 |
| 2025 | Event-Enhanced Snapshot Mosaic Hyperspectral Frame DeblurringabstractSnapshot Mosaic Hyperspectral Cameras (SMHCs) are popular hyperspectral imaging devices for acquiring both color and motion details of scenes. However, the narrow-band spectral filters in SMHCs may negatively impact their motion perception ability, resulting in blurry SMHC frames. In this paper, we propose a hardware-software collaborative approach to address the blurring issue of SMHCs. Our approach involves integrating SMHCs with neuromorphic event cameras for efficient event-enhanced SMHC frame deblurring. To achieve spectral information recovery guided by event signals, we formulate a spectral-aware Event-based Double Integral (sEDI) model that links SMHC frames and events from a spectral perspective, providing principled model design insights. Then, we develop a Diffusion-guided Noise Awareness (DNA) training framework that utilizes diffusion models to learn noise-aware features and promote model robustness towards camera noise. Furthermore, we design an Event-enhanced Hyperspectral frame Deblurring Network (EvHDNet) based on sEDI, which is trained with DNA and features improved spatial-spectral learning and modality interaction for reliable SMHC frame deblurring. Experiments on both synthetic data and real data show that the proposed DNA + EvHDNet outperforms state-of-the-art methods on both spatial and spectral fidelity. The code and dataset will be made publicly available. Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Building Non-Uniform Degradation Model for Position-Aware Hyperspectral Image FusionabstractThe fusion of low-spatial-resolution hyperspectral image (LR-HSI) with high-spatial-resolution multispectral image (HR-MSI) has become an effective way to obtain the high-spatial-resolution hyperspectral image (HR-HSI). Currently, learning-based methods have emerged as the mainstream solution in this field. However, these methods typically rely on predefined or simplified degradation models during fusion training, resulting in inaccurate supervision of the fusion networks. Meanwhile, most methods overlook the degradation characteristics in designing the fusion networks, leading to a mismatch between the degradation and fusion processes. These limitations ultimately result in unsatisfactory fusion performance on real data. To enhance the practicality of learning-based methods, accurate degradation modeling and effective network design have become the critical priorities. We observe that, in practical scenarios, the degree of pixel degradation varies across different positions due to the unforeseen factors such as illumination variations and imaging system fluctuations. Considering this, we propose a non-uniform degradation model (NUD), which introduces non-uniformity into the degradation processes of LR-HSI and HR-MSI. In addition, we emphasize that the essence of fusion is to reverse the degradation process. Therefore, to align with the non-uniform degradation process, the fusion process should exhibit similar positional specificity. For this purpose, we propose a position-aware fusion network (PAF), which employs positional encoding to endow the fusion process with the position-aware attribute. Experimental results show that our proposed methods provide an effective solution for HSI fusion in practical scenarios. Lizhi Wang 0001, Lin Zhu 0012, Renwei Dian, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Continuous-Time Object Segmentation Using High Temporal Resolution Event CameraabstractEvent cameras are novel bio-inspired sensors, where individual pixels operate independently and asynchronously, generating intensity changes as events. Leveraging the microsecond resolution (no motion blur) and high dynamic range (compatible with extreme light conditions) of events, there is considerable promise in directly segmenting objects from sparse and asynchronous event streams in various applications. However, different from the rich cues in video object segmentation, it is challenging to segment complete objects from the sparse event stream. In this paper, we present the first framework for continuous-time object segmentation from event stream. Given the object mask at the initial time, our task aims to segment the complete object at any subsequent time in event streams. Specifically, our framework consists of a Recurrent Temporal Embedding Extraction (RTEE) module based on a novel ResLSTM, a Cross-time Spatiotemporal Feature Modeling (CSFM) module which is a transformer architecture with long-term and short-term matching modules, and a segmentation head. The historical events and masks (reference sets) are recurrently fed into our framework along with current-time events. The temporal embedding is updated as new events are input, enabling our framework to continuously process the event stream. To train and test our model, we construct both real-world and simulated event-based object segmentation datasets, each comprising event streams, APS images, and object annotations. Extensive experiments on our datasets demonstrate the effectiveness of the proposed recurrent architecture. Lin Zhu 0012, Xianzhang Chen, Lizhi Wang 0001, Xiao Wang 0014, Yonghong Tian 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | GT-HAD: Gated Transformer for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to distinguish between the background and anomalies in a scene, which has been widely adopted in various applications. Deep neural network (DNN)-based methods have emerged as the predominant solution, wherein the standard paradigm is to discern the background and anomalies based on the error of self-supervised hyperspectral image (HSI) reconstruction. However, current DNN-based methods cannot guarantee correspondence between the background, anomalies, and reconstruction error, which limits the performance of HAD. In this article, we propose a novel gated transformer network for HAD (GT-HAD). Our key observation is that the spatial-spectral similarity in HSI can effectively distinguish between the background and anomalies, which aligns with the fundamental definition of HAD. Consequently, we develop GT-HAD to exploit the spatial-spectral similarity during HSI reconstruction. GT-HAD consists of two distinct branches that model the features of the background and anomalies, respectively, with content similarity as constraints. Furthermore, we introduce an adaptive gating unit to regulate the activation states of these two branches based on a content-matching method (CMM). Extensive experimental results demonstrate the superior performance of GT-HAD. The original code is publicly available at https://github.com/jeline0110/ GT-HAD, along with a comprehensive benchmark of state-of-the-art HAD methods. Lizhi Wang 0001, He Sun 0009, Hua Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | SpikeNeRF: Learning Neural Radiance Fields from Continuous Spike StreamabstractSpike cameras, leveraging spike-based integration sampling and high temporal resolution, offer distinct advan-tages over standard cameras. However, existing approaches reliant on spike cameras often assume optimal illumination, a condition frequently unmet in real-world scenarios. To address this, we introduce SpikeNeRF, the first work that derives a NeRF-based volumetric scene representation from spike camera data. Our approach leverages NeRF's multi-view consistency to establish robust self-supervision, effectively eliminating erroneous measurements and uncovering coherent structures within exceedingly noisy input amidst diverse real-world illumination scenarios. The framework comprises two core elements: a spike generation model incorporating an integrate-and-fire neuron layer and parameters accounting for non-idealities, such as threshold variation, and a spike rendering loss capable of general-izing across varying illumination conditions. We describe how to effectively optimize neural radiance fields to render photorealistic novel views from the novel continuous spike stream, demonstrating advantages over other vision sen-sors in certain scenes. Empirical evaluations conducted on both real and novel realistically simulated sequences affirm the efficacy of our methodology. The dataset and source code are released at https://github.com/BIT-Vision/SpikeNeRF. Lin Zhu 0012, Kangmin Jia, Yifan Zhao 0002, Yunshan Qi, Lizhi Wang 0001, Hua Huang 0001 |
CVPR | 5 |
| 2024 | Event-Based Visible and Infrared Fusion via Multi-Task CollaborationabstractVisible and Infrared image Fusion (VIF) offers a comprehensive scene description by combining thermal infrared images with the rich textures from visible cameras. However, conventional VIF systems may capture over/under exposure or blurry images in extreme lighting and high dynamic motion scenarios, leading to degraded fusion results. To address these problems, we propose a novel Event-based Visible and Infrared Fusion (EVIF) system that employs a visible event camera as an alternative to traditional frame-based cameras for the VIF task. With extremely low latency and high dynamic range, event cameras can effectively address blurriness and are robust against diverse luminous ranges. To produce high-quality fused images, we develop a multitask collaborative framework that simultaneously performs event-based visible texture reconstruction, event-guided infrared image deblurring, and visible-infrared fusion. Rather than independently learning these tasks, our framework capitalizes on their synergy, leveraging cross-task event enhancement for efficient deblurring and bi-level min-max mutual information optimization to achieve higher fusion quality. Experiments on both synthetic and real data show that EVIF achieves remarkable performance in dealing with extreme lighting conditions and high-dynamic scenes, ensuring high-quality fused images across a broad range of practical scenarios. Mengyue Geng, Lin Zhu 0012, Lizhi Wang 0001, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
CVPR | 3 |
| 2024 | In2SET: Intra-Inter Similarity Exploiting Transformer for Dual-Camera Compressive Hyperspectral ImagingabstractDual-camera compressive hyperspectral imaging (DC-CHI) offers the capability to reconstruct 3D hyperspectral image (HSI) by fusing compressive and panchromatic (PAN) image, which has shown great potential for snapshot hyperspectral imaging in practice. In this paper, we introduce a novel DCCHI reconstruction network, intra-inter similarity exploiting Transformer (In2SET). Our key insight is to make full use of the PAN image to assist the reconstruction. To this end, we propose to use the intra-similarity within the PAN image as a proxy for approximating the intra-similarity in the original HSI, thereby offering an enhanced content prior for more accurate HSI reconstruction. Furthermore, we propose to use the inter-similarity to align the features between HSI and PAN images, thereby maintaining semantic consistency between the two modalities during the reconstruction process. By integrating In2SET into a PAN-guided deep unrolling (PGDU)framework, our method substantially enhances the spatial-spectral fidelity and detail of the reconstructed images, providing a more comprehensive and accurate depiction of the scene. Experiments conducted on both real and simulated datasets demonstrate that our approach consistently outperforms existing state-of-the-art methods in terms of reconstruction quality and computational complexity. The code is available at https://github.com/2JONAS/In2SET. Lizhi Wang 0001, Xiangtian Ma, Maoqing Zhang, Lin Zhu 0012, Hua Huang 0001 |
CVPR | 2 |
| 2024 | Learning Exhaustive Correlation for Spectral Super-Resolution: Where Spatial-Spectral Attention Meets Linear Dependence
Lizhi Wang 0001, Chang Chen 0004, Fenglong Song, Youliang Yan |
ECCV (25) | 2 |
| 2024 | Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
Lin Zhu 0012, Yunlong Zheng, Yijun Zhang 0003, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001 |
ECCV (40) | 5 |
| 2024 | MLP Embedded Inverse Tone MappingabstractThe advent of High Dynamic Range/Wide Color Gamut (HDR/WCG) display technology has made significant progress in providing exceptional richness and vibrancy for the human visual experience. However, the widespread adoption of HDR/WCG images is hindered by their substantial storage requirements, imposing significant bandwidth challenges during distribution. Besides, HDR/WCG images are often tone-mapped into Standard Dynamic Range (SDR) versions for compatibility, necessitating the usage of inverse Tone Mapping (iTM) techniques to reconstruct their original representation. In this work, we propose a meta-transfer learning framework for practical HDR/WCG media transmission by embedding image-wise metadata into their SDR counterparts for later iTM reconstruction. Specifically, we devise a meta-learning strategy to pre-train a lightweight multilayer perceptron (MLP) model that maps SDR pixels to HDR/WCG ones on an external dataset, resulting in a domain-wise iTM model. Subsequently, for the transfer learning process of each HDR/WCG image, we present a spatial-aware online mining mechanism to select challenging training pairs to adapt the meta-trained model to an image-wise iTM model. Finally, the adapted MLP, embedded as metadata, is transmitted alongside the SDR image, facilitating the reconstruction of the original image on HDR/WCG displays. We conduct extensive experiments and evaluate the proposed framework with diverse metrics. Compared with existing solutions, our framework shows superior performance in fidelity, minimal latency, and negligible overhead. The codes are available at https://github.com/pjliu3/MLP_iTM. Panjun Liu, Jiacheng Li 0004, Lizhi Wang 0001, Zhengjun Zha, Zhiwei Xiong |
ACM Multimedia | 3 |
| 2024 | Learning a physics-based filter attachment for hyperspectral imaging with RGB cameras
Maoqing Zhang, Lizhi Wang 0001, Lin Zhu 0012, Hua Huang 0001 |
Neurocomputing | 2 |
| 2024 | Learnability Enhancement for Low-Light Raw Image Denoising: A Data PerspectiveabstractLow-light raw image denoising is an essential task in computational photography, to which the learning-based method has become the mainstream solution. The standard paradigm of the learning-based method is to learn the mapping between the paired real data, i.e., the low-light noisy image and its clean counterpart. However, the limited data volume, complicated noise model, and underdeveloped data quality have constituted the learnability bottleneck of the data mapping between paired real data, which limits the performance of the learning-based method. To break through the bottleneck, we introduce a learnability enhancement strategy for low-light raw image denoising by reforming paired real data according to noise modeling. Our learnability enhancement strategy integrates three efficient methods: shot noise augmentation (SNA), dark shading correction (DSC) and a developed image acquisition protocol. Specifically, SNA promotes the precision of data mapping by increasing the data volume of paired real data, DSC promotes the accuracy of data mapping by reducing the noise complexity, and the developed image acquisition protocol promotes the reliability of data mapping by improving the data quality of paired real data. Meanwhile, based on the developed image acquisition protocol, we build a new dataset for low-light raw image denoising. Experiments on public datasets and our dataset demonstrate the superiority of the learnability enhancement strategy. Hansen Feng, Lizhi Wang 0001, Yuzhi Wang, Haoqiang Fan, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and EnsemblingabstractImage denoising is a fundamental problem in computational photography, where achieving high perception with low distortion is highly demanding. Current methods either struggle with perceptual quality or suffer from significant distortion. Recently, the emerging diffusion model has achieved state-of-the-art performance in various tasks and demonstrates great potential for image denoising. However, stimulating diffusion models for image denoising is not straightforward and requires solving several critical problems. For one thing, the input inconsistency hinders the connection between diffusion models and image denoising. For another, the content inconsistency between the generated image and the desired denoised image introduces distortion. To tackle these problems, we present a novel strategy called the Diffusion Model for Image Denoising (DMID) by understanding and rethinking the diffusion model from a denoising perspective. Our DMID strategy includes an adaptive embedding method that embeds the noisy image into a pre-trained unconditional diffusion model and an adaptive ensembling method that reduces distortion in the denoised image. Our DMID strategy achieves state-of-the-art performance on both distortion-based and perception-based metrics, for both Gaussian and real-world image denoising. Tong Li 0016, Hansen Feng, Lizhi Wang 0001, Lin Zhu 0012, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Non-Serial Quantization-Aware Deep Optics for Snapshot Hyperspectral ImagingabstractDeep optics has been endeavoring to capture hyperspectral images of dynamic scenes, where the optical encoder plays an essential role in deciding the imaging performance. Our key insight is that the optical encoder of a deep optics system is expected to keep fabrication-friendliness and decoder-friendliness, to be faithfully realized in the implementation phase and fully interacted with the decoder in the design phase, respectively. In this paper, we propose the non-serial quantization-aware deep optics (NSQDO), which consists of the fabrication-friendly quantization-aware model (QAM) and the decoder-friendly non-serial manner (NSM). The QAM integrates the quantization process into the optimization and adaptively adjusts the physical height of each quantization level, reducing the deviation of the physical encoder from the numerical simulation through the awareness of and adaptation to the quantization operation of the DOE physical structure. The NSM bridges the encoder and the decoder with full interaction through bidirectional hint connections and flexibilize the connections with a gating mechanism, boosting the power of joint optimization in deep optics. The proposed NSQDO improves the fabrication-friendliness and decoder-friendliness of the encoder and develops the deep optics framework to be more practical and powerful. Extensive synthetic simulation and real hardware experiments demonstrate the superior performance of the proposed method. Lizhi Wang 0001, Lingen Li, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Hyperbolic Space-Based Autoencoder for Hyperspectral Anomaly DetectionabstractDeep-learning (DL)-based methods have been shown to be effective on the hyperspectral image (HSI) anomaly detection task because of their feature extraction ability. However, current DL-based methods lack an effective means of regularizing the background information. In this article, the hyperbolic space-based autoencoder (HSAE) is proposed for the hyperspectral anomaly detection task. We assume that an effective hierarchical structural representation can better model the HSI in the spatial domain, and this enables the background information to be effectively regularized. Motivated by this idea, the HSAE embeds the HSI into hyperbolic space, which is a non-Euclidean geometry with a constant negative curvature and an exponential growth distance between points. Using a wrapped normal prior distribution, the training of the hidden representation is supervised to preserve more hierarchical features. After the training process, a hyperbolic distance-based anomaly detector (HDB) is introduced to discover anomalies in a more robust way. Experimental results on several popular HSI benchmarks fully demonstrate the superiority of our HSAE. He Sun 0009, Lizhi Wang 0001, Lei Zhang 0021, Lianru Gao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Mutual-Guided Dynamic Network for Image FusionabstractImage fusion aims to generate a high-quality image from multiple images captured under varying conditions. The key problem of this task is to preserve complementary information while filtering out irrelevant information for the fused result. However, existing methods address this problem by leveraging static convolutional neural networks (CNNs), suffering two inherent limitations during feature extraction,i.e., being unable to handle spatial-variant contents and lacking guidance from multiple inputs. In this paper, we propose a novel mutual-guided dynamic network (MGDN) for image fusion, which allows for effective information utilization across different locations and inputs. Specifically, we design a mutual-guided dynamic filter (MGDF) for adaptive feature extraction, composed of a mutual-guided cross-attention (MGCA) module and a dynamic filter predictor, where the former incorporates additional guidance from different inputs and the latter generates spatial-variant kernels for different locations. In addition, we introduce a parallel feature fusion (PFF) module to effectively fuse local and global information of the extracted features. To further reduce the redundancy among the extracted features while simultaneously preserving their shared structural information, we devise a novel loss function that combines the minimization of normalized mutual information (NMI) with an estimated gradient mask. Experimental results on five benchmark datasets demonstrate that our proposed method outperforms existing methods on four image fusion tasks. The code and model are publicly available at: https://github.com/Guanys-dar/MGDN. Yuanshen Guan, Ruikang Xu, Mingde Yao, Lizhi Wang 0001, Zhiwei Xiong |
ACM Multimedia | 4 |
| 2023 | Learning Spectral-wise Correlation for Spectral Super-Resolution: Where Similarity Meets ParticularityabstractHyperspectral images consist of multiple spectral channels, and the task of spectral super-resolution is to reconstruct hyperspectral images from 3-channel RGB images, where modeling spectral-wise correlation is of great importance. Based on the analysis of the physical process of this task, we distinguish the spectral-wise correlation into two aspects: similarity and particularity. The Existing Transformer model cannot accurately capture spectral-wise similarity due to the inappropriate spectral-wise fully connected linear mapping acting on input spectral feature maps, which results in spectral feature maps mixing. Moreover, the token normalization operation in the existing Transformer model also results in its inability to capture spectral-wise particularity and thus fails to extract key spectral feature maps. To address these issues, we propose a novel Hybrid Spectral-wise Attention Transformer (HySAT). The key module of HySAT is Plausible Spectral-wise self-Attention (PSA), which can simultaneously model spectral-wise similarity and particularity. Specifically, we propose a Token Independent Mapping (TIM) mechanism to reasonably model spectral-wise similarity, where a linear mapping shared by spectral feature maps is applied on input spectral feature maps. Moreover, we propose a Spectral-wise Re-Calibration (SRC) mechanism to model spectral-wise particularity and effectively capture significant spectral feature maps. Experimental results show that our method achieves state-of-the-art performance in the field of spectral super-resolution with the lowest error and computational costs. Lizhi Wang 0001, Chang Chen 0004, Fenglong Song, Hua Huang 0001 |
ACM Multimedia | 2 |
| 2023 | Recurrent Spike-based Image Restoration under General IlluminationabstractSpike camera is a new type of bio-inspired vision sensor that records light intensity in the form of a spike array with high temporal resolution (20,000 Hz). This new paradigm of vision sensor offers significant advantages for many vision tasks such as high speed image reconstruction. However, existing spike-based approaches typically assume that the scenes are with sufficient light intensity, which is usually unavailable in many real-world scenarios such as rainy days or dusk scenes. To unlock more spike-based application scenarios, we propose a Recurrent Spike-based Image Restoration (RSIR) network, which is the first work towards restoring clear images from spike arrays under general illumination. Specifically, to accurately describe the noise distribution under different illuminations, we build a physical-based spike noise model according to the sampling process of the spike camera. Based on the noise model, we design our RSIR network which consists of an adaptive spike transformation module, a recurrent temporal feature fusion module, and a frequency-based spike denoising module. Our RSIR can process the spike array in a recursive manner to ensure that the spike temporal information is well utilized. In the training process, we generate the simulated spike data based on our noise model to train our network. Extensive experiments on real-world datasets with different illuminations demonstrate the effectiveness of the proposed network. The code and dataset are released at https://github.com/BIT-Vision/RSIR. Lin Zhu 0012, Yunlong Zheng, Mengyue Geng, Lizhi Wang 0001, Hua Huang 0001 |
ACM Multimedia | 4 |
| 2022 | Quantization-aware Deep Optics for Diffractive Snapshot Hyperspectral ImagingabstractDiffractive snapshot hyperspectral imaging based on the deep optics framework has been striving to capture the spectral images of dynamic scenes. However, existing deep optics frameworks all suffer from the mismatch between the optical hardware and the reconstruction algorithm due to the quantization operation in the diffractive optical element (DOE) fabrication, leading to the limited performance of hyperspectral imaging in practice. In this paper, we propose the quantization-aware deep optics for diffractive snapshot hyperspectral imaging. Our key observation is that common lithography techniques used in fabricating DOEs need to quantize the DOE height map to a few levels, and can freely set the height for each level. Therefore, we propose to integrate the quantization operation into the DOE height map optimization and design an adaptive mechanism to adjust the physical height of each quantization level. According to the optimization, we fabricate the quantized DOE directly and build a diffractive hyperspectral snapshot imaging system. Our method develops the deep optics framework to be more practical through the awareness of and adaptation to the quantization operation of the DOE physical structure, making the fabricated DOE and the reconstruction algorithm match each other systematically. Extensive synthetic simulation and real hardware experiments validate the superior performance of our method. Lingen Li, Lizhi Wang 0001, Lei Zhang 0021, Zhiwei Xiong, Hua Huang 0001 |
CVPR | 2 |
| 2022 | Learnability Enhancement for Low-light Raw Denoising: Where Paired Real Data Meets Noise ModelingabstractLow-light raw denoising is an important and valuable task in computational photography where learning-based methods trained with paired real data are mainstream. However, the limited data volume and complicated noise distribution have constituted a learnability bottleneck for paired real data, which limits the denoising performance of learning-based methods. To address this issue, we present a learnability enhancement strategy to reform paired real data according to noise modeling. Our strategy consists of two efficient techniques: shot noise augmentation (SNA) and dark shading correction (DSC). Through noise model decoupling, SNA improves the precision of data mapping by increasing the data volume and DSC reduces the complexity of data mapping by reducing the noise complexity. Extensive results on the public datasets and real imaging scenarios collectively demonstrate the state-of-the-art performance of our method. Hansen Feng, Lizhi Wang 0001, Yuzhi Wang, Hua Huang 0001 |
ACM Multimedia | 2 |
| 2022 | Coded Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hyperspectral cameras often suffer from, coded hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Specifically, we first develop a CNN-based channel attention reconstruction network to effectively exploit the spatial-spectral correlation of the HSI. Then, the reconstruction network is learned by leveraging an arbitrary external hyperspectral dataset to exploit the general spatial-spectral correlation under adversarial loss. Finally, we customize the network by internal learning with spatial-spectral constraint and total variation regularization for each coded image, which can make use of the internal imaging model to learn specific prior for current desirable image and effectively avoids overfitting. Experimental results using both synthetic data and real images show that our method outperforms the state-of-the-art methods on several popular coded hyperspectral imaging systems under both comprehensive quantitative metrics and perceptive quality. Ying Fu 0001, Tao Zhang 0042, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Stochastic gate-based autoencoder for unsupervised hyperspectral band selection
He Sun 0009, Lei Zhang 0021, Lizhi Wang 0001, Hua Huang 0001 |
Pattern Recognit. | 3 |
| 2021 | Learning Tensor Low-Rank Prior for Hyperspectral Image ReconstructionabstractSnapshot hyperspectral imaging has been developed to capture the spectral information of dynamic scenes. In this paper, we propose a deep neural network by learning the tensor low-rank prior of hyperspectral images (HSI) in the feature domain to promote the reconstruction quality. Our method is inspired by the canonical-polyadic (CP) decomposition theory, where a low-rank tensor can be expressed as a weight summation of several rank-1 component tensors. Specifically, we first learn the tensor low-rank prior of the image features with two steps: (a) we generate rank-1 tensors with discriminative components to collect the contextual information from both spatial and channel dimensions of the image features; (b) we aggregate those rank-1 tensors into a low-rank tensor as a 3D attention map to exploit the global correlation and refine the image features. Then, we integrate the learned tensor low-rank prior into an iterative optimization algorithm to obtain an end-to-end HSI reconstruction. Experiments on both synthetic and real data demonstrate the superiority of our method. Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001 |
CVPR | 2 |
| 2021 | Adaptive Dimension-Discriminative Low-Rank Tensor Recovery for Computational Hyperspectral Imaging
Lizhi Wang 0001, Hua Huang 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Sparse additive discriminant canonical correlation analysis for multiple features fusion
Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
Neurocomputing | 2 |
| 2021 | Shared Low-Rank Correlation Embedding for Multiple Feature FusionabstractThe diversity of multimedia data in the real world usually forms heterogeneous types of feature sets. How to explore the structure information and the relationships among multiple features is still an open problem. In this paper, we propose an unsupervised subspace learning method, named the shared low-rank correlation embedding (SLRCE) for multiple feature fusion. First, in the learned subspace, we implement the low-rank representation on each feature set and enforce a shared low-rank constraint to uncover the common structure information of multiple features. Second, we develop an enhanced correlation analysis in the learned subspace for simultaneously removing the redundancy of each feature set and exploring the correlation of multiple features. Finally, we incorporate the shared low-rank representation and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple features, but also assists robust subspace learning. Our method is robust to noise in practice and can be extended to the kernel case to handle the nonlinear feature fusion. Experimental results on several typical datasets demonstrate the superior performance of the proposed methods. Zhan Wang 0007, Lizhi Wang 0001, Jun Wan 0001, Hua Huang 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | DNU: Deep Non-Local Unrolling for Computational Spectral ImagingabstractComputational spectral imaging has been striving to capture the spectral information of the dynamic world in the last few decades. In this paper, we propose an interpretable neural network for computational spectral imaging. First, we introduce a novel data-driven prior that can adaptively exploit both the local and non-local correlations among the spectral image. Our data-driven prior is integrated as a regularizer into the reconstruction problem. Then, we propose to unroll the reconstruction problem into an optimization-inspired deep neural network. The architecture of the network has high interpretability by explicitly characterizing the image correlation and the system imaging model. Finally, we learn the complete parameters in the network through end-to-end training, enabling robust performance with high spatial-spectral fidelity. Extensive simulation and hardware experiments validate the superior performance of our method over state-of-the-art methods. Lizhi Wang 0001, Maoqing Zhang, Ying Fu 0001, Hua Huang 0001 |
CVPR | 1 |
| 2020 | Structure Preserving Multi-View Dimensionality ReductionabstractThe multi-view features from multimedia data in the real-world are usually high-dimensional. How to simultaneously reduce their dimensions and explore the complementary information among multi-view features is of vital importance but challenging. In this paper, we propose a novel unsupervised method named structure preserving multi-view dimensionality reduction (SPMDR). We first propose a bilinear low-rank representation with an orthogonal constraint in the learning subspace. Then, we construct a 3-order rotated tensor among the low-rank coefficient matrices and utilize tensor nuclear norm to capture complementary information among multi-view representations. Finally, we develop a numerical algorithm for solving the proposed model. Our method is robust to noisy data and can capture the complex correlations among multi-view features. Experimental results on recognition tasks demonstrate the superior performance of SPMDR. Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
ICME | 2 |
| 2020 | Snapshot Hyperspectral Imaging Based on Weighted High-order Singular Value RegularizationabstractSnapshot hyperspectral imaging can capture the 3D hyperspectral image (HSI) with a single 2D measurement and has attracted increasing attention recently. Recovering the underlying HSI from the compressive measurement is an ill-posed problem and exploiting the image prior is essential for solving this ill-posed problem. However, existing reconstruction methods always start from modeling image prior with the 1D vector or 2D matrix and cannot fully exploit the structurally spectral-spatial nature in 3D HSI, thus leading to a poor fidelity. In this paper, we propose an effective high-order tensor optimization based method to boost the reconstruction fidelity for snapshot hyperspectral imaging. We first build high-order tensors by exploiting the spatial-spectral correlation in HSI. Then, we propose a weight high-order singular value regularization (WHOSVR) based low-rank tensor recovery model to characterize the structure prior of HSI. By integrating the structure prior in WHOSVR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented on two representative systems demonstrate that our method outperforms state-of-the-art methods. Niankai Cheng, Hua Huang 0001, Lei Zhang 0021, Lizhi Wang 0001 |
ICPR | 4 |
| 2020 | Embedding shared low-rank and feature correlation for multi-view data analysisabstractThe diversity of multimedia data in the real-world usually forms multi-view features. How to explore the structure information and correlations among multi-view features is still a challenging problem. In this paper, we propose a novel multi-view subspace learning method, named embedding shared low-rank and feature correlation (ESLRFC), for multi-view data analysis. First, in the embedding subspace, we propose a robust low-rank model on each feature set and enforce a shared low-rank constraint to characterize the common structure information of multiple feature data. Second, we develop an enhanced correlation analysis in the embedding subspace for simultaneously removing the redundancy of each feature set and exploring the correlations of multiple feature data. Finally, we incorporate the low-rank model and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple feature data, but also assists robust subspace learning. Experimental results on recognition tasks demonstrate the superior performance and noise robustness of the proposed method. Zhan Wang 0007, Lizhi Wang 0001, Lei Zhang 0021, Hua Huang 0001 |
ICPR | 2 |
| 2020 | Joint low rank embedded multiple features learning for audio-visual emotion recognition
Zhan Wang 0007, Lizhi Wang 0001, Hua Huang 0001 |
Neurocomputing | 2 |
| 2019 | Hyperspectral Image Reconstruction Using a Deep Spatial-Spectral PriorabstractRegularization is a fundamental technique to solve an ill-posed optimization problem robustly and is essential to reconstruct compressive hyperspectral images. Various hand-crafted priors have been employed as a regularizer but are often insufficient to handle the wide variety of spectra of natural hyperspectral images, resulting in poor reconstruction quality. Moreover, the prior-regularized optimization requires manual tweaking of its weight parameters to achieve a balance between the spatial and spectral fidelity of result images. In this paper, we present a novel hyperspectral image reconstruction algorithm that substitutes the traditional hand-crafted prior with a data-driven prior, based on an optimization-inspired network. Our method consists of two main parts: First, we learn a novel data-driven prior that regularizes the optimization problem with a goal to boost the spatial-spectral fidelity. Our data-driven prior learns both local coherence and dynamic characteristics of natural hyperspectral images. Second, we combine our regularizer with an optimization-inspired network to overcome the heavy computation problem in the traditional iterative optimization methods. We learn the complete parameters in the network through end-to-end training, enabling robust performance with high accuracy. Extensive simulation and hardware experiments validate the superior performance of our method over the state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Min H. Kim 0001, Hua Huang 0001 |
CVPR | 1 |
| 2019 | Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hypelrspectral cameras often suffer from, coded snapshot hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Our method can effectively exploit spatial-spectral correlation and sufficiently represent the variety nature of HSIs. Experimental results show our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Tao Zhang 0042, Ying Fu 0001, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 3 |
| 2019 | Computational Hyperspectral Imaging Based on Dimension-Discriminative Low-Rank Tensor RecoveryabstractExploiting the prior information is fundamental for the image reconstruction in computational hyperspectral imaging. Existing methods usually unfold the 3D signal as a 1D vector and treat the prior information within different dimensions in an indiscriminative manner, which ignores the high-dimensionality nature of hyperspectral image (HSI) and thus results in poor quality reconstruction. In this paper, we propose to make full use of the high-dimensionality structure of the desired HSI to boost the reconstruction quality. We first build a high-order tensor by exploiting the nonlocal similarity in HSI. Then, we propose a dimension-discriminative low-rank tensor recovery (DLTR) model to characterize the structure prior adaptively in each dimension. By integrating the structure prior in DLTR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented with both synthetic and real data demonstrate that our method outperforms state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Xiaoming Zhong, Hua Huang 0001 |
ICCV | 2 |
| 2019 | High-Speed Hyperspectral Video Acquisition By Combining Nyquist and Compressive SamplingabstractWe propose a novel hybrid imaging system to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. The proposed system consists of two branches: one branch performs Nyquist sampling in the temporal dimension while integrating the whole spectrum, resulting in a high-frame-rate panchromatic video; the other branch performs compressive sampling in the spectral dimension with longer exposures, resulting in a low-frame-rate hyperspectral video. Owing to the high light throughput and complementary sampling, these two branches jointly provide reliable measurements for recovering the underlying HSHS video. Moreover, the panchromatic video can be used to learn an over-complete 3D dictionary to represent each band-wise video sparsely, thanks to the inherent structural similarity in the spectral dimension. Based on the joint measurements and the self-adaptive dictionary, we further propose a simultaneous spectral sparse (3S) model to reinforce the structural similarity across different bands and develop an efficient computational reconstruction algorithm to recover the HSHS video. Both simulation and hardware experiments validate the effectiveness of the proposed approach. To the best of our knowledge, this is the first time that hyperspectral videos can be acquired at a frame rate up to 100fps with commodity optical elements and under ordinary indoor illumination. Lizhi Wang 0001, Zhiwei Xiong, Hua Huang 0001, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | HyperReconNet: Joint Coded Aperture Optimization and Image Reconstruction for Compressive Hyperspectral ImagingabstractCoded aperture snapshot spectral imaging (CASSI) system encodes the 3D hyperspectral image (HSI) within a single 2D compressive image and then reconstructs the underlying HSI by employing an inverse optimization algorithm, which equips with the distinct advantage of snapshot but usually results in low reconstruction accuracy. To improve the accuracy, existing methods attempt to design either alternative coded apertures or advanced reconstruction methods, but cannot connect these two aspects via a unified framework, which limits the accuracy improvement. In this paper, we propose a convolution neural network (CNN) based endto- end method to boost the accuracy by jointly optimizing the coded aperture and the reconstruction method. On the one hand, based on the nature of CASSI forward model, we design a repeated pattern for the coded aperture, whose entities are learned by acting as the network weights. On the other hand, we conduct the reconstruction through simultaneously exploiting intrinsic properties within HSI - the extensive correlations across the spatial and the spectral dimensions. By leveraging the power of deep learning, the coded aperture design and the image reconstruction are connected and optimized via a unified framework. Experimental results show that our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Lizhi Wang 0001, Tao Zhang 0042, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | HSVCNN: CNN-Based Hyperspectral Reconstruction from RGB VideosabstractHyperspectral video acquisition usually requires high complexity hardware and reconstruction algorithms. In this paper, we propose a low complexity CNN-based method for hyperspectral reconstruction from ubiquitous RGB videos, which effectively exploits the temporal redundancies within RGB videos and generates high-quality hyperspectral output. Specifically, given an RGB video, we first design an efficient motion compensation network to align the RGB frames and reduce the large motion. Then, we design a temporal-adaptive fusion network to exploit the inter-frame correlation. The fusion network has the ability to determine the optimum temporal dependency within successive frames, which further promotes the hyperspectral reconstruction fidelity. Preliminary experimental results validate the superior performance of the proposed method over previous learning-based methods. To the best of our knowledge, this is the first time that RGB videos are utilized for hyperspectral reconstruction through deep learning. Huiqun Li, Zhiwei Xiong, Lizhi Wang 0001, Dong Liu 0002, Feng Wu 0001 |
ICIP | 4 |
| 2018 | Simultaneous Depth and Spectral Imaging With a Cross-Modal Stereo SystemabstractThis letter presents a novel approach for simultaneous depth and spectral imaging with a cross-modal stereo system. Two images of the target scene are captured at the same time: one compressively sampled hyperspectral measurement and one panchromatic measurement. The underlying hyperspectral cube is first reconstructed by leveraging the compressive sensing theory, during which a self-adaptive dictionary is learned from the panchromatic measurement to facilitate the reconstruction. The depth information of the scene is then recovered by estimating a disparity map between the hyperspectral cube and the panchromatic measurement through stereo matching. This disparity map, once obtained, is used to align the hyperspectral and panchromatic measurements to boost the hyperspectral reconstruction in an iterative manner. Through hardware experiments, for the first time to our knowledge, we demonstrate a snapshot system that allows for simultaneous depth and spectral imaging. The proposed system is capable of recording depth and spectral videos of dynamic scenes. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Snapshot Hyperspectral Light Field ImagingabstractThis paper presents the first snapshot hyperspectral light field imager in practice. Specifically, we design a novel hybrid camera system to obtain two complementary measurements that sample the angular and spectral dimensions respectively. To recover the full 5D hyperspectral light field from the severely undersampled measurements, we then propose an efficient computational reconstruction algorithm by exploiting the large correlations across the angular and spectral dimensions through self-learned dictionaries. Simulation on an elaborate hyperspectral light field dataset validates the effectiveness of the proposed approach. Hardware experimental results demonstrate that, for the first time to our knowledge, a 5D hyperspectral light field containing 9x9 angular views and 27 spectral bands can be acquired in a single shot. Zhiwei Xiong, Lizhi Wang 0001, Huiqun Li, Dong Liu 0002, Feng Wu 0001 |
CVPR | 2 |
| 2017 | Adaptive Nonlocal Sparse Representation for Dual-Camera Compressive Hyperspectral ImagingabstractLeveraging the compressive sensing (CS) theory, coded aperture snapshot spectral imaging (CASSI) provides an efficient solution to recover 3D hyperspectral data from a 2D measurement. The dual-camera design of CASSI, by adding an uncoded panchromatic measurement, enhances the reconstruction fidelity while maintaining the snapshot advantage. In this paper, we propose an adaptive nonlocal sparse representation (ANSR) model to boost the performance of dual-camera compressive hyperspectral imaging (DCCHI). Specifically, the CS reconstruction problem is formulated as a 3D cube based sparse representation to make full use of the nonlocal similarity in both the spatial and spectral domains. Our key observation is that, the panchromatic image, besides playing the role of direct measurement, can be further exploited to help the nonlocal similarity estimation. Therefore, we design a joint similarity metric by adaptively combining the internal similarity within the reconstructed hyperspectral image and the external similarity within the panchromatic image. In this way, the fidelity of CS reconstruction is greatly enhanced. Both simulation and hardware experimental results show significant improvement of the proposed method over the state-of-the-art. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Compressive hyperspectral imaging with complementary RGB measurementsabstractCoded aperture snapshot spectral imaging (CASSI) has been demonstrated as a feasible solution to recover a 3D hyperspectral image by using a single 2D measurement. In this paper, we propose a new hybrid camera design for CASSI to capture high quality hyperspectral images while maintaining the snapshot advantage. Specifically, we employ a complementary RGB camera in conjunction with the CASSI system. The recorded RGB image can provide reliable spectral clue of the scene. By combining the coded hyperspectral information from the CASSI branch and the uncoded color information from the RGB branch, hyperspectral images can be reconstructed with high fidelity. Furthermore, by conducting demosaicing on the raw RGB image as a preprocessing procedure, even better performance can be achieved. Both theoretical analysis and simulation results show improved accuracy of the proposed method compared to the state-of-the-arts. Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
VCIP | 1 |
| 2015 | High-speed hyperspectral video acquisition with a dual-camera architectureabstractWe propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hyperspectral video at a low frame rate, which jointly provide reliable projections for the underlying HSHS video. Second, we exploit the panchromatic video to learn an over-complete 3D dictionary to represent each band-wise video sparsely, and a robust computational reconstruction is then employed to recover the HSHS video based on the joint videos and the self-learned dictionary. Experimental results demonstrate that, for the first time to our knowledge, the hyperspectral video frame rate reaches up to 100fps with decent quality, even when the incident light is not strong. Lizhi Wang 0001, Zhiwei Xiong, Dahua Gao, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001 |
CVPR | 1 |