VLDB 2026 Research / reviewers in the wild / expert
Xin Yuan 0002
dblp:78/713-2
· DBLP profile ↗
144ranked-venue papers
14as first author
90since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 12 first-author · 55 since 2021Artificial intelligence and machine learning · 69 · 4 first-author · 55 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCIGaussian-D: Dynamic Scene Reconstruction from a Single Snapshot Compressive ImageabstractIn this paper, we explore the potential of snapshot compressive imaging (SCI) for dynamic 3D scene reconstruction from a single temporal compressed image. SCI is a low-cost imaging technique that captures high-dimensional information-such as temporal data-using$2 D$sensors and coded masks, significantly reducing data bandwidth while offering inherent privacy advantages. While recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have enabled 3D reconstruction from SCI measurements, these methods are fundamentally limited to static scenes and fail to generalize to dynamic content. To address this, we propose SCIGaussian-D, a novel framework that enables dynamic 3D reconstruction from a single SCI image. Our method represents the scene with 3D Gaussians defined in a canonical space and models motion using learnable deformation fields. By incorporating the SCI imaging model into the training loop, SCIGaussian-D directly reconstructs the dynamic 3D scene and recovers the corresponding camera motion from a single SCI. We evaluate our method on both synthetic and real SCI datasets, demonstrating significant improvements in reconstruction quality over existing baselines. Our results establish a new state of the art for dynamic scene reconstruction within the SCI framework, paving the way for practical applications in high-speed imaging and real-time scene rendering. Yuze Yang, Xin Yuan 0002, Peidong Liu 0001 |
3DV | 6 |
| 2026 | Breaking Measurement Barriers: From Compressed Sensing to Deep ReconstructionabstractDeep learning methods have achieved remarkable success in image compressed sensing (CS) task, namely reconstructing a high-fidelity image from its compressed measurement. However, existing methods are deficient in incoherent compressed measurement at sensing phase and implicit measurement representations at reconstruction phase, limiting the overall performance. In this work, we answer two questions: (i) how to improve the measurement incoherence for decreasing the ill-posedness; (ii) how to learn informative representations from measurements. To this end, we propose a novel asymmetric Kronecker CS (AKCS) model and theoretically present its better incoherence than previous Kronecker CS with minimal increase of complexity. Moreover, apart from the explicit measurement representations in gradient descent projection in unfolding networks, we further propose a measurement-aware cross attention (MACA) mechanism to learn implicit measurement representations. We integrate AKCS and MACA into a widely-used unfolding architecture to get a measurement-enhanced unfolding network (MEUNet). Extensive experiments demonstrate that the proposed MEUNet achieves state-of-the-art (SOTA) performance in reconstruction accuracy with high efficiency. Gang Qu 0005, Ping Wang 0029, Siming Zheng, Xin Yuan 0002 |
AAAI | 4 |
| 2026 | Realism Control One-step Diffusion for Real-world Image Super ResolutionabstractPre-trained diffusion models have shown great potential in real-world image super-resolution (Real-ISR) tasks by enabling high-resolution reconstructions. While one-step diffusion (OSD) methods significantly improve efficiency compared to traditional multi-step approaches, they still have limitations in balancing fidelity and realism across diverse scenarios. Since the OSDs for SR are usually trained or distilled by a single timestep, they lack flexible control mechanisms to adaptively prioritize these competing objectives, which are inherently manageable in multi-step methods through adjusting sampling steps. To address this challenge, we propose a Realism Controlled One-step Diffusion (RCOD) framework for Real-ISR. RCOD provides a latent domain grouping strategy that enables explicit control over fidelity-realism trade-offs during the noise prediction phase with minimal training paradigm modifications and original training data. A degradation-aware sampling strategy is also introduced to align distillation regularization with the grouping strategy and enhance the controlling of trade-offs. Moreover, a visual prompt injection module is used to replace conventional text prompts with degradation-aware visual tokens, enhancing both restoration accuracy and semantic consistency. Our method achieves superior fidelity and perceptual quality while maintaining computational efficiency. Extensive experiments demonstrate that RCOD outperforms state-of-the-art OSD methods in both quantitative metrics and visual qualities, with flexible realism control capabilities in the inference stage. Zongliang Wu, Siming Zheng, Peng-Tao Jiang, Xin Yuan 0002 |
AAAI | 4 |
| 2026 | High-Speed FHD Full-Color Video Computer-Generated HolographyabstractComputer-generated holography (CGH) is a promising technology for next-generation displays. However, generating high-speed, high-quality holographic video requires both high frame rate display and efficient computation, but is constrained by two key limitations: (i) Learning-based models often produce over-smoothed phases with narrow angular spectra, causing severe color crosstalk in high frame rate full-color displays such as depth-division multiplexing and thus resulting in a trade-off between frame rate and color fidelity. (ii) Existing frame-by-frame optimization methods typically optimize frames independently, neglecting spatial-temporal correlations between consecutive frames and leading to computationally inefficient solutions. To overcome these challenges, in this paper, we propose a novel high-speed full-color video CGH generation scheme. First, we introduce Spectrum-Guided Depth Division Multiplexing (SGDDM), which optimizes phase distributions via frequency modulation, enabling high-fidelity full-color display at high frame rates. Second, we present HoloMamba, a lightweight asymmetric Mamba-Unet architecture that explicitly models spatial-temporal correlations across video sequences to enhance reconstruction quality and computational efficiency. Extensive simulated and real-world experiments demonstrate that SGDDM achieves high-fidelity full-color display without compromise in frame rate, while HoloMamba generates FHD (1080p) full-color holographic video at over 260 FPS, more than 2.6 times faster than the prior state-of-the-art Divide-Conquer-and-Merge Strategy. Haomiao Zhang, Yanling Piao, Zhangyuan Li, Ping Wang 0029, Xin Yuan 0002 |
AAAI | 9 |
| 2026 | Degradation learning adaptive deep unfolding network for spectral compressive imaging
Lei Liu 0067, Xin Yuan 0002, Haifeng Zheng |
Pattern Recognit. | 3 |
| 2026 | 2D-Slice and 3D-Cube Mamba Network for Snapshot Spectral Compressive ImagingabstractHyperspectral image (HSI) reconstruction algorithms are fundamental to coded aperture snapshot spectral imaging (CASSI) systems. Recently, deep unfolding networks (DUNs) have emerged as a dominant solution, seamlessly combining traditional optimization frameworks with the strengths of deep learning. Among these, Mamba stands out as a prominent method for modeling long-range dependencies. However, its reliance on one-dimensional (1D) spatial scanning often compromises spectral consistency and spatial coherence, leading to misalignment of neighboring pixels within sequences. To address these limitations, we propose a novel multi-view framework based on 2D-slice modeling, which ensures spatial-spectral continuity in 1D sequences while maintaining computational efficiency. Furthermore, motivated by the need for precise local patch modeling in 2D images, we develop a 3D-cube Mamba model for HSI reconstruction. By integrating the UNet architecture, this model enhances spatial and spectral detail representation through multi-scale receptive field modeling, using fixed cube sizes to dynamically adjust pixel distances. These advancements are incorporated into the A-HQS-accelerated deep unfolding framework, synergistically combining the strengths of 2D-slice and 3D-cube MambaNet to achieve state-of-the-art HSI reconstruction performance. Experimental evaluations on simulated and real-world CASSI datasets demonstrate the efficacy of the proposed approach, achieving superior spectral fidelity and detailed feature representation. The source code is available at: https://github.com/fengyuchao97/SCM-DUN. Yuchao Feng, Zongliang Wu, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | MePAT: Meta-Prior Aided Transformer for Adverse Weather Condition RestorationabstractImage restoration under adverse weather conditions is critical for real-world applications. However, existing approaches mainly suffer from two fundamental limitations,i) the impractical requirement of prior degradation knowledge for task-specific model selection andii) performance degradation when handling with in-the-wild corruptions. To address the above issues, in this paper, we propose a novelMeta-prior Aided Transformerrestoration framework, MePAT, to synergize dynamic feature modulation with optimal transport (OT) theory. Specifically, we first architect an efficient attention mechanism,rectified self-channel attention(RSCA) to capture long-range associations along the channel dimension. Then, to adaptively tackle different conditions, we design atask-shared prior learning network(TPLN) to generate content-adaptive weather embeddings and serve as feature modulators to direct a more flexible and robust restoration process. In addition to learn discriminative task features, we propose an weakly-supervised OT-driven contrastive loss to measure the discrepancy between different weather corruptions. During the inference process, through the shared TPLN, we derive image-oriented vectors for unseen corruptions and then perform image restoration. The superior experimental results on three synthetic benchmarks demonstrate the effectiveness of MePAT. We also conduct experiments on real-world applications to verify the generalization ability and robustness. The code and pre-trained models will be made available. Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002, Chunhui Qu, Hongwei Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Sparse Transformer for Ultra-Sparse Sampled Video Compressive SensingabstractDigital cameras consume$\sim 0.1$microjoule per pixel to capture and encode video, resulting in a power usage of$\sim 20$W for a 4K sensor operating at 30 fps. Imagining gigapixel cameras operating at 100-1000 fps, the current processing model is unsustainable. To address this, physical layer compressive measurement has been proposed to reduce power consumption per pixel by 10-100×. Video Snapshot Compressive Imaging (SCI) introduces high frequency modulation in the optical sensor layer to increase effective frame rate. A commonly used sampling strategy of video SCI is Random Sampling (RS) where each mask element value is randomly set to be 0 or 1. Similarly, image inpainting (I2P) has demonstrated that images can be recovered from a fraction of the image pixels. Inspired by I2P, we propose Ultra-Sparse Sampling (USS) regime, where at each spatial location, only one sub-frame is set to 1 and all others are set to 0. We then build a Digital Micro-mirror Device (DMD) encoding system to verify the effectiveness of our USS strategy. Ideally, we can decompose the USS measurement into sub-measurements for which we can utilize I2P algorithms to recover high-speed frames. However, due to the mismatch between the DMD and CCD, the USS measurement cannot be perfectly decomposed. To this end, we proposeBSTFormer, a sparse TransFormer that utilizes local Block attention, global Sparse attention, and global Temporal attention to exploit the sparsity of the USS measurement. Extensive results on both simulated and real-world data show that our method significantly outperforms all previous state-of-the-art algorithms. Additionally, an essential advantage of the USS strategy is its higher dynamic range than that of the RS strategy. Finally, from the application perspective, the USS strategy is a good choice to implement a complete video SCI system on chip due to its fixed exposure time. Code is available athttps://github.com/mcao92/BSTFormer. Siming Zheng, Lishun Wang, David J. Brady, Xin Yuan 0002 |
IEEE Trans. Multim. | 6 |
| 2025 | Detail Matters: Mamba-Inspired Joint Unfolding Network for Snapshot Spectral Compressive ImagingabstractIn the coded aperture snapshot spectral imaging system, Deep Unfolding Networks (DUNs) have made impressive progress in recovering 3D hyperspectral images (HSIs) from a single 2D measurement. However, the inherent nonlinear and ill-posed characteristics of HSI reconstruction still pose challenges to existing methods in terms of accuracy and stability. To address this issue, we propose a Mamba-inspired Joint Unfolding Network (MiJUN), which integrates physics-embedded DUNs with learning-based HSI imaging. Firstly, leveraging the concept of trapezoid discretization to expand the representation space of unfolding networks, we introduce an accelerated unfolding network scheme. This approach can be interpreted as a generalized accelerated half-quadratic splitting with a second-order differential equation, which reduces the reliance on initial optimization stages and addresses challenges related to long-range interactions. Crucially, within the Mamba framework, we restructure the Mamba-inspired global-to-local attention mechanism by incorporating a selective state space model and an attention mechanism. This effectively reinterprets Mamba as a variant of the Transformer architecture, improving its adaptability and efficiency. Furthermore, we refine the scanning strategy with Mamba by integrating the tensor mode-k unfolding into the Mamba network. This approach emphasizes the low-rank properties of tensors along various modes, while conveniently facilitating 12 scanning directions. Numerical and visual comparisons on both simulation and real datasets demonstrate the superiority of our proposed MiJUN, and achieving overwhelming detail representation. Yuchao Feng, Zongliang Wu, Yulun Zhang 0001, Xin Yuan 0002 |
AAAI | 5 |
| 2025 | Prior-guided Hierarchical Harmonization Network for Efficient Image DehazingabstractImage dehazing is a crucial task that involves the enhancement of degraded images to recover their sharpness and textures. While vision Transformers have exhibited impressive results in diverse dehazing tasks, their quadratic complexity and lack of dehazing priors pose significant drawbacks for real-world applications. In this paper, guided by triple priors, Bright Channel Prior (BCP), Dark Channel Prior (DCP), and Histogram Equalization (HE), we propose a Prior-guided Hierarchical Harmonization Network (PGHHNet) for image dehazing. PGHNet is built upon the UNet-like architecture with an efficient encoder and decoder, consisting of two module types: (1) Prior aggregation module that injects BCP/DCP and selects diverse contexts with gating attention. (2) Feature harmonization modules that subtract low-frequency components from spatial and channel aspects and learn more informative feature distributions to equalize the feature maps. Inspired by observing the sparsity of BCP/DCP and the histogram equalization, we harmonize the deep features using a histogram equation-guided module and further leverage BCP/DCP to guide spatial attention through a sandwich module as the bottleneck. Comprehensive experiments demonstrate that our model efficiently attains the highest level of performance among existing methods across four different datasets for image dehazing tasks. Xiongfei Su, Yuning Cui 0001, Yulun Zhang 0001, Zheng Chen 0014, Zongliang Wu, Zedong Wang, Yuanlong Zhang, Xin Yuan 0002 |
AAAI | 10 |
| 2025 | Dual-branch Graph Feature Learning for NLOS ImagingabstractThe domain of non-line-of-sight (NLOS) imaging is advancing rapidly, offering the capability to reveal occluded scenes that are not directly visible. However, contemporary NLOS systems face several significant challenges: (1) The computational and storage requirements are profound due to the inherent three-dimensional grid data structure, which restricts practical application. (2) The simultaneous reconstruction of albedo and depth information requires a delicate balance using hyperparameters in the loss function, rendering the concurrent reconstruction of texture and depth information difficult. This paper introduces the innovative methodology, DG-NLOS, which integrates an albedo-focused reconstruction branch dedicated to albedo information recovery and a depth-focused reconstruction branch that extracts geometrical structure, to overcome these obstacles. The dual-branch framework segregates content delivery to the respective reconstructions, thereby enhancing the quality of the retrieved data. To our knowledge, we are the first to employ the GNN as a fundamental component to transform dense NLOS grid data into sparse structural features for efficient reconstruction. Comprehensive experiments demonstrate that our method attains the highest level of performance among existing methods across synthetic and real data. Xiongfei Su, Lina Liu 0010, Zheng Chen 0014, Yulun Zhang 0001, Juntian Ye, Feihu Xu, Xin Yuan 0002 |
AAAI | 9 |
| 2025 | Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging
Ping Wang 0029, Lishun Wang, Gang Qu 0005, Xiaodong Wang 0026, Yulun Zhang 0001, Xin Yuan 0002 |
CVPR | 6 |
| 2025 | SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive ImageabstractSnapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality limitations are general issues that have limited wider adoption. This paper introduces SCI-Gaussian, the first 3D-aware SCI reconstruction based on 3D Gaussian splatting (3D-GS). This method utilizes an explicit 3D representation to achieve efficient and high-quality scene reconstruction. The motivation stems from the highly efficient representation and surprising quality of 3D-GS, despite when applied to SCI system, it encounters difficulties in generating point initialization for explicit Gaussians and accurate pose recovery from a single SCI measured image. Specifically, we effectively initialize these Gaussians through sampling a coarsely trained NeRF at various hash structures, then model the physical formation of the SCI measurement and jointly optimize Gaussians and camera trajectories with a bundle adjustment formulation during exposure time. Extensive experiments on synthetic and real-world datasets demonstrate that SCI-Gaussian outperforms the state-of-the-art (SOTA) methods, achieving comparable or better results with significantly 10× faster training and 1000× faster rendering speed than the most recent NeRF-based method. Xiaodong Wang 0026, Xin Yuan 0002, Mark D. Butala, Gaoang Wang |
ICASSP | 4 |
| 2025 | Improving the Performance of Compressive Spectral Imaging with Bayer Color Filter ArrayabstractReconstructing hyperspectral images (HSIs) from coded measurements from coded aperture snapshot spectral imaging (CASSI) system is essential for capturing images that offer superior spectral resolution over traditional RGB images. Nevertheless, the reconstruction of HSIs is difficult because of the aliasing of spatial and spectral information in coded measurements. To address this issue, we explore the possibility of utilizing an RGB camera with a Bayer color filter array (CFA) as an alternative to the gray-scale camera for implementing CASSI, which we term Bayer-CASSI. We first formulate the mathematical model of Bayer-CASSI, and then we evaluate the performance of Bayer-CASSI in simulation data and real data captured by our self-built Bayer-CASSI optical system. Experiment results in both simulation and real data show the effectiveness of Bayer-CASSI. Zijun He, Ziyi Meng 0001, Xin Yuan 0002 |
ICIP | 4 |
| 2025 | Motion-Aware Reconstruction for Video Snapshot Compressive ImagingabstractVideo Snapshot Compressive Imaging (SCI) provides an elegant solution for recording fast-dynamics motion through optical compression and computational reconstruction. However, existing video reconstruction algorithms are computationally intensive and typically require substantial computing and storage resources to large areas of static background with minimal information, leading to significant resource waste. To this end, this paper proposes a motion-aware reconstruction paradigm, in which moving objects are separated from the static background to identify Regions Of Interest (ROI) in the measurement domain and then video reconstruction is confined to ROI to save reconstruction costs. Specifically, a Motion Decomposition Network (MoDeNet) is proposed to get ROI by deep unfolding of Half-Quadratic Splitting (HQS). Subsequently, high-performance reconstruction network is used to recover ROI videos. Extensive experiments demonstrate that the proposed paradigm reduces the reconstruction costs of video SCI effectively and efficiently, and the proposed MoDeNet outperforms previous decomposition algorithms in terms of accuracy and speed. Zhangyuan Li, Ping Wang 0029, Haomiao Zhang, Xin Yuan 0002 |
ICIP | 5 |
| 2025 | Texture-aware Intrinsic Image Decomposition with Model- and Learning-based PriorsabstractThis paper aims to recover the intrinsic reflectance layer and shading layer given a single image. Though this intrinsic image decomposition problem has been studied for decades, it remains a significant challenge in cases of complex scenes, i.e. spatially-varying lighting effect and rich textures. In this paper, we propose a novel method for handling severe lighting and rich textures in intrinsic image decomposition, which enables to produce high-quality intrinsic images for real-world images. Specifically, we observe that previous learning-based methods tend to produce texture-less and over-smoothing intrinsic images, which can be used to infer the lighting and texture information given a RGB image. In this way, we design a texture-guided regularization term and formulate the decomposition problem into an optimization framework, to separate the material textures and lighting effect. We demonstrate that combining the novel texture-aware prior can produce superior results to existing approaches. Code is available at https://github.com/xiaodongwo/Efficient IID. Xiaodong Wang 0026, Zijun He, Xin Yuan 0002 |
ICME | 3 |
| 2025 | Unfolding Framework with Complex-Valued Deformable Attention for High-Quality Computer-Generated Hologram GenerationabstractComputer-generated holography (CGH) has gained wide attention with deep learning-based algorithms. However, due to its nonlinear and ill-posed nature, challenges remain in achieving accurate and stable reconstruction. Specifically, (i) the widely used end-to-end networks treat the reconstruction model as a black box, ignoring underlying physical relationships, which reduces interpretability and flexibility. (ii) CNN-based CGH algorithms have limited receptive fields, hindering their ability to capture long-range dependencies and global context. (iii) Angular spectrum method (ASM)-based models are constrained to finite near-fields. In this paper, we propose a Deep Unfolding Network (DUN) that decomposes gradient descent into two modules: an adaptive bandwidth-preserving model (ABPM) and a phase-domain complex-valued denoiser (PCD), providing more flexibility. ABPM allows for wider working distances compared to ASM-based methods. At the same time, PCD leverages its complex-valued deformable self-attention module to capture global features and enhance performance, achieving a PSNR over 35 dB. Experiments on simulated and real data show state-of-the-art results. Code is available at https://github.com/HannahZhang1926/Complex-Valued-Deformable-Transformer-for-CGH. Haomiao Zhang, Zhangyuan Li, Yanling Piao, Xiaodong Wang 0026, Xiongfei Su, Xin Yuan 0002 |
ICME | 9 |
| 2025 | KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color EditingabstractRecent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundaries, difficulties in maintaining multi-view consistency, and over-reliance on annotated data. To address these limitations, in this paper, we propose a novel weakly-supervised method called KaRF for local color editing, which facilitates high-fidelity and realistic appearance edits in arbitrary regions of 3D scenes. At the core of the proposed KaRF approach is a unified two-stage Kolmogorov-Arnold Networks (KANs)-based radiance fields framework, comprising a segmentation stage followed by a local recoloring stage. This architecture seamlessly integrates geometric priors from NeRF to achieve weakly-supervised learning, leading to superior performance. More specifically, we propose a residual adaptive gating KAN structure, which integrates KAN with residual connections, adaptive parameters, and gating mechanisms to effectively enhance segmentation accuracy and refine specific editing effects. Additionally, we propose a palette-adaptive reconstruction loss, which can enhance the accuracy of additive mixing results. Extensive experiments demonstrate that the proposed KaRF algorithm significantly outperforms many state-of-the-art methods both qualitatively and quantitatively. Our code and more results are available at: https://github.com/PaiDii/KARF.git. Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Zipei Fan, Ce Zhu |
NeurIPS | 5 |
| 2025 | DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-ResolutionabstractDiffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training overhead on video data and stringent fidelity demands. To tackle the above issues, we propose DOVE, an efficient one-step diffusion model for real-world VSR. DOVE is obtained by fine-tuning a pretrained video diffusion model (*i.e.*, CogVideoX). To effectively train DOVE, we introduce the latent-pixel training strategy. The strategy employs a two-stage scheme to gradually adapt the model to the video super-resolution task. Meanwhile, we design a video processing pipeline to construct a high-quality dataset tailored for VSR, termed HQ-VSR. Fine-tuning on this dataset further enhances the restoration capability of DOVE. Extensive experiments show that DOVE exhibits comparable or superior performance to multi-step diffusion-based VSR methods. It also offers outstanding inference efficiency, achieving up to a **28$\times$** speed-up over existing methods such as MGLD-VSR. Code is available at: https://github.com/zhengchen1999/DOVE. Zheng Chen 0014, Zichen Zou, Xiongfei Su, Xin Yuan 0002, Yulun Zhang 0001 |
NeurIPS | 5 |
| 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion DiscriminatorabstractDiffusion models have demonstrated excellent performance for real-world image super-resolution (Real-ISR), albeit at high computational costs. Most existing methods are trying to derive one-step diffusion models from multi-step counterparts through knowledge distillation (KD) or variational score distillation (VSD). However, these methods are limited by the capabilities of the teacher model, especially if the teacher model itself is not sufficiently strong. To tackle these issues, we propose a new One-Step \textbf{D}iffusion model with a larger-scale \textbf{D}iffusion \textbf{D}iscriminator for SR, called D$^3$SR. Our discriminator is able to distill noisy features from any time step of diffusion models in the latent space. In this way, our diffusion discriminator breaks through the potential limitations imposed by the presence of a teacher model. Additionally, we improve the perceptual loss with edge-aware DISTS (EA-DISTS) to enhance the model's ability to generate fine details. Our experiments demonstrate that, compared with previous diffusion-based methods requiring dozens or even hundreds of steps, our D$^3$SR attains comparable or even superior results in both quantitative metrics and qualitative evaluations. Moreover, compared with other methods, D$^3$SR achieves at least $3\times$ faster inference speed and reduces parameters by at least 30\%. Jianze Li, Jiezhang Cao, Zichen Zou, Xiongfei Su, Xin Yuan 0002, Yulun Zhang 0001, Xiaokang Yang 0001 |
NeurIPS | 5 |
| 2025 | Spectral Compressive Imaging via Chromaticity-Intensity DecompositionabstractIn coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it difficult to recover the intrinsic spectral reflectance that remains invariant to lighting conditions. To address these challenges, we propose a chromaticity-intensity decomposition framework, which disentangles an HSI into a spatially smooth intensity map and a spectrally variant chromaticity cube. The chromaticity encodes lighting-invariant reflectance, enriched with high-frequency spatial details and local spectral sparsity. Building on this decomposition, we develop CIDNet—a Chromaticity-Intensity Decomposition unfolding network within a dual-camera CASSI system. CIDNet integrates a hybrid spatial-spectral Transformer tailored to reconstruct fine-grained and sparse spectral chromaticity and a degradation-aware, spatially-adaptive noise estimation module that captures anisotropic noise across iterative stages. Extensive experiments on both synthetic and real-world CASSI datasets demonstrate that our method achieves superior performance in both spectral and chromaticity fidelity. Code is released at: \url{https://github.com/xiaodongwo/CIDNet}. Xiaodong Wang 0026, Zijun He, Ping Wang 0029, Lishun Wang, Xin Yuan 0002 |
NeurIPS | 6 |
| 2025 | Self-supervised Learning with Spectral Low-Rank Prior for Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction from coded measurement is significant for acquiring images with higher spectral resolution than traditional RGB images. Current advanced neural networks have already shown impressive performance in some datasets like CAVE and KAIST. However, these networks rely on a large amount of simulated ground truth, measurement pairs. Unfortunately, in some scenarios, it is hard to obtain a sufficient high-quality HSI training set, resulting in low generalization ability. Although iterative algorithms show good generalization ability, they are limited by slow speed and low reconstruction quality. To address this challenge, in this paper, we propose a self-supervised learning framework, which can train and fine-tune networks using measurements without ground truth. Besides, we propose the spectral low-rank loss function that enables networks to learn the signal model of HSI. Finally, we train and fine-tune a representative deep unfolding network, GAP-net, using our proposed framework. Extensive simulation and real data results show that the proposed self-supervised framework is capable of achieving results competitive with those of supervised networks. Code is available at https://github.com/zjhe02/CASSI-SSL. Zijun He, Lishun Wang, Ziyi Meng 0001, Xin Yuan 0002 |
WACV | 4 |
| 2025 | X-clustering beyond contextual representations
Tianyi Huang, Zhengjun Zhang, Xin Yuan 0002, Stan Z. Li, Naixue Xiong, Shenghui Cheng |
Inf. Sci. | 3 |
| 2025 | $S^{2}$S2-Transformer for Mask-Aware Hyperspectral Image ReconstructionabstractSnapshot compressive imaging (SCI) surges as a novel way of capturing hyperspectral images. It operates an optical encoder to compress the 3D data into a 2D measurement and adopts a software decoder for the signal reconstruction. Recently, a representative SCI set-up of coded aperture snapshot compressive imager (CASSI) with Transformer reconstruction backend remarks high-fidelity sensing performance. However, dominant spatial and spectral attention designs show limitations in hyperspectral modeling. The spatial attention values describe the inter-pixel correlation but overlook the across-spectra variation within each pixel. The spectral attention size is unscalable to the token spatial size and thus bottlenecks information allocation. Besides, CASSI entangles the spatial and spectral information into a 2D measurement, placing a barrier for information disentanglement and modeling. In addition, CASSI blocks the light with a physical binary mask, yielding the masked data loss. To tackle above challenges, we propose a spatial-spectral ($S^{2}$S2-) Transformer implemented by a paralleled attention design and a mask-aware learning strategy. First, we systematically explore pros and cons of different spatial (-spectral) attention designs, based on which we find performing both attentions in parallel well disentangles and models the blended information. Second, the masked pixels induce higher prediction difficulty and should be treated differently from unmasked ones. We adaptively prioritize the loss penalty attributing to the mask structure by referring to the mask-encoded prediction as an uncertainty estimator. We theoretically discuss the distinct convergence tendencies between masked/unmasked regions of the proposed learning strategy. Extensive experiments demonstrate that on average, the results of the proposed method are superior over the state-of-the-art methods. We empirically visualize and reason the behaviour of spatial and spectral attentions, and comprehensively examine the impact of the mask-aware learning, both of which advances the physics-driven deep network design for the reconstruction with CASSI. Jiamian Wang, Yulun Zhang 0001, Xin Yuan 0002, Zhiqiang Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Rethinking Semantic-Level Building Change Detection: Ensemble Learning and Dynamic InteractionabstractBuilding change detection (BCD) of multi-temporal images plays a significant role in urban expansion and area internal change analysis. However, current BCD methods remain stagnant at binary-level predictions due to the scarcity of detectable changes and the imbalance between new constructions and demolitions. To advance semantic-level BCD, we propose a dynamic interaction ensemble learning network (DIELNet) using a collaborative training paradigm across multiple datasets and tasks. Firstly, we create a simulated BCD dataset, Inria-CD, derived from the building segmentation dataset. It features complex structures, large scale, and balanced ratios with both binary- and semantic-level labels. Importantly, we shift the traditional single-dataset and single-task BCD learning paradigm by introducing ensemble learning. This mechanism feeds multiple datasets into the model to obtain binary- and semantic-level predictions through a single training process, accommodating partial samples without semantic-level labels. In addition, our DIELNet incorporates bitemporal dynamic interactions during data processing and feature extraction. The former generates progressive sequences by swapping mutual high-frequency components during the Fourier transformation, while the latter is achieved through Mamba-structure modules, which integrate local convolution with dynamic-static kernels and long-range dependencies via state space models. Numerical and visual comparisons demonstrate the superiority of DIELNet. Moreover, existing algorithms can also benefit significantly from our ensemble learning approach. Datasets and codes are available at: https://github.com/fengyuchao97/DIELNet. Yuchao Feng, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | LCTC: Lightweight Convolutional Thresholding Sparse Coding Network Prior for Compressive Hyperspectral ImagingabstractCompressive spectral imaging has garnered significant attention for its ability to effectively enhance the captured spatial and spectral information. Predominant methods, based on compressive sensing, typically formulate the imaging task as a constrained optimization problem and rely on hand-crafted priors to model the sparsity of spectral images. However, these approaches often suffer from suboptimal performance due to the inherent difficulty of identifying an appropriate transform space where spectral images exhibit sparsity. To overcome this limitation, we propose a novel convolutional sparse coding-inspired untrained network prior for fast and adaptive identification of the sparse transform domain and compressible signal. Specifically, a Lightweight Convolutional Thresholding sparse Coding (LCTC) network is designed as the sparse transform domain, with its inputs interpreted as sparse coefficients. Crucially, both the transform domain and its coefficients are solved in a self-supervised learning manner. Furthermore, we demonstrate that LCTC prior can be seamlessly incorporated into the iterative optimization algorithm as a Plug-and-Play (PnP) regularization. Both the LCTC and PnP-LCTC exhibit superior performance compared to previous methods. Experiments under various scenarios validate the effectiveness and efficiency of our approach. Yurong Chen 0003, Yaonan Wang 0001, Xiaodong Wang 0026, Xin Yuan 0002, Hui Zhang 0023 |
IEEE Trans. Image Process. | 4 |
| 2025 | Texture-Consistent 3D Scene Style Transfer via Transformer-Guided Neural Radiance FieldsabstractRecent advancements have suggested that neural radiance fields (NeRFs) show great potential in 3D style transfer. However, most existing NeRF-based style transfer methods still face considerable challenges in generating stylized images that simultaneously preserve clear scene textures and maintain strong cross-view consistency. To address these limitations, in this paper, we propose a novel transformer-guided approach for 3D scene style transfer. Specifically, we first design a transformer-based style transfer network to capture long-range dependencies and generate 2D stylized images with initial consistency, which serve as supervision for the 3D stylized generation. To enable fine-grained control over style, we propose a latent style vector as a conditional feature and design a style network that projects this style information into the 3D space. We further develop a merge network that integrates style features with scene geometry to render 3D stylized images that are both visually coherent and stylistically consistent. In addition, we propose a texture consistency loss to preserve scene structure and enhance texture fidelity across views. Extensive quantitative and qualitative experimental results demonstrate that our proposed approach outperforms many state-of-the-art methods in terms of visual perception, image quality and multi-view consistency. Our code and more results are available at: https://github.com/PaiDii/TGTC-Style.git. Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 6 |
| 2025 | Restoration of Images Taken Through a Dirty Window Using Optics-Guided TransformerabstractTaking photographs through windows is an inevitable scenario in the real world, but glass windows are not ideally clean in most cases. Although there exists various raindrop removal methods, the occlusion of dirt, as another dirty window case, has not been well valued. The vital reasons include i) the limitation of the optical imaging model proposed in previous methods, and ii) the shortage of a practical dataset for sufficient types of dirty glass windows. To fill this research gap, in this paper, we first propose a general optical imaging model that fits widely used dirty window cases. Following this, training and testing synthetic datasets are generated, and real-world dirty window data are collected to evaluate the effectiveness of our imaging model and synthetic data. For the methodology part, we propose an optics-guided Transformer network to solve this special image restoration problem, i.e., the dirt removal for images taken through a dirty window. Experimental results demonstrate that our imaging model is effective and robust. Our proposed network leads to higher performance than existing methods on both synthetic and real-world dirty window images. Code and data are available at https://github.com/Zongliang-Wu/ReDNet. Zongliang Wu, Juzheng Zhang, Ying Fu 0001, Yulun Zhang 0001, Xin Yuan 0002 |
IEEE Trans. Image Process. | 5 |
| 2025 | Efficient High-Fidelity Global Low-Rank Optimization for Multispectral DemosaicingabstractThe nonlocal low-rank (NLR) optimization has shown promise for generalized multispectral filter array (MSFA) demosaicing. However, it faces challenges in balancing efficiency and accuracy. To tackle these challenges, we report here the multi-channel global low-rank optimization technique, achieving efficient high-fidelity MSFA demosaicing. Inspired by the cross-band correlations of natural multispectral images, we introduce the multi-channel matching and low-rank strategies that jointly optimize image patches of all channels, exhibiting higher efficiency and accuracy than existing approaches. Furthermore, we present global structural matching (GSM) which performs structure-aware multi-channel matching across the entire multispectral image. GSM extracts structurally important patches and efficiently searches their similar patches via parallel correlation, providing an order-of-magnitude improvement in efficiency. By combining the aforementioned techniques, we have achieved superior performance over the state-of-the-art NLR demosaicing technique, leading to up to 3.9 dB peak signal-to-noise ratio (PSNR) gain and over a 150-fold increase in computational speed. Experiments validated that the technique outperforms existing methods in reconstructing fine textures and details and exhibits superior robustness to noise. Daoyu Li, Xin Yuan 0002, Liheng Bian |
IEEE Trans. Multim. | 4 |
| 2025 | Advancing Hyperspectral and Multispectral Image Fusion: An Information-Aware Transformer-Based Unfolding NetworkabstractIn hyperspectral image (HSI) processing, the fusion of the high-resolution multispectral image (HR-MSI) and the low-resolution HSI (LR-HSI) on the same scene, known as MSI-HSI fusion, is a crucial step in obtaining the desired high-resolution HSI (HR-HSI). With the powerful representation ability, convolutional neural network (CNN)-based deep unfolding methods have demonstrated promising performances. However, limited receptive fields of CNN often lead to inaccurate long-range spatial features, and inherent input and output images for each stage in unfolding networks restrict the feature transmission, thus limiting the overall performance. To this end, we propose a novel and efficient information-aware transformer-based unfolding network (ITU-Net) to model the long-range dependencies and transfer more information across the stages. Specifically, we employ a customized transformer block to learn representations from both the spatial and frequency domains as well as avoid the quadratic complexity with respect to the input length. For spatial feature extractions, we develop an information transfer guided linearized attention (ITLA), which transmits high-throughput information between adjacent stages and extracts contextual features along the spatial dimension in linear complexity. Moreover, we introduce frequency domain learning in the feedforward network (FFN) to capture token variations of the image and narrow the frequency gap. Via integrating our proposed transformer blocks with the unfolding framework, our ITU-Net achieves state-of-the-art (SOTA) performance on both synthetic and real hyperspectral datasets. Bo Chen 0001, Ruiying Lu, Ziheng Cheng 0001, Chunhui Qu, Xin Yuan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | SCINeRF: Neural Radiance Fields from a Snapshot Compressive ImageabstractIn this paper, we explore the potential of Snapshot Compressive Imaging (SCI) technique for recovering the underlying 3D scene representation from a single temporal compressed image. SCI is a cost-effective method that enables the recording of high-dimensional data, such as hyperspectral or temporal information, into a single image using low- cost 2D imaging sensors. To achieve this, a series of specially designed 2D masks are usually employed, which not only reduces storage requirements but also offers potential privacy protection. Inspired by this, to take one step further, our approach builds upon the powerful 3D scene representation capabilities of neural radiance fields (NeRF). Specifically, we formulate the physical imaging process of SCI as part of the training of NeRF, allowing us to exploit its impressive performance in capturing complex scene structures. To assess the effectiveness of our method, we conduct extensive evaluations using both synthetic data and real data captured by our SCI system. Extensive experimental results demonstrate that our proposed approach sur- passes the state-of-the-art methods in terms of image re- construction and novel view image synthesis. Moreover, our method also exhibits the ability to restore high frame- rate multi-view consistent images by leveraging SCI and the rendering capabilities of NeRF. The code is available at https://github.com/WU-CVGL/SCINeRF. Xiaodong Wang 0026, Ping Wang 0029, Xin Yuan 0002, Peidong Liu 0001 |
CVPR | 4 |
| 2024 | Dual-Scale Transformer for Large-Scale Single-Pixel ImagingabstractSingle-pixel imaging (SPI) is a potential computational imaging technique which produces image by solving an ill-posed reconstruction problem from few measurements captured by a single-pixel detector. Deep learning has achieved impressive success on SPI reconstruction. However, previ-ous poor reconstruction performance and impractical imaging model limit its real-world applications. In this paper, we propose a deep unfolding network with hybrid-attention Transformer on Kronecker SPI model, dubbed HATNet, to im-prove the imaging quality of real SPI cameras. Specifically, we unfold the computation graph of the iterative shrinkage-thresholding algorithm (ISTA) into two alternative modules: efficient tensor gradient descent and hybrid-attention multi-scale denoising. By virtue of Kronecker SPI, the gradient descent module can avoid high computational overheads rooted in previous gradient descent modules based on vector-ized SPI. The denoising module is an encoder-decoder archi-tecture powered by dual-scale spatial attention for high- and low-frequency aggregation and channel attention for global information recalibration. Moreover, we build a SPI proto-type to verify the effectiveness of the proposed method. Ex-tensive experiments on synthetic and real data demonstrate that our method achieves the state-of-the-art performance. The source code and pre-trained models are available at https://github.com/Gang-Qu/HATNet-SPI. Gang Qu 0005, Ping Wang 0029, Xin Yuan 0002 |
CVPR | 3 |
| 2024 | Binarized Low-Light Raw Video EnhancementabstractRecently, deep neural networks have achieved excellent performance on low-light raw video enhancement. How-ever, they often come with high computational complexity and large memory costs, which hinder their applications on resource-limited devices. In this paper, we explore the feasibility of applying the extremely compact binary neural network (BNN) to low-light raw video enhancement. Nev-ertheless, there are two main issues with binarizing video enhancement models. One is how to fuse the temporal in-formation to improve low-light denoising without complex modules. The other is how to narrow the performance gap between binary convolutions with the full precision ones. To address the first issue, we introduce a spatial-temporal shift operation, which is easy-to-binarize and effective. The temporal shift efficiently aggregates the features of neigh-bor frames and the spatial shift handles the misalignment caused by the large motion in videos. For the second issue, we present a distribution-aware binary convolution, which captures the distribution characteristics of real-valued in-put and incorporates them into plain binary convolutions to alleviate the degradation in performance. Extensive quantitative and qualitative experiments have shown our high-efficiency binarized low-light raw video enhancement method can attain a promising performance. The code is available at https://github.com/ying-fuIBRVE. Gengchen Zhang, Yulun Zhang 0001, Xin Yuan 0002, Ying Fu 0001 |
CVPR | 3 |
| 2024 | A Simple Low-Bit Quantization Framework for Video Snapshot Compressive Imaging
Lishun Wang, Huan Wang 0014, Xin Yuan 0002 |
ECCV (52) | 4 |
| 2024 | Hierarchical Separable Video Transformer for Snapshot Compressive Imaging
Ping Wang 0029, Yulun Zhang 0001, Lishun Wang, Xin Yuan 0002 |
ECCV (81) | 4 |
| 2024 | Latent Diffusion Prior Enhanced Deep Unfolding for Snapshot Spectral Compressive Imaging
Zongliang Wu, Ruiying Lu, Ying Fu 0001, Xin Yuan 0002 |
ECCV (33) | 4 |
| 2024 | Coarse-Fine Spectral-Aware Deformable Convolution for Hyperspectral Image ReconstructionabstractWe study the inverse problem of Coded Aperture Snapshot Spectral Imaging (CASSI), which captures a spatial-spectral data cube using snapshot 2D measurements and uses algorithms to reconstruct 3D hyperspectral images (HSI). However, current methods based on Convolutional Neural Networks (CNNs) struggle to capture long-range dependencies and non-local similarities. The recently popular Transformerbased methods are poorly deployed on downstream tasks due to the high computational cost caused by self-attention. In this paper, we propose Coarse-Fine Spectral-Aware Deformable Convolution Network (CFSDCN), applying deformable convolutional networks (DCN) to this task for the first time. Considering the sparsity of HSI, we design a deformable convolution module that exploits its deformability to capture long-range dependencies and non-local similarities. In addition, we propose a new spectral information interaction module that considers both coarse-grained and fine-grained spectral similarities. Extensive experiments demonstrate that our CFSDCN significantly outperforms previous state-of-the-art (SOTA) methods on both simulated and real HSI datasets. Lishun Wang, Huan Wang 0014, Yinping Zhao, Xin Yuan 0002 |
ICIP | 6 |
| 2024 | Towards Real-time Video Compressive Sensing on Mobile Devices
Lishun Wang, Huan Wang 0014, Guoqing Wang 0001, Xin Yuan 0002 |
ACM Multimedia | 5 |
| 2024 | Binarized Diffusion Model for Image Super-ResolutionabstractAdvanced diffusion models (DMs) perform impressively in image super-resolution (SR), but the high memory and computational costs hinder their deployment. Binarization, an ultra-compression algorithm, offers the potential for effectively accelerating DMs. Nonetheless, due to the model structure and the multi-step iterative attribute of DMs, existing binarization methods result in significant performance degradation. In this paper, we introduce a novel binarized diffusion model, BI-DiffSR, for image SR. First, for the model structure, we design a UNet architecture optimized for binarization. We propose the consistent-pixel-downsample (CP-Down) and consistent-pixel-upsample (CP-Up) to maintain dimension consistent and facilitate the full-precision information transfer. Meanwhile, we design the channel-shuffle-fusion (CS-Fusion) to enhance feature fusion in skip connection. Second, for the activation difference across timestep, we design the timestep-aware redistribution (TaR) and activation function (TaA). The TaR and TaA dynamically adjust the distribution of activations based on different timesteps, improving the flexibility and representation alability of the binarized module. Comprehensive experiments demonstrate that our BI-DiffSR outperforms existing binarization methods. Code is released at: https://github.com/zhengchen1999/BI-DiffSR. Zheng Chen 0014, Haotong Qin, Xiongfei Su, Xin Yuan 0002, Linghe Kong, Yulun Zhang 0001 |
NeurIPS | 5 |
| 2024 | 2DQuant: Low-bit Post-Training Quantization for Image Super-ResolutionabstractLow-bit quantization has become widespread for compressing image super-resolution (SR) models for edge deployment, which allows advanced SR models to enjoy compact low-bit parameters and efficient integer/bitwise constructions for storage compression and inference acceleration, respectively. However, it is notorious that low-bit quantization degrades the accuracy of SR models compared to their full-precision (FP) counterparts. Despite several efforts to alleviate the degradation, the transformer-based SR model still suffers severe degradation due to its distinctive activation distribution. In this work, we present a dual-stage low-bit post-training quantization (PTQ) method for image super-resolution, namely 2DQuant, which achieves efficient and accurate SR under low-bit quantization. The proposed method first investigates the weight and activation and finds that the distribution is characterized by coexisting symmetry and asymmetry, long tails. Specifically, we propose Distribution-Oriented Bound Initialization (DOBI), using different searching strategies to search a coarse bound for quantizers. To obtain refined quantizer parameters, we further propose Distillation Quantization Calibration (DQC), which employs a distillation approach to make the quantized model learn from its FP counterpart. Through extensive experiments on different bits and scaling factors, the performance of DOBI can reach the state-of-the-art (SOTA) while after stage two, our method surpasses existing PTQ in both metrics and visual effects. 2DQuant gains an increase in PSNR as high as 4.52dB on Set5 (x2) compared with SOTA when quantized to 2-bit and enjoys a 3.60x compression ratio and 5.08x speedup ratio. The code and models are available at https://github.com/Kai-Liu001/2DQuant. Kai Liu 0034, Haotong Qin, Xin Yuan 0002, Linghe Kong, Guihai Chen, Yulun Zhang 0001 |
NeurIPS | 4 |
| 2024 | Cooperative Hardware-Prompt Learning for Snapshot Compressive ImagingabstractExisting reconstruction models in snapshot compressive imaging systems (SCI) are trained with a single well-calibrated hardware instance, making their perfor- mance vulnerable to hardware shifts and limited in adapting to multiple hardware configurations. To facilitate cross-hardware learning, previous efforts attempt to directly collect multi-hardware data and perform centralized training, which is impractical due to severe user data privacy concerns and hardware heterogeneity across different platforms/institutions. In this study, we explicitly consider data privacy and heterogeneity in cooperatively optimizing SCI systems by proposing a Federated Hardware-Prompt learning (FedHP) framework. Rather than mitigating the client drift by rectifying the gradients, which only takes effect on the learning manifold but fails to solve the heterogeneity rooted in the input data space, FedHP learns a hardware-conditioned prompter to align inconsistent data distribution across clients, serving as an indicator of the data inconsistency among different hardware (e.g., coded apertures). Extensive experimental results demonstrate that the proposed FedHP coordinates the pre-trained model to multiple hardware con- figurations, outperforming prevalent FL frameworks for 0.35dB under challenging heterogeneous settings. Moreover, a Snapshot Spectral Heterogeneous Dataset has been built upon multiple practical SCI systems. Data and code are aveilable at https://github.com/Jiamian-Wang/FedHP-Snapshot-Compressive-Imaging.git Jiamian Wang, Zongliang Wu, Yulun Zhang 0001, Xin Yuan 0002, Zhiqiang Tao |
NeurIPS | 4 |
| 2024 | Hybrid CNN-Transformer Architecture for Efficient Large-Scale Video Snapshot Compressive Imaging
Lishun Wang, Xin Yuan 0002 |
Int. J. Comput. Vis. | 4 |
| 2024 | Motion-Aware Dynamic Graph Neural Network for Video Compressive SensingabstractVideo snapshot compressive imaging (SCI) utilizes a 2D detector to capture sequential video frames and compress them into a single measurement. Various reconstruction methods have been developed to recover the high-speed video frames from the snapshot measurement. However, most existing reconstruction methods are incapable of efficiently capturing long-range spatial and temporal dependencies, which are critical for video processing. In this paper, we propose a flexible and robust approach based on the graph neural network (GNN) to efficiently model non-local interactions between pixels in space and time regardless of the distance. Specifically, we develop a motion-aware dynamic GNN for better video representation, i.e., represent each node as the aggregation of relative neighbors under the guidance of frame-by-frame motions, which consists of motion-aware dynamic sampling, cross-scale node sampling, global knowledge integration, and graph aggregation. Extensive results on both simulation and real data demonstrate both the effectiveness and efficiency of the proposed approach, and the visualization illustrates the intrinsic dynamic sampling operations of our proposed model for boosting the video SCI reconstruction results. The code and model will be released. Ruiying Lu, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Structured residual sparsity for video compressive sensing reconstruction
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
Signal Process. | 3 |
| 2024 | Multiple Complementary Priors for Multispectral Image Compressive Sensing ReconstructionabstractCompressive sensing (CS) techniques using a few compressed measurements have drawn considerable interest in reconstructing multispectral imagery (MSI). Nonlocal-based tensor methods have been widely used for MSI-CS reconstruction, which employ the nonlocal self-similarity (NSS) property of MSI to obtain satisfactory results. However, such methods only consider the internal priors of MSI while ignoring important external image information, for example deep-driven priors learned from a corpus of natural image datasets. Meanwhile, they usually suffer from annoying ringing artifacts due to the aggregation of overlapping patches. In this article, we propose a novel approach for highly effective MSI-CS reconstruction using multiple complementary priors (MCPs). The proposed MCP jointly exploits nonlocal low-rank and deep image priors under a hybrid plug-and-play framework, which contains multiple pairs of complementary priors, namely, internal and external, shallow and deep, and NSS and local spatial priors. To make the optimization tractable, a well-known alternating direction method of multiplier (ADMM) algorithm based on the alternating minimization framework is developed to solve the proposed MCP-based MSI-CS reconstruction problem. Extensive experimental results demonstrate that the proposed MCP algorithm outperforms many state-of-the-art CS techniques in MSI reconstruction. The source code of the proposed MCP-based MSI-CS reconstruction algorithm is available at: https://github.com/zhazhiyuan/MCP_MSI_CS_Demo.git. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Xudong Jiang 0001, Ce Zhu |
IEEE Trans. Cybern. | 3 |
| 2024 | Degradation-Aware Dynamic Fourier-Based Network for Spectral Compressive ImagingabstractWe consider the problem of hyperspectral image (HSI) reconstruction, which aims to recover 3D hyperspectral data from 2D compressive HSI measurements acquired by a coded aperture snapshot spectral imaging (CASSI) system. Existing deep learning methods have achieved acceptable results in HSI reconstruction. However, these methods did not consider the imaging system degradation pattern. In this article, based on observing the initialized HSIs obtained by shifting and splitting the measurements, we propose a dynamic Fourier network based on degradation learning, called the degradation-aware dynamic Fourier-based network (DADF-Net). We estimate the degradation feature maps from the degraded hyperspectral images to realize the linear transformation and dynamic processing of the features. In particular, we use the Fourier transform to extract the HSI non-local features. Extensive experimental results show that the proposed model outperforms state-of-the-art algorithms on simulation and real-world HSI datasets. Lei Liu 0067, Haifeng Zheng, Xin Yuan 0002, Lingyun Xue |
IEEE Trans. Multim. | 4 |
| 2024 | Plug-and-Play Algorithms for Dynamic Non-line-of-sight ImagingabstractNon-line-of-sight (NLOS) imaging has the ability to recover 3D images of scenes outside the direct line of sight, which is of growing interest for diverse applications. Despite the remarkable progress, NLOS imaging of dynamic objects is still challenging. It requires a large amount of multibounce photons for the reconstruction of single-frame data. To overcome this obstacle, we develop a computational framework for dynamic time-of-flight NLOS imaging based on plug-and-play (PnP) algorithms. By combining imaging forward model with the deep denoising network from the computer vision community, we show a 4 frames-per-second (fps) 3D NLOS video recovery (128 × 128 × 512) in post-processing. Our method leverages the temporal similarity among adjacent frames and incorporates sparse priors and frequency filtering. This enables higher-quality reconstructions for complex scenes. Extensive experiments are conducted to verify the superior performance of our proposed algorithm both through simulations and real data. Juntian Ye, Yu Hong 0004, Xiongfei Su, Xin Yuan 0002, Feihu Xu |
ACM Trans. Graph. | 4 |
| 2023 | Deep Equilibrium Models for Snapshot Compressive ImagingabstractThe ability of snapshot compressive imaging (SCI) systems to efficiently capture high-dimensional (HD) data has led to an inverse problem, which consists of recovering the HD signal from the compressed and noisy measurement. While reconstruction algorithms grow fast to solve it with the recent advances of deep learning, the fundamental issue of accurate and stable recovery remains. To this end, we propose deep equilibrium models (DEQ) for video SCI, fusing data-driven regularization and stable convergence in a theoretically sound manner. Each equilibrium model implicitly learns a nonexpansive operator and analytically computes the fixed point, thus enabling unlimited iterative steps and infinite network depth with only a constant memory requirement in training and testing. Specifically, we demonstrate how DEQ can be applied to two existing models for video SCI reconstruction: recurrent neural networks (RNN) and Plug-and-Play (PnP) algorithms. On a variety of datasets and real data, both quantitative and qualitative evaluations of our results demonstrate the effectiveness and stability of our proposed method. The code and models are available at: https://github.com/IndigoPurple/DEQSCI. Siming Zheng, Xin Yuan 0002 |
AAAI | 3 |
| 2023 | EfficientSCI: Densely Connected Network with Space-time Factorization for Large-scale Video Snapshot Compressive ImagingabstractVideo snapshot compressive imaging (SCI) uses a twodimensional detector to capture consecutive video frames during a single exposure time. Following this, an efficient reconstruction algorithm needs to be designed to reconstruct the desired video frames. Although recent deep learning-based state-of-the-art (SOTA) reconstruction algorithms have achieved good results in most tasks, they still face the following challenges due to excessive model complexity and GPU memory limitations: 1) these models need high computational cost, and 2) they are usually unable to reconstruct large-scale video frames at high compression ratios. To address these issues, we develop an efficient network for video SCI by using dense connections and space-time factorization mechanism within a single residual block, dubbed EfficientSCI. The EfficientSCI network can well establish spatial-temporal correlation by using convolution in the spatial domain and Transformer in the temporal domain, respectively. We are the first time to show that an UHD color video with high compression ratio can be reconstructed from a snapshot 2D measurement using a single end-to-end deep learning model with PSNR above 32 dB. Extensive results on both simulation and real data show that our method significantly outperforms all previous SOTA algorithms with better real-time performance. The code is at https://github.com/ucaswangls/EfficientSCI.git. Lishun Wang, Xin Yuan 0002 |
CVPR | 3 |
| 2023 | Hyperspectral Image Denoising Via Nonlocal Rank Residual ModelingabstractNonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an LR penalty on each nonlocal full-band group. However, in most existing methods, the LR tensor is only approximated directly from the degraded nonlocal full-band tensor, which is subject to certain issues (e.g., in heavy noise environments) in obtaining a suboptimal tensor approximation, and thus leading to unsatisfactory denoising results. In this paper, we propose a novel nonlocal rank residual (NRR) approach for highly effective HSI denoising, which progressively approximates the underlying L-R tensor via minimizing the rank residual. Towards this end, we first obtain a good estimate of the original nonlocal full-band group by using the NSS prior, and then the rank residual between the de-graded nonlocal full-band group with the corresponding estimated nonlocal full-band group is minimized to achieve a more accurate LR tensor. Moreover, the global spectral LR prior is employed to reduce the spectral redundancy of HSI in the proposed denoising framework. Finally, we develop a simple yet effective alternating minimization algorithm to jointly refine global spectral information and nonlocal full-band groups. Experimental results clearly show that the proposed NRR algorithm outperforms many state-of-the-art HSI denoising methods. The source code of the proposed NRR algorithm for HSI denoising is available at: https://github.com/zhazhiyuan/NRR_HSI_Denoising_Demo.git. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICASSP | 3 |
| 2023 | Deep Optics for Video Snapshot Compressive ImagingabstractVideo snapshot compressive imaging (SCI) aims to capture a sequence of video frames with only a single shot of a 2D detector, whose backbones rest in optical modulation patterns (also known as masks) and a computational reconstruction algorithm. Advanced deep learning algorithms and mature hardware are putting video SCI into practical applications. Yet, there are two clouds in the sunshine of SCI: i) low dynamic range as a victim of high temporal multiplexing, and ii) existing deep learning algorithms’ degradation on real system. To address these challenges, this paper presents a deep optics framework to jointly optimize masks and a reconstruction network. Specifically, we first propose a new type of structural mask to realize motionaware and full-dynamic-range measurement. Considering the motion awareness property in measurement domain, we develop an efficient network for video SCI reconstruction using Transformer to capture long-term temporal dependencies, dubbed Res2former. Moreover, sensor response is introduced into the forward model of video SCI to guarantee end-to-end model training close to real system. Finally, we implement the learned structural masks on a digital micro-mirror device. Experimental results on synthetic and real data validate the effectiveness of the proposed frame-work. We believe this is a milestone for real-world video SCI. The source code and data are available at https://github.com/pwangcs/DeepOpticsSCI. Ping Wang 0029, Lishun Wang, Xin Yuan 0002 |
ICCV | 3 |
| 2023 | Unfolding Framework with Prior of Convolution-Transformer Mixture and Uncertainty Estimation for Video Snapshot Compressive ImagingabstractWe consider the problem of video snapshot compressive imaging (SCI), where sequential high-speed frames are modulated by different masks and captured by a single measurement. The underlying principle of reconstructing multi-frame images from only one single measurement is to solve an ill-posed problem. By combining optimization algorithms and neural networks, deep unfolding networks (DUNs) score tremendous achievements in solving inverse problems. In this paper, our proposed model is under the DUN framework and we propose a 3D Convolution-Transformer Mixture (CTM) module with a 3D efficient and scalable attention model plugged in, which helps fully learn the correlation between temporal and spatial dimensions by virtue of Transformer. To our best knowledge, this is the first time that Transformer is employed to video SCI reconstruction. Besides, to further investigate the high-frequency information during the reconstruction process which are neglected in previous studies, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art (SOTA) (with a 1.2dB gain in PSNR over previous SOTA algorithm) results. Code can be found on https://github.com/zsm1211/CTM-SCI. Siming Zheng, Xin Yuan 0002 |
ICCV | 2 |
| 2023 | Accurate Image Restoration with Attention Retractable Transformer
Yulun Zhang 0001, Jinjin Gu, Yongbing Zhang 0002, Linghe Kong, Xin Yuan 0002 |
ICLR | 6 |
| 2023 | SAUNet: Spatial-Attention Unfolding Network for Image Compressive SensingabstractImage Compressive Sensing (CS) enables compressed capture of natural images via a spatial multiplexing camera and accurate reconstruction from few measurements via an advanced algorithm. Deep learning, especially deep unfolding, has recently achieved impressive success in image CS reconstruction. However, existing learning-based methods have been developed for block (usually with 33 X 33 pixels) CS instead of full image CS. Apart from the difficulties in hardware implementation, block CS breaks the global pixel interactions, limiting the overall performance. In this paper, we propose the first two-dimensional deep unfolding framework, and further develop a Spatial-Attention Unfolding Network (SAUNet) for full image CS reconstruction by alternately performing a spatially-adaptive gradient descent module and a cross-stage multi-scale denoising module. The gradient descent module has the spatial self-adaptation to the degradation of in-process image. The denoising module is a three-level U-shaped structure powered by Convolutional Self-Attention (CSA) mechanism. Inspired by Transformer, CSA is designed to adaptively aggregate spatially local information and adaptively recalibrate channel-wise global information with only normal convolutional operator. Extensive experiments demonstrate that SAUNet outperforms the state-of-the-art methods by a large margin. The source code and pre-trained models are available at https://github.com/pwangcs/SAUNet. Ping Wang 0029, Xin Yuan 0002 |
ACM Multimedia | 2 |
| 2023 | Binarized Spectral Compressive ImagingabstractExisting deep learning models for hyperspectral image (HSI) reconstruction achieve good performance but require powerful hardwares with enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited mobile devices. In this paper, we propose a novel method, Binarized Spectral-Redistribution Network (BiSRNet), for efficient and practical HSI restoration from compressed measurement in snapshot compressive imaging (SCI) systems. Firstly, we redesign a compact and easy-to-deploy base model to be binarized. Then we present the basic unit, Binarized Spectral-Redistribution Convolution (BiSR-Conv). BiSR-Conv can adaptively redistribute the HSI representations before binarizing activation and uses a scalable hyperbolic tangent function to closer approximate the Sign function in backpropagation. Based on our BiSR-Conv, we customize four binarized convolutional modules to address the dimension mismatch and propagate full-precision information throughout the whole network. Finally, our BiSRNet is derived by using the proposed techniques to binarize the base model. Comprehensive quantitative and qualitative experiments manifest that our proposed BiSRNet outperforms state-of-the-art binarization algorithms. Code and models are publicly available at https://github.com/caiyuanhao1998/BiSCI Yuanhao Cai, Xin Yuan 0002, Yulun Zhang 0001, Haoqian Wang |
NeurIPS | 4 |
| 2023 | Hierarchical Integration Diffusion Model for Realistic Image DeblurringabstractDiffusion models (DMs) have recently been introduced in image deblurring and exhibited promising performance, particularly in terms of details reconstruction. However, the diffusion model requires a large number of inference iterations to recover the clean image from pure Gaussian noise, which consumes massive computational resources. Moreover, the distribution synthesized by the diffusion model is often misaligned with the target results, leading to restrictions in distortion-based metrics. To address the above issues, we propose the Hierarchical Integration Diffusion Model (HI-Diff), for realistic image deblurring. Specifically, we perform the DM in a highly compacted latent space to generate the prior feature for the deblurring process. The deblurring process is implemented by a regression-based method to obtain better distortion accuracy. Meanwhile, the highly compact latent space ensures the efficiency of the DM. Furthermore, we design the hierarchical integration module to fuse the prior into the regression-based model from multiple scales, enabling better generalization in complex blurry scenarios. Comprehensive experiments on synthetic and real-world blur datasets demonstrate that our HI-Diff outperforms state-of-the-art methods. Code and trained models are available at https://github.com/zhengchen1999/HI-Diff. Zheng Chen 0014, Yulun Zhang 0001, Ding Liu 0001, Bin Xia 0014, Jinjin Gu, Linghe Kong, Xin Yuan 0002 |
NeurIPS | 7 |
| 2023 | Multi-scale Iterative Model-guided Unfolding Network for NLOS ReconstructionabstractAbstract Non‐line‐of‐sight (NLOS) imaging can reconstruct hidden objects by analyzing diffuse reflection of relay surfaces, and is potentially used in autonomous driving, medical imaging and national defense. Despite the challenges of low signal‐to‐noise ratio (SNR) and ill‐conditioned problem, NLOS imaging has developed rapidly in recent years. While deep neural networks have achieved impressive success in NLOS imaging, most of them lack flexibility when dealing with multiple spatial‐temporal resolution and multi‐scene images in practical applications. To bridge the gap between learning methods and physical priors, we present a novel end‐to‐end Multi‐scale Iterative Model‐guided Unfolding (MIMU), with superior performance and strong flexibility. Furthermore, we overcome the lack of real training data with a general architecture that can be trained in simulation. Unlike existing encoder‐decoder architectures and generative adversarial networks, the proposed method allows for only one trained model adaptive for various dimensions, such as various sampling time resolution, various spatial resolution and multiple channels for colorful scenes. Simulation and real‐data experiments verify that the proposed method achieves better reconstruction results both in quality and quantity than existing methods. Xiongfei Su, Yu Hong 0004, Juntian Ye, Feihu Xu, Xin Yuan 0002 |
Comput. Graph. Forum | 5 |
| 2023 | Deep Unfolding for Snapshot Compressive Imaging
Ziyi Meng 0001, Xin Yuan 0002, Shirin Jalali |
Int. J. Comput. Vis. | 2 |
| 2023 | Adaptive Deep PnP Algorithm for Video Snapshot Compressive Imaging
Zongliang Wu, Chengshuai Yang, Xiongfei Su, Xin Yuan 0002 |
Int. J. Comput. Vis. | 4 |
| 2023 | Recurrent Neural Networks for Snapshot Compressive ImagingabstractConventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT. Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Deep Gaussian Scale Mixture Prior for Image ReconstructionabstractImage reconstruction from partial observations has attracted increasing attention. Conventional image reconstruction methods with hand-crafted priors often fail to recover fine image details due to the poor representation capability of the hand-crafted priors. Deep learning methods attack this problem by directly learning mapping functions between the observations and the targeted images can achieve much better results. However, most powerful deep networks lack transparency and are nontrivial to design heuristically. This paper proposes a novel image reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Unlike existing unfolding methods that only estimate the image means (i.e., the denoising prior) but neglected the variances, we propose characterizing images by the GSM models with learned means and variances through a deep network. Furthermore, to learn the long-range dependencies of images, we develop an enhanced variant based on the Swin Transformer for learning GSM models. All parameters of the MAP estimator and the deep network are jointly optimized through end-to-end training. Extensive simulation and real data experimental results on spectral compressive imaging and image super-resolution demonstrate that the proposed method outperforms existing state-of-the-art methods. Xin Yuan 0002, Weisheng Dong, Jinjian Wu, Guangming Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Spatial-Temporal Transformer for Video Snapshot Compressive ImagingabstractVideo snapshot compressive imaging (SCI) captures multiple sequential video frames by a single measurement using the idea of computational imaging. The underlying principle is to modulate high-speed frames through different masks and these modulated frames are summed to a single measurement captured by a low-speed 2D sensor (dubbed optical encoder); following this, algorithms are employed to reconstruct the desired high-speed frames (dubbed software decoder) if needed. In this article, we consider the reconstruction algorithm in video SCI, i.e., recovering a series of video frames from a compressed measurement. Specifically, we propose a Spatial-Temporal transFormer (STFormer) to exploit the correlation in both spatial and temporal domains. STFormer network is composed of a token generation block, a video reconstruction block, and these two blocks are connected by a series of STFormer blocks. Each STFormer block consists of a spatial self-attention branch, a temporal self-attention branch and the outputs of these two branches are integrated by a fusion network. Extensive results on both simulated and real data demonstrate the state-of-the-art performance of STFormer. The code and models are publicly available at https://github.com/ucaswangls/STFormer. Lishun Wang, Yong Zhong, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Computational Imaging and Artificial Intelligence: The Next Revolution of Mobile VisionabstractSignal capture is at the forefront of perceiving and understanding the environment; thus, imaging plays a pivotal role in mobile vision. Recent unprecedented progress in artificial intelligence (AI) has shown great potential in the development of advanced mobile platforms with new imaging devices. Traditional imaging systems based on the “capturing images first and processing afterward” mechanism cannot meet this explosive demand. On the other hand, computational imaging (CI) systems are designed to capture high-dimensional data in an encoded manner to provide more information for mobile vision systems. Thanks to AI, CI can now be used in real-life systems by integrating deep learning algorithms into the mobile vision platform to achieve a closed loop of intelligent acquisition, processing, and decision-making, thus leading to the next revolution of mobile vision. Starting from the history of mobile vision using digital cameras, this work first introduces the advancement of CI in diverse applications and then conducts a comprehensive review of current research topics combining CI and AI. Although new-generation mobile platforms, represented by smart mobile phones, have deeply integrated CI and AI for better image acquisition and processing, most mobile vision platforms, such as self-driving cars and drones only loosely connect CI and AI, and are calling for a closer integration. Motivated by this fact, at the end of this work, we propose some potential technologies and disciplines that aid the deep integration of CI and AI and shed light on new directions in the future generation of mobile vision platforms. Jin-Li Suo, Jin Gong, Xin Yuan 0002, David J. Brady, Qionghai Dai |
Proc. IEEE | 4 |
| 2023 | Universal Domain Adaptation for Remote Sensing Image Scene ClassificationabstractThe domain adaptation (DA) approaches available to date are usually not well suited for practical DA scenarios of remote sensing image classification since these methods (such as unsupervised DA) rely on rich prior knowledge about the relationship between label sets of source and target domains, and source data are often not accessible due to privacy or confidentiality issues. To this end, we propose a practical universal DA (UniDA) setting for remote sensing image scene classification that requires no prior knowledge on the label sets. Furthermore, a novel UniDA method without source data is proposed for cases when the source data are unavailable. The architecture of the model is divided into two parts: the source data generation stage and the model adaptation stage. The first stage estimates the conditional distribution of source data from the pretrained model using the knowledge of class separability in the source domain and then synthesizes the source data. With this synthetic source data in hand, it becomes a UniDA task to classify a target sample correctly if it belongs to any category in the source label set or mark it as “unknown” otherwise. In the second stage, a novel transferable weight that distinguishes the shared and private label sets in each domain promotes the adaptation in the automatically discovered shared label set and recognizes the “unknown” samples successfully. Empirical results show that the proposed model is effective and practical for remote sensing image scene classification, regardless of whether the source data are available or not. The code is available athttps://github.com/zhu-xlab/UniDA. Qingsong Xu 0001, Yilei Shi, Xin Yuan 0002, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Nonlocal Structured Sparsity Regularization Modeling for Hyperspectral Image DenoisingabstractThe non-local-based model for hyperspectral image (HSI) denoising first uses non-local self-similarity (NSS) prior to group similar full-band patches into three-dimensional non-local full-band groups (tensors) using a block matching (BM) operation, and then a low-rank (LR) penalty is typically applied to each non-local full-band group to reduce noise. While non-local-based methods have shown promising performance in HSI denoising, most existing methods have only considered the LR property of the non-local full-band group while ignoring the strong correlation between sparse coefficients. Moreover, such methods often result in unsatisfactory visual artifacts due to the noise sensitivity of BM operations, while requiring expensive computations. To address these limitations, this paper proposes a novel non-local structured sparsity regularization (NLSSR) approach for HSI denoising. First, to mitigate the noise sensitivity of the BM operation, we propose a graph-based domain distance scheme to index similar full-band patches to form the non-local full-band group. Second, we design an adaptive unidirectional low-rank (LR) dictionary with low complexity that takes into account the differences in intrinsic structure correlation among different modes of the non-local full-band tensor. Third, we utilize a global spectral LR prior to reduce spectral redundancy. Fourth, we develop a generalized soft-thresholding (GST) algorithm based on the alternating minimization framework to solve the NLSSR-based HSI denoising problem. We perform extensive experiments on both simulated and real data to show that the proposed NLSSR algorithm outperforms many popular or state-of-the-art HSI denoising methods in both quantitative and visual evaluations. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Yilong Lu, Ce Zhu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Reinforcement Learning for Adaptive Video Compressive SensingabstractWe apply reinforcement learning to video compressive sensing to adapt the compression ratio. Specifically, video snapshot compressive imaging (SCI), which captures high-speed video using a low-speed camera is considered in this work, in which multiple ( B ) video frames can be reconstructed from a snapshot measurement. One research gap in previous studies is how to adapt B in the video SCI system for different scenes. In this article, we fill this gap utilizing reinforcement learning (RL). An RL model, as well as various convolutional neural networks for reconstruction, are learned to achieve adaptive sensing of video SCI systems. Furthermore, the performance of an object detection network using directly the video SCI measurements without reconstruction is also used to perform RL-based adaptive video compressive sensing. Our proposed adaptive SCI method can thus be implemented in low cost and real time. Our work takes the technology one step further towards real applications of video SCI. Sidi Lu, Xin Yuan 0002, Aggelos K. Katsaggelos, Weisong Shi |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | Low-Rankness Guided Group Sparse Representation for Image RestorationabstractAs a spotlighted nonlocal image representation model, group sparse representation (GSR) has demonstrated a great potential in diverse image restoration tasks. Most of the existing GSR-based image restoration approaches exploit the nonlocal self-similarity (NSS) prior by clustering similar patches into groups and imposing sparsity to each group coefficient, which can effectively preserve image texture information. However, these methods have imposed only plain sparsity over each individual patch of the group, while neglecting other beneficial image properties, e.g., low-rankness (LR), leads to degraded image restoration results. In this article, we propose a novel low-rankness guided group sparse representation (LGSR) model for highly effective image restoration applications. The proposed LGSR jointly utilizes the sparsity and LR priors of each group of similar patches under a unified framework. The two priors serve as the complementary priors in LGSR for effectively preserving the texture and structure information of natural images. Moreover, we apply an alternating minimization algorithm with an adaptively adjusted parameter scheme to solve the proposed LGSR-based image restoration problem. Extensive experiments are conducted to demonstrate that the proposed LGSR achieves superior results compared with many popular or state-of-the-art algorithms in various image restoration tasks, including denoising, inpainting, and compressive sensing (CS). Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactions is beneficial for HSI reconstruction. However, existing CNN-based methods show limitations in capturing spectral-wise similarity and long-range dependencies. Besides, the HSI information is modulated by a coded aperture (physical mask) in CASSI. Nonetheless, current algorithms have not fully explored the guidance effect of the mask for HSI restoration. In this paper, we propose a novel framework, Mask-guided Spectral-wise Transformer (MST), for HSI reconstruction. Specifically, we present a Spectral-wise Multi-head Self-Attention (S-MSA) that treats each spectral feature as a token and calculates self-attention along the spectral dimension. In addition, we customize a Mask-guided Mechanism (MM) that directs S- MSA to pay attention to spatial regions with high-fidelity spectral representations. Extensive experiments show that our MST significantly outperforms state-of-the-art (SOTA) methods on simulation and real HSI datasets while requiring dramatically cheaper computational and memory costs. https://github.com/caiyuanhao1998/MST/ Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
CVPR | 5 |
| 2022 | HDNet: High-resolution Dual-domain Learning for Spectral Compressive ImagingabstractThe rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance against complexity, losing fine-grained high-resolution (HR) features. Secondly, even if the optimization focusing on spatial-spectral domain learning (SDL) converges to the ideal solution, there is still a significant visual difference between the reconstructed HSI and the truth. So we propose a high-resolution dual-domain learning network (HDNet) for HSI reconstruction. On the one hand, the proposed HR spatial-spectral attention module with its efficient feature fusion provides continuous and fine pixel-level features. On the other hand, frequency domain learning (FDL) is introduced for HSI reconstruction to narrow the frequency domain discrepancy. Dynamic FDL supervision forces the model to reconstruct fine-grained frequencies and compensate for excessive smoothing and distortion caused by pixel-level losses. The HR pixel-level attention and frequency-level refinement in our HDNet mutually promote HSI perceptual quality. Extensive quantitative and qualitative experiments show that our method achieves SOTA performance on simulated and real HSI datasets. https://github.com/Huxiaowan/HDNet Xiaowan Hu, Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
CVPR | 5 |
| 2022 | Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Xin Yuan 0002, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
ECCV (17) | 5 |
| 2022 | Modeling Mask Uncertainty in Hyperspectral Image Reconstruction
Jiamian Wang, Yulun Zhang 0001, Xin Yuan 0002, Ziyi Meng 0001, Zhiqiang Tao |
ECCV (19) | 3 |
| 2022 | Ensemble Learning Priors Driven Deep Unfolding for Scalable Video Snapshot Compressive Imaging
Chengshuai Yang, Xin Yuan 0002 |
ECCV (23) | 3 |
| 2022 | Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson DenoisingabstractPoisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultaneous nonlocal low-rank and deep priors (SNLDP) for Poisson denoising. The proposed SNLD-P simultaneously employs nonlocal self-similarity and deep image priors under the hybrid plug and play framework, which comprises multiple pairs of complementary priors, namely, nonlocal and local, shallow and deep, and internal and external. To make the optimization tractable, an effective alternating direction method of multiplier (ADMM) algorithm under the alternative minimization framework is provided to solve the proposed SNLDP-based Poisson denoising problem. Experimental results demonstrate the superiority of the proposed SNLDP over many popular or state-of-the-art Poisson denoising algorithms in terms of quantitative and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICASSP | 3 |
| 2022 | Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingabstractIn coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer from two issues. Firstly, they do not estimate the degradation patterns and ill-posedness degree from CASSI to guide the iterative learning. Secondly, they are mainly CNN-based, showing limitations in capturing long-range dependencies. In this paper, we propose a principled Degradation-Aware Unfolding Framework (DAUF) that estimates parameters from the compressed image and physical mask, and then uses these parameters to control each iteration. Moreover, we customize a novel Half-Shuffle Transformer (HST) that simultaneously captures local contents and non-local dependencies. By plugging HST into DAUF, we establish the first Transformer-based deep unfolding method, Degradation-Aware Unfolding Half-Shuffle Transformer (DAUHST), for HSI reconstruction. Experiments show that DAUHST surpasses state-of-the-art methods while requiring cheaper computational and memory costs. Code and models are publicly available at https://github.com/caiyuanhao1998/MST Yuanhao Cai, Haoqian Wang, Xin Yuan 0002, Henghui Ding, Yulun Zhang 0001, Radu Timofte, Luc Van Gool |
NeurIPS | 4 |
| 2022 | Cross Aggregation Transformer for Image RestorationabstractRecently, Transformer architecture has been introduced into image restoration to replace convolution neural network (CNN) with surprising results. Considering the high computational complexity of Transformer with global attention, some methods use the local square window to limit the scope of self-attention. However, these methods lack direct interaction among different windows, which limits the establishment of long-range dependencies. To address the above issue, we propose a new image restoration model, Cross Aggregation Transformer (CAT). The core of our CAT is the Rectangle-Window Self-Attention (Rwin-SA), which utilizes horizontal and vertical rectangle window attention in different heads parallelly to expand the attention area and aggregate the features cross different windows. We also introduce the Axial-Shift operation for different window interactions. Furthermore, we propose the Locality Complementary Module to complement the self-attention mechanism, which incorporates the inductive bias of CNN (e.g., translation invariance and locality) into Transformer, enabling global-local coupling. Extensive experiments demonstrate that our CAT outperforms recent state-of-the-art methods on several image restoration applications. The code and models are available at https://github.com/zhengchen1999/CAT. Zheng Chen 0014, Yulun Zhang 0001, Jinjin Gu, Yongbing Zhang 0002, Linghe Kong, Xin Yuan 0002 |
NeurIPS | 6 |
| 2022 | Plug-and-Play Algorithms for Video Snapshot Compressive ImagingabstractWe consider the reconstruction problem of video snapshot compressive imaging (SCI), which captures high-speed videos using a low-speed 2D sensor (detector). The underlying principle of SCI is to modulate sequential high-speed frames with different masks and then these encoded frames are integrated into a snapshot on the sensor and thus the sensor can be of low-speed. On one hand, video SCI enjoys the advantages of low-bandwidth, low-power and low-cost. On the other hand, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging and one of the bottlenecks lies in the reconstruction algorithm. Existing algorithms are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload. We first employ the image deep denoising priors to show that PnP can recover a UHD color video with 30 frames from a snapshot measurement. Since videos have strong temporal correlation, by employing the video deep denoising priors, we achieve a significant improvement in the results. Furthermore, we extend the proposed PnP algorithms to the color SCI system using mosaic sensors, where each pixel only captures the red, green or blue channels. A joint reconstruction and demosaicing paradigm is developed for flexible and high quality reconstruction of color video SCI systems. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm. Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Frédo Durand, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Nonconvex Structural Sparsity Residual Constraint for Image RestorationabstractThis article proposes a novel nonconvex structural sparsity residual constraint (NSSRC) model for image restoration, which integrates structural sparse representation (SSR) with nonconvex sparsity residual constraint (NC-SRC). Although SSR itself is powerful for image restoration by combining the local sparsity and nonlocal self-similarity in natural images, in this work, we explicitly incorporate the novel NC-SRC prior into SSR. Our proposed approach provides more effective sparse modeling for natural images by applying a more flexible sparse representation scheme, leading to high-quality restored images. Moreover, an alternating minimizing framework is developed to solve the proposed NSSRC-based image restoration problems. Extensive experimental results on image denoising and image deblocking validate that the proposed NSSRC achieves better results than many popular or state-of-the-art methods over several publicly available datasets. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Ce Zhu |
IEEE Trans. Cybern. | 2 |
| 2022 | Class-Aware Domain Adaptation for Semantic Segmentation of Remote Sensing ImagesabstractUnsupervised domain adaptation (UDA) for the semantic segmentation of remote sensing images is challenging since the same class of objects may have different spectra while the different class of objects may have the same spectrum. To address this issue, we propose a class-aware generative adversarial network (CaGAN) for UDA semantic segmentation of multisource remote sensing images, which explicitly models the discrepancies of intraclass and the interclass between the source domain images with labels and the target domain images without labels. Specifically, first, to enhance the global domain alignment (GDA), we propose a transferable attention alignment (TAA) procedure to add more fine-grained features into the adversarial learning framework. Then, we propose a novel class-aware domain alignment (CDA) approach in semantic segmentation. CDA mainly includes two parts: the first one is adaptive category selection, which is to alleviate the class imbalance and select the reliable per-category centers in the source and target domains; the second one is adaptive category alignment, which is to model the intraclass compactness and interclass separability from source-only, target-only, and joint source and target images. Finally, the CDA plays as a penalty of GDA to train GaGAN in an alternating and iterative manner. Experiments on domain adaptation of space to space, spectrum to spectrum, both space-to-space and spectrum-to-spectrum data sets demonstrate that CaGAN outperforms the current state-of-the-art methods, which may serve as a starting point and baseline for the comprehensive applications of semantic segmentation in cross-space and cross-spectrum remote sensing images. Qingsong Xu 0001, Xin Yuan 0002, Chaojun Ouyang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Hybrid Structural Sparsification Error Model for Image RestorationabstractRecent works on structural sparse representation (SSR), which exploit image nonlocal self-similarity (NSS) prior by grouping similar patches for processing, have demonstrated promising performance in various image restoration applications. However, conventional SSR-based image restoration methods directly fit the dictionaries or transforms to the internal (corrupted) image data. The trained internal models inevitably suffer from overfitting to data corruption, thus generating the degraded restoration results. In this article, we propose a novel hybrid structural sparsification error (HSSE) model for image restoration, which jointly exploits image NSS prior using both the internal and external image data that provide complementary information. Furthermore, we propose a general image restoration scheme based on the HSSE model, and an alternating minimization algorithm for a range of image restoration applications, including image inpainting, image compressive sensing and image deblocking. Extensive experiments are conducted to demonstrate that the proposed HSSE-based scheme outperforms many popular or state-of-the-art image restoration methods in terms of both objective metrics and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Alex Chichung Kot |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Memory-Efficient Network for Large-Scale Video Compressive SensingabstractVideo snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net. Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002 |
CVPR | 7 |
| 2021 | Deep Gaussian Scale Mixture Prior for Spectral Compressive ImagingabstractIn coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achieved limited success due to the poor representation capability of these hand-crafted priors. Deep learning based methods learning the mappings between the compressive images and the HSIs directly achieved much better results. Yet, it is nontrivial to design a powerful deep network heuristically for achieving satisfied results. In this paper, we propose a novel HSI reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Different from existing GSM models using hand-crafted scale priors (e.g., the Jeffrey’s prior), we propose to learn the scale prior through a deep convolutional neural network (DCNN). Furthermore, we also propose to estimate the local means of the GSM models by the DCNN. All the parameters of the MAP estimation algorithm and the DCNN parameters are jointly optimized through end-to-end training. Extensive experimental results on both synthetic and real datasets demonstrate that the proposed method outperforms existing state-of-the-art methods. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/DGSM-SCI.htm. Weisheng Dong, Xin Yuan 0002, Jinjian Wu, Guangming Shi |
CVPR | 3 |
| 2021 | MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive SensingabstractTo capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-speed frames, where the state-of-the-art results are achieved by deep learning networks. However, these networks are usually trained for specific small-scale masks and often have high demands of training time and GPU memory, which are hence not flexible to i) a new mask with the same size and ii) a larger-scale mask. We address these challenges by developing a Meta Modulated Convolutional Network for SCI reconstruction, dubbed MetaSCI. MetaSCI is composed of a shared backbone for different masks, and light-weight meta-modulation parameters to evolve to different modulation parameters for each mask, thus having the properties of fast adaptation to new masks (or systems) and ready to scale to large data. Extensive simulation and real data results demonstrate the superior performance of our proposed approach. Our code is available at https://github.com/xyvirtualgroup/MetaSCI-CVPR2021. Zhengjue Wang, Hao Zhang 0050, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002 |
CVPR | 5 |
| 2021 | Universal and Flexible Optical Aberration Correction Using Deep-Prior Based DeconvolutionabstractHigh quality imaging usually requires bulky and expensive lenses to compensate geometric and chromatic aberrations. This poses high constraints on the optical hash or low cost applications. Although one can utilize algorithmic reconstruction to remove the artifacts of low-end lenses, the degeneration from optical aberrations is spatially varying and the computation has to trade off efficiency for performance. For example, we need to conduct patch-wise optimization or train a large set of local deep neural networks to achieve high reconstruction performance across the whole image. In this paper, we propose a PSF aware deep network, which takes the aberrant image and PSF map as input and produces the latent high quality version via incorporating deep priors, thus leading to a universal and flexible optical aberration correction method. Specifically, we pre-train a base model from a set of diverse lenses and then adapt it to a given lens by quickly refining the parameters, which largely alleviates the time and memory consumption of model learning. The approach is of high efficiency in both training and testing stages. Extensive results verify the promising applications of our proposed approach for compact low-end cameras. The code is available at https://github.com/leehsiu/UABC Xiu Li 0003, Jin-Li Suo, Xin Yuan 0002, Qionghai Dai |
ICCV | 4 |
| 2021 | Self-supervised Neural Networks for Spectral Snapshot Compressive ImagingabstractWe consider using untrained neural networks to solve the reconstruction problem of snapshot compressive imaging (SCI), which uses a two-dimensional (2D) detector to capture a high-dimensional (usually 3D) data-cube in a compressed manner. Various SCI systems have been built in recent years to capture data such as high-speed videos, hyperspectral images, and the state-of-the-art reconstruction is obtained by the deep neural networks. However, most of these networks are trained in an end-to-end manner by a large amount of corpus with sometimes simulated ground truth, measurement pairs. In this paper, inspired by the untrained neural networks such as deep image priors (DIP) and deep decoders, we develop a framework by integrating DIP into the plug-and-play regime, leading to a self-supervised network for spectral SCI reconstruction. Extensive synthetic and real data results show that the proposed algorithm without training is capable of achieving competitive results to the training based networks. Furthermore, by integrating the proposed method with a pre-trained deep denoising prior, we have achieved state-of-the-art results. Our code is available at https://github.com/mengziyi64/CASSI-Self-Supervised. Ziyi Meng 0001, Zhenming Yu, Kun Xu 0008, Xin Yuan 0002 |
ICCV | 4 |
| 2021 | Perception Inspired Deep Neural Networks For Spectral Snapshot Compressive ImagingabstractWe consider the inverse problem of coded aperture snapshot spectral imaging (CASSI), which captures the spatio-spectral data-cube using a snapshot 2D measurement and reconstructs the 3D hyperspectral images using algorithms. Recent advances of deep learning have boosted the image quality of the reconstructed hyperspectral images significantly, and this leads to an end-to-end real-time capture and reconstruction system. However, the network design for CASSI reconstruction is still at the incubation stage and usually an off-the-shelf network is employed and re-purposed. In this work, from a different perspective, inspired by the fact that most existing hyperspectral images are still in the visible bandwidth, we introduce the perceptual loss into the deep neural network for CASSI reconstruction. Extensive results on both simulation and real data demonstrate that with this small change, the reconstructed image quality can be improved dramatically using the same network. Ziyi Meng 0001, Xin Yuan 0002 |
ICIP | 2 |
| 2021 | Low-Rank Regularized Joint Sparsity for Image DenoisingabstractNonlocal sparse representation models such as group sparse representation (GSR), low-rankness and joint sparsity (JS) have shown great potentials in image denoising studies, by effectively exploiting image nonlocal self-similarity (NSS) property. Popular dictionary-based JS algorithms apply convex JS penalties in their objective functions, which avoid NP-hard sparse coding step, but lead to only approximately sparse representation. Such approximated JS models fail to impose low-rankness of the underlying image data, resulting in degraded quality in image restoration. To simultaneously exploit the low-rank and JS priors, we propose a novel low-rank regularized joint sparsity model, dubbed LRJS, to enhance the dependency (i. e., low-rankness) of similar patches, thus better suppress independent noise. Moreover, to make the optimization tractable and robust, an alternating minimization algorithm with an adaptive parameter adjustment strategy is developed to solve the proposed LRJS-based image denoising problem. Experimental results demonstrate that the proposed LRJS outperforms many popular or state-of-the-art denoising algorithms in terms of both objective and visual perception met- Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICIP | 3 |
| 2021 | Dual-view Snapshot Compressive Imaging via Optical Flow Aided Recurrent Neural Network
Ruiying Lu, Bo Chen 0001, Guanliang Liu, Ziheng Cheng 0001, Xin Yuan 0002 |
Int. J. Comput. Vis. | 6 |
| 2021 | Fast Hyperspectral Image Recovery of Dual-Camera Compressive Hyperspectral Imaging via Non-Iterative Subspace-Based FusionabstractCoded aperture snapshot spectral imaging (CASSI) is a promising technique for capturing three-dimensional hyperspectral images (HSIs), in which algorithms are used to perform the inverse problem of HSI reconstruction from a single coded two-dimensional (2D) measurement. Due to the ill-posed nature of this problem, various regularizers have been exploited to reconstruct 3D data from 2D measurements. Unfortunately, the accuracy and computational complexity are unsatisfactory. One feasible solution is to utilize additional information such as the RGB measurement in CASSI. Considering the combined CASSI and RGB measurements, in this paper, we propose a fusion model for HSI reconstruction. Specifically, we investigate the low-dimensional spectral subspace property of HSIs composed of a spectral basis and spatial coefficients. In particular, the RGB measurement is utilized to estimate the coefficients, while the CASSI measurement is adopted to provide the spectral basis. We further propose a patch processing strategy to enhance the spectral low-rank property of HSIs. The optimization of the proposed model requires neither iteration nor the spectral sensing matrix of the RGB detector. Extensive experiments on both simulated and real HSI datasets demonstrate that our proposed method not only outperforms previous state-of-the-art (iterative algorithms) methods in quality but also speeds up the reconstruction by more than 5000 times. Wei He 0003, Naoto Yokoya, Xin Yuan 0002 |
IEEE Trans. Image Process. | 3 |
| 2021 | Image Restoration via Reconciliation of Group Sparsity and Low-Rank ModelsabstractImage nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, e.g., JS enforces the sparse codes to share the same support, or too general, e.g., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely, low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. An alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem for different image restoration tasks, including image denoising, image deblocking, image inpainting, and image compressive sensing. Extensive experimental results demonstrate that the proposed LR-GSC algorithm outperforms many popular or state-of-the-art methods in terms of objective and perceptual metrics. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 3 |
| 2021 | Triply Complementary Priors for Image RestorationabstractRecent works that utilized deep models have achieved superior results in various image restoration (IR) applications. Such approach is typically supervised, which requires a corpus of training images with distributions similar to the images to be recovered. On the other hand, the shallow methods, which are usually unsupervised remain promising performance in many inverse problems, e.g., image deblurring and image compressive sensing (CS), as they can effectively leverage nonlocal self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various artifacts due to naive patch aggregation in addition to the slow speed. Using either approach alone usually limits performance and generalizability in IR tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely, internal and external, shallow and deep, and non-local and local priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for IR. Following this, a simple yet effective algorithm is developed to solve the proposed H-PnP based IR problems. Extensive experimental results on several representative IR tasks, including image deblurring, image CS and image deblocking, demonstrate that the proposed H-PnP algorithm achieves favorable performance compared to many popular or state-of-the-art IR methods in terms of both objective and visual perception. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 3 |
| 2020 | Plug-and-Play Algorithms for Large-Scale Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) aims to capture the high-dimensional (usually 3D) images using a 2D sensor (detector) in a single snapshot. Though enjoying the advantages of low-bandwidth, low-power and low-cost, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging. The bottleneck lies in the reconstruction algorithms; they are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the widely used PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload and prove the {global convergence} of PnP-GAP under the SCI hardware constraints. By employing deep denoising priors, we first time show that PnP can recover a UHD color video (3840×1644×48 with PNSR above 30dB) from a snapshot 2D measurement. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm. Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Qionghai Dai |
CVPR | 1 |
| 2020 | BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002 |
ECCV (24) | 7 |
| 2020 | End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention
Ziyi Meng 0001, Jiawei Ma, Xin Yuan 0002 |
ECCV (23) | 3 |
| 2020 | A Hybrid Structural Sparse Error Model for Image DeblockingabstractInspired by the image nonlocal self-similarity (NSS) prior, structural sparse representation (SSR) models exploit each group as the basic unit for sparse representation, which have achieved promising results in various image restoration applications. However, conventional SSR models only exploited the group within the input degraded (internal) image for image restoration, which can be limited by over-fitting to data corruption. In this paper, we propose a novel hybrid structural sparse error (HSSE) model for image deblocking. The proposed HSSE model exploits image NSS prior over both the internal image and external image corpus, which can be complementary in both feature space and image plane. Moreover, we develop an alternating minimization with an adaptive parameter setting strategy to solve the proposed HSSE model. Experimental results demonstrate that the proposed HSSE-based image deblocking algorithm outperforms many state-of-the-art image deblocking methods in terms of objective and visual perception. Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen |
ICASSP | 2 |
| 2020 | Drcas: Deep Restoration Network For Hardware Based Compressive Acquisition SchemeabstractWe propose a novel image acquisition scheme HCAS (Hardware based Compressed Acquisition Scheme) using hardware-based binning (downsampling), bit truncation and JPEG compression and develop a deep learning based reconstruction network DRCAS (Deep Restoration network for hardware based Compressed Acquisition Scheme) for images acquired using HCAS. DRCAS to our best knowledge is the first work proposed in the literature for the restoration of images acquired using acquisition scheme like HCAS. It is also superior in performance than state-of-the-art super resolution networks while being much smaller. We show that HCAS and DRCAS techniques will enable us to design much simpler and power efficient image acquisition pipelines. Pravir Singh Gupta, Xin Yuan 0002, Gwan S. Choi |
ICIP | 2 |
| 2020 | The Power Of Triply Complementary Priors For Image Compressive SensingabstractRecent works that utilized deep models have achieved superior results in various image restoration applications. Such approach is typically supervised which requires a corpus of training images with distribution similar to the images to be recovered. On the other hand, the shallow methods which are usually unsupervised remain promising performance in many inverse problems, e.g., image compressive sensing (CS), as they can effectively leverage non-local self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various ringing artifacts due to naive patch aggregation. Using either approach alone usually limits performance and generalizability in image restoration tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely external and internal, deep and shallow, and local and nonlocal priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for image CS. To make the optimization tractable, a simple yet effective algorithm is proposed to solve the proposed H-PnP based image CS problem. Extensive experimental results demonstrate that the proposed H-PnP algorithm significantly outperforms the state-of-the-art techniques for image CS recovery such as SCSNet and WNNM. Zhiyuan Zha, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Bihan Wen, Ce Zhu |
ICIP | 2 |
| 2020 | Reconciliation Of Group Sparsity And Low-Rank Models For Image RestorationabstractImage nonlocal self-similarity (NSS) property has been widely exploited via various sparsity models such as joint sparsity (JS) and group sparse coding (GSC). However, the existing NSS-based sparsity models are either too restrictive, i.e., JS enforces the sparse codes to share the same support, or too general, i.e., GSC imposes only plain sparsity on the group coefficients, which limit their effectiveness for modeling real images. In this paper, we propose a novel NSS-based sparsity model, namely low-rank regularized group sparse coding (LR-GSC), to bridge the gap between the popular GSC and JS. The proposed LR-GSC model simultaneously exploits the sparsity and low-rankness of the dictionary-domain coefficients for each group of similar patches. To make the proposed scheme tractable and robust, an alternating minimization with an adaptive adjusted parameter strategy is developed to solve the proposed optimization problem. Experimental results on both image deblocking and denoising demonstrate that the proposed LR-GSC image restoration algorithms outperform many popular or state-of-the-art methods, in terms of both the objective and perceptual quality. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu |
ICME | 3 |
| 2020 | Edge Compression: An Integrated Framework for Compressive Imaging Processing on CAVsabstractMachine vision is the key to the successful deployment of many Advanced Driver Assistant System (ADAS) / Automated Driving System (ADS) functions, which require accurate high-resolution video processing in a real-time manner. Conventional approaches are either to reduce the frame rate or reduce the related frame size of the conventional camera videos, which lead to undesired consequences such as losing informative high-speed information and/or small objects in the video frames.Unlike conventional cameras, Compressive Imaging (CI) cameras are the promising implications of Compressive Sensing, which is an emerging field with the revelation that the optical domain compressed signal (a small number of linear projections of the original video image data) contains sufficient high-speed information for reconstruction and processing. Yet, CI cameras usually need complicated algorithms to retrieve the desired signal, leading to the corresponding high energy consumption. In this paper, we take a step further to the real applications of CI cameras in connected and autonomous vehicles (CAVs), with the primary goal of accelerating accurate video analysis and decreasing energy consumption. We propose a novel Vehicle Edge Server-Cloud closed-loop framework called Edge Compression for CI processing on CAVs. Our comprehensive experiments with four public datasets demonstrate that the detection accuracy of the compressed video images (named measurements) generated by the CI camera is close to the accuracy on reconstructed videos and comparable to the true value, which paves the way of applying CI in CAVs. Finally, six important observations with supporting evidence and analysis are presented to provide practical implications for researchers and domain experts. The code to reproduce our results is available at https://www.thecarlab.oryoutcomes/software. Sidi Lu, Xin Yuan 0002, Weisong Shi |
SEC | 2 |
| 2020 | Shearlet Enhanced Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) is a promising approach to capture high-dimensional data with low dimensional sensors. With modest modifications to off-the-shelf cameras, SCI cameras encode multiple frames into a single measurement frame. These correlated frames can then be retrieved by reconstruction algorithms. Existing reconstruction algorithms suffer from low speed or low fidelity. In this paper, we propose a novel reconstruction algorithm, namely, Shearlet enhanced Snapshot Compressive Imaging (SeSCI), which exploits the sparsity of the image representation in both frequency domain and shearlet domain. Towards this end, we first derive our SeSCI algorithm under the alternating direction method of multipliers (ADMM) framework. We then propose an efficient solution of SeSCI algorithm. Moreover, we prove that the improved SeSCI algorithm converges to a fixed point. Experimental results on both synthetic data and real data captured by SCI cameras demonstrate the significant advantages of SeSCI, which outperforms the conventional algorithms by more than 2dB in PSNR. At the same time, the SeSCI achieves a speed-up more than 100× over the state-of-the-art algorithm. Peihao Yang, Linghe Kong, Xiao-Yang Liu, Xin Yuan 0002, Guihai Chen |
IEEE Trans. Image Process. | 4 |
| 2020 | Group Sparsity Residual Constraint With Non-Local Priors for Image RestorationabstractGroup sparse representation (GSR) has made great strides in image restoration producing superior performance, realized through employing a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. However, due to some form of degradation (e.g., noise, down-sampling or pixels missing), traditional GSR models may fail to faithfully estimate sparsity of each group in an image, thus resulting in a distorted reconstruction of the original image. This motivates us to design a simple yet effective model that aims to address the above mentioned problem. Specifically, we propose group sparsity residual constraint with nonlocal priors (GSRC-NLP) for image restoration. Through introducing the group sparsity residual constraint, the problem of image restoration is further defined and simplified through attempts at reducing the group sparsity residual. Towards this end, we first obtain a good estimation of the group sparse coefficient of each original image group by exploiting the image nonlocal self-similarity (NSS) prior along with self-supervised learning scheme, and then the group sparse coefficient of the corresponding degraded image group is enforced to approximate the estimation. To make the proposed scheme tractable and robust, two algorithms, i.e., iterative shrinkage/thresholding (IST) and alternating direction method of multipliers (ADMM), are employed to solve the proposed optimization problems for different image restoration tasks. Experimental results on image denoising, image inpainting and image compressive sensing (CS) recovery, demonstrate that the proposed GSRC-NLP based image restoration algorithm is comparable to state-of-the-art denoising methods and outperforms several state-of-the-art image inpainting and image CS recovery methods in terms of both objective and perceptual quality metrics. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2020 | From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image RestorationabstractIn this paper, we propose a novel approach for the rank minimization problem, termed rank residual constraint (RRC). Different from existing low-rank based approaches, such as the well-known nuclear norm minimization (NNM) and the weighted nuclear norm minimization (WNNM), which estimate the underlying low-rank matrix directly from the corrupted observation, we progressively approximate (approach) the underlying low-rank matrix via minimizing the rank residual. Through integrating the image nonlocal self-similarity (NSS) prior with the proposed RRC model, we apply it to image restoration tasks, including image denoising and image compression artifacts reduction. Toward this end, we first obtain a good reference of the original image groups by using the image NSS prior, and then the rank residual of the image groups between this reference and the degraded image is minimized to achieve a better estimate to the desired image. In this manner, both the reference and the estimated image in each iteration are improved gradually and jointly. Based on the group-based sparse representation model, we further provide a theoretical analysis on the feasibility of the proposed RRC model. Experimental results demonstrate that the proposed RRC model outperforms many state-of-the-art schemes in both the objective and perceptual qualities. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2020 | A Benchmark for Sparse Coding: When Group Sparsity Meets Rank MinimizationabstractSparse coding has achieved a great success in various image processing tasks. However, a benchmark to measure the sparsity of image patch/group is missing since sparse coding is essentially an NP-hard problem. This work attempts to fill the gap from the perspective of rank minimization. We firstly design an adaptive dictionary to bridge the gap between group-based sparse coding (GSC) and rank minimization. Then, we show that under the designed dictionary, GSC and the rank minimization problems are equivalent, and therefore the sparse coefficients of each patch group can be measured by estimating the singular values of each patch group. We thus earn a benchmark to measure the sparsity of each patch group because the singular values of the original image patch groups can be easily computed by the singular value decomposition (SVD). This benchmark can be used to evaluate performance of any kind of norm minimization methods in sparse coding through analyzing their corresponding rank minimization counterparts. Towards this end, we exploit four well-known rank minimization methods to study the sparsity of each patch group and the weighted Schatten p-norm minimization (WSNM) is found to be the closest one to the real singular values of each patch group. Inspired by the aforementioned equivalence regime of rank minimization and GSC, WSNM can be translated into a non-convex weighted ℓp-norm minimization problem in GSC. By using the earned benchmark in sparse coding, the weighted ℓp-norm minimization is expected to obtain better performance than the three other norm minimization methods, i.e., ℓ1-norm, ℓp-norm and weighted ℓ1-norm. To verify the feasibility of the proposed benchmark, we compare the weighted ℓp-norm minimization against the three aforementioned norm minimization methods in sparse coding. Experimental results on image restoration applications, namely image inpainting and image compressive sensing recovery, demonstrate that the proposed scheme is feasible and outperforms many state-of-the-art methods. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2020 | Image Restoration Using Joint Patch-Group-Based Sparse RepresentationabstractSparse representation has achieved great success in various image processing and computer vision tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models lean to produce over-smooth effects. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides an effective mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR to image restoration tasks, including image inpainting and image deblocking. An iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR based image restoration problems. Experimental results demonstrate that the proposed JPG-SR is effective and outperforms many state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2020 | Image Restoration via Simultaneous Nonlocal Self-Similarity PriorsabstractThrough exploiting the image nonlocal self-similarity (NSS) prior by clustering similar patches to construct patch groups, recent studies have revealed that structural sparse representation (SSR) models can achieve promising performance in various image restoration tasks. However, most existing SSR methods only exploit the NSS prior from the input degraded (internal) image, and few methods utilize the NSS prior from external clean image corpus; how to jointly exploit the NSS priors of internal image and external clean image corpus is still an open problem. In this paper, we propose a novel approach for image restoration by simultaneously considering internal and external nonlocal self-similarity (SNSS) priors that offer mutually complementary information. Specifically, we first group nonlocal similar patches from images of a training corpus. Then a group-based Gaussian mixture model (GMM) learning algorithm is applied to learn an external NSS prior. We exploit the SSR model by integrating the NSS priors of both internal and external image data. An alternating minimization with an adaptive parameter adjusting strategy is developed to solve the proposed SNSS-based image restoration problems, which makes the entire algorithm more stable and practical. Experimental results on three image restoration applications, namely image denoising, deblocking and deblurring, demonstrate that the proposed SNSS produces superior results compared to many popular or state-of-the-art methods in both objective and perceptual quality measurements. Zhiyuan Zha, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu, Bihan Wen |
IEEE Trans. Image Process. | 2 |
| 2020 | Image Compression Based on Compressive Sensing: End-to-End Comparison With JPEGabstractWe present an end-to-end image compression system based on compressive sensing. The presented system integrates the conventional scheme of compressive sampling (on the entire image) and reconstruction with quantization and entropy coding. The compression performance, in terms of decoded image quality versus data rate, is shown to be comparable with JPEG and significantly better at the low rate range. We study the parameters that influence the system performance, including (i) the choice of sensing matrix, (ii) the trade-off between quantization and compression ratio, and (iii) the reconstruction algorithms. We propose an effective method to select, among all possible combinations of quantization step and compression ratio, the ones that yield the near-best quality at any given bit rate. Furthermore, our proposed image compression system can be directly used in the compressive sensing camera, e.g., the single pixel camera, to construct a hardware compressive sampling system. Xin Yuan 0002, Raziel Haimi-Cohen |
IEEE Trans. Multim. | 1 |
| 2019 | Deep Tensor ADMM-Net for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) systems have been developed to capture high-dimensional (≥ 3) signals using low-dimensional off-the-shelf sensors, i.e., mapping multiple video frames into a single measurement frame. One key module of a SCI system is an accurate decoder that recovers the original video frames. However, existing model-based decoding algorithms require exhaustive parameter tuning with prior knowledge and cannot support practical applications due to the extremely long running time. In this paper, we propose a deep tensor ADMM-Net for video SCI systems that provides high-quality decoding in seconds. Firstly, we start with a standard tensor ADMM algorithm, unfold its inference iterations into a layer-wise structure, and design a deep neural network based on tensor operations. Secondly, instead of relying on a pre-specified sparse representation domain, the network learns the domain of low-rank tensor through stochastic gradient descent. It is worth noting that the proposed deep tensor ADMM-Net has potentially mathematical interpretations. On public video data, the simulation results show the proposed method achieves average 0.8 ~ 2.5 dB improvement in PSNR and 0.07 ~ 0.1 in SSIM, and 1500× ~ 3600× speedups over the state-of-the-art methods. On real data captured by SCI cameras, the experimental results show comparable visual results with the state-of-the-art methods but in much shorter running time. Jiawei Ma, Xiao-Yang Liu, Xin Yuan 0002 |
ICCV | 4 |
| 2019 | lambda-Net: Reconstruct Hyperspectral Images From a Snapshot MeasurementabstractWe propose the λ-net, which reconstructs hyperspectral images (e.g., with 24 spectral channels) from a single shot measurement. This task is usually termed snapshot compressive-spectral imaging (SCI), which enjoys low cost, low bandwidth and high-speed sensing rate via capturing the three-dimensional (3D) signal i.e., (x, y, λ), using a 2D snapshot. Though proposed more than a decade ago, the poor quality and low-speed of reconstruction algorithms preclude wide applications of SCI. To address this challenge, in this paper, we develop a dual-stage generative model to reconstruct the desired 3D signal in SCI, dubbed λ-net. Results on both simulation and real datasets demonstrate the significant advantages of λ-net, which leads to >4dB improvement in PSNR for real-mask-in-the-loop simulation data compared to the current state-of-the-art. Furthermore, λ-net can finish the reconstruction task within sub-seconds instead of hours taken by the most recently proposed DeSCI algorithm, thus speeding up the reconstruction >1000 times. Xin Yuan 0002, Yunchen Pu, Vassilis Athitsos |
ICCV | 2 |
| 2019 | Simultaneous Nonlocal Self-Similarity Prior for Image DenoisingabstractNonlocal image representation has achieved great success in various image processing tasks such as image denoising, image deblurring and image deblocking. Particularly, by exploiting the image nonlo-cal self-similarity (NSS) prior, many nonlocal similar patches can be searched across the whole image for a given patch, which has significantly boosted the performance of image restoration. To the best of our knowledge, most existing methods only consider the NSS prior of the input degraded image, while few methods exploit the NSS prior from external clean image corpus. However, how to utilize the NSS priors of input degraded image and external clean image corpus simultaneously is still an open problem. In this paper, we propose a novel approach for image denoising, which exploits simultaneous nonlocal self-similarity (SNSS) by integrating the NSS priors of both the input degraded image and external clean image corpus. Firstly, we search and group nonlocal similar patches from a clean image corpus, and a group-based Gaussian Mixture Model (GMM) learning algorithm is developed to learn an external NSS prior. Then, an optimal group is selected from the best suitable Gaussian component for a group of the noisy image. By integrating the group of the noisy image and the corresponding group of the Gaussian component with a low-rank constraint, an iterative algorithm is developed to solve the proposed SNSS model. Experimental results demonstrate that the proposed SNSS-based denoising method produces superior results compared with many state-of-the-art denoising methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
ICIP | 2 |
| 2019 | Solving linear inverse problems using generative modelsabstractCompressed sensing (CS) algorithms recover a signal from its under-determined linear measurements via exploiting its structure. Starting from sparsity, recovery methods have steadily moved towards more complex structures. Emerging machine learning tools, e.g., generative models that are based on neural nets, potentially learn general complex structures from training data. Inspired by the success of such models in various computer vision tasks, researchers in CS have recently started to employ them to design efficient recovery methods. Consider a generative model defined by function g : Uk→ Rn, where U denotes a bounded subset of R. Assume that the function g is trained such that it can describe the class of desired signals Q c Rn. The standard problem in noiseless CS is to recover x ∈ Q from under-determined linear measurements y = Ax, where y ∈ Rmand mu∈Uk||g(u) - x||. Finally, using projected gradient descent to solve the aforementioned optimization, some preliminary numerical results are reported. Shirin Jalali, Xin Yuan 0002 |
ISIT | 2 |
| 2019 | Rank Minimization for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) refers to compressive imaging systems where multiple frames are mapped into a single measurement, with video compressive imaging and hyperspectral compressive imaging as two representative applications. Though exciting results of high-speed videos and hyperspectral images have been demonstrated, the poor reconstruction quality precludes SCI from wide applications. This paper aims to boost the reconstruction quality of SCI via exploiting the high-dimensional structure in the desired signal. We build a joint model to integrate the nonlocal self-similarity of video/hyperspectral frames and the rank minimization approach with the SCI sensing process. Following this, an alternating minimization algorithm is developed to solve this non-convex problem. We further investigate the special structure of the sampling process in SCI to tackle the computational workload and memory issues in SCI reconstruction. Both simulation and real data (captured by four different SCI cameras) results demonstrate that our proposed algorithm leads to significant improvements compared with current state-of-the-art algorithms. We hope our results will encourage the researchers and engineers to pursue further in compressive imaging for real applications. Yang Liu 0146, Xin Yuan 0002, Jin-Li Suo, David J. Brady, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Snapshot Compressed Sensing: Performance Bounds and AlgorithmsabstractSnapshot compressed sensing (CS) refers to compressive imaging systems in which multiple frames are mapped into a single measurement frame. Each pixel in the acquired frame is a noisy linear mapping of the corresponding pixels in the frames that are combined together. While the problem can be cast as a CS problem, due to the very special structure of the sensing matrix, standard CS theory cannot be employed to study such systems. In this paper, a compression-based framework is employed for theoretical analysis of snapshot CS systems. It is shown that this framework leads to two novel, computationally-efficient and theoretically-analyzable compression-based recovery algorithms. The proposed methods are iterative and employ compression codes to define and impose the structure of the desired signal. Theoretical convergence guarantees are derived for both algorithms. In the simulations, it is shown that, in the cases of both noise-free and noisy measurements, combining the proposed algorithms with a customized video compression code, designed to exploit nonlocal structures of video frames, significantly improves the state-of-the-art performance. Shirin Jalali, Xin Yuan 0002 |
IEEE Trans. Inf. Theory | 2 |
| 2018 | Joint Patch-Group Based Sparse Representation for Image InpaintingabstractSparse representation has achieved great successes in various machine learning and image processing tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models produce over-smooth phenomena. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR model to a low-level vision problem, namely, image inpainting. To make the proposed scheme tractable and robust, an iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR model. Experimental results demonstrate that the proposed model is efficient and outperforms several state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Ce Zhu |
ACML | 2 |
| 2018 | Nonlocal Low-Rank Tensor Factor Analysis for Image RestorationabstractLow-rank signal modeling has been widely leveraged to capture non-local correlation in image processing applications. We propose a new method that employs low-rank tensor factor analysis for tensors generated by grouped image patches. The low-rank tensors are fed into the alternative direction multiplier method (ADMM) to further improve image reconstruction. The motivating application is compressive sensing (CS), and a deep convolutional architecture is adopted to approximate the expensive matrix inversion in CS applications. An iterative algorithm based on this low-rank tensor factorization strategy, called NLR-TFA, is presented in detail. Experimental results on noiseless and noisy CS measurements demonstrate the superiority of the proposed approach, especially at low CS sampling rates. Xinyuan Zhang 0001, Xin Yuan 0002, Lawrence Carin |
CVPR | 2 |
| 2018 | Group Sparsity Residual with Non-Local Samples for Image DenoisingabstractInspired by group-based sparse coding, recently proposed group sparsity residual (GSR) scheme has demonstrated superior performance in image processing. However, one challenge in GSR is to estimate the residual by using a proper reference of the group-based sparse coding (GSC), which is desired to be as close to the truth as possible. Previous researches utilized the estimations from other algorithms (i.e., GMM or BM3D), which are either not accurate or too slow. In this paper, we propose to use the Non-Local Samples (NL-S) as reference in the GSR regime for image denoising, thus termed GSR-NLS. More specifically, we first obtain a good estimation of the group sparse coefficients by the image nonlocal self-similarity, and then solve the GSR model by an effective iterative shrinkage algorithm. Experimental results demonstrate that the proposed GSR-NLS not only outperforms many state-of-the-art methods, but also delivers the competitive advantage of speed. Zhiyuan Zha, Xinggan Zhang, Qiong Wang 0002, Yechao Bai, Lan Tang, Xin Yuan 0002 |
ICASSP | 6 |
| 2018 | Compressive Imaging Via One-Shot MeasurementsabstractOne-shot measurement systems combine multiple frames of a signal such as a video into a single frame of the same dimensions. Each element (pixel) in the measured frame is a linear combination of the corresponding elements (pixels) in the combined frames. Such systems are a crucial part of various modern compressive imaging systems with applications ranging from high-speed videos to high-dimensional medical images. In this paper, employing ideas from compression-based compressed sensing, a new theoretical framework for such one-shot compressive imaging systems is proposed. This new framework enables us to show that it is possible to recover frames that are combined under the one-shot measurement paradigm. The number of frames that can be combined and later deconvolved is connected with the level of structured-ness of the original multi-frame signal. This theoretical analysis aims at filling the gap between existing practical compressive imaging systems and traditional compressed sensing theory developed for random and dense sensing matrices. Shirin Jalali, Xin Yuan 0002 |
ISIT | 2 |
| 2018 | Non-convex weighted ℓp nuclear norm based ADMM framework for image restoration
Zhiyuan Zha, Xinggan Zhang, Yu Wu 0022, Qiong Wang 0002, Xin Liu 0012, Lan Tang, Xin Yuan 0002 |
Neurocomputing | 7 |
| 2018 | Adaptive step-size iterative algorithm for sparse signal recovery
Xin Yuan 0002 |
Signal Process. | 1 |
| 2018 | A New Nested MIMO Array With Increased Degrees of Freedom and Hole-Free Difference CoarrayabstractWe propose anew antenna array design approach for a multiple-input and multiple-output (MIMO) radar, which has closed-form expressions for the sensor locations and the number of achievable degrees of freedom (DOFs). This new approach utilizes the nested array as transmitting and receiving arrays. We employ the difference coarray of the sum coarray (DCSC) of the MIMO radar to obtain more DOFs for direction-of-arrival (DOA) estimation. Via properly designing the interelement spacings of the transmitting and receiving arrays, we can obtain a hole-free DCSC. The characteristics of array geometries are analyzed and the optimal numbers of sensors in transmitting/receiving antenna array are derived when given the total number of physical sensors. Simulations are conducted to demonstrate the advantages of the proposed array in terms of the number of DOFs, the number of resolvable sources, and the DOA estimation performance over the coprime MIMO array. Minglei Yang 0001, Xin Yuan 0002, Baixiao Chen |
IEEE Signal Process. Lett. | 3 |
| 2017 | Hyperspectral image super-resolution via convolutional neural networkabstractDue to the tradeoff between spatial and spectral resolution in remote sensing imaging, hyperspectral images are often acquired with a relative low spatial resolution, which limits their applications in many areas. Inspired by recent achievements in convolutional neural network (CNN) based super resolution (SR), a novel CNN based framework is constructed for SR of hyperspectral images by considering both spatial context and spectral correlation. As a result, the spectral distortion incurred by directly applying traditional SR algorithms to hyperspectral images is alleviated. Experimental results on several benchmark hyperspectral datasets have demonstrated that higher quality of reconstruction and spectral fidelity can be achieved, compared to band-wise manner based algorithms. Shaohui Mei, Xin Yuan 0002, Jingyu Ji, Shuai Wan, Junhui Hou, Qian Du 0001 |
ICIP | 2 |
| 2017 | Block-wise lensless compressive cameraabstractThe existing lensless compressive camera (L2C2) [1] suffers from low capture rates, resulting in low resolution images when acquired over a short time. In this work, we propose a new regime to mitigate these drawbacks. We replace the global-based compressive sensing used in the existing L2C2by the local block (patch) based compressive sensing. We use a single sensor for each block, rather than for the entire image, thus forming a multiple but spatially parallel sensor L2C2. This new camera retains the advantages of existing L2C2while leading to the additional benefits of fast acquisition, real time reconstruction, and the flexibility to image size. We develop multiple geometries of this block-wise L2C2in this paper. We have built prototypes of the proposed blockwise L2C2and demonstrated excellent results of real data. Xin Yuan 0002, Gang Huang 0003, Hong Jiang 0002, Paul A. Wilford |
ICIP | 1 |
| 2017 | Convolutional factor analysis inspired compressive sensingabstractWe solve the compressive sensing problem via convolutional factor analysis, where the convolutional dictionaries are learned in situ from the compressed measurements. An alternating direction method of multipliers (ADMM) paradigm for compressive sensing inversion based on convolutional factor analysis is developed. The proposed algorithm provides reconstructed images as well as features, which can be directly used for recognition (e.g., classification) tasks. We demonstrate that using ∼ 30% (relative to pixel numbers) compressed measurements, the proposed model achieves the classification accuracy comparable to the original data on MNIST. Xin Yuan 0002, Yunchen Pu |
ICIP | 1 |
| 2016 | A Deep Generative Deconvolutional Image ModelabstractA deep generative model is developed for representation and analysis of images, based on a hierarchical convolutional dictionary-learning framework. Stochastic unpooling is employed to link consecutive layers in the model, yielding top-down image generation. A Bayesian support vector machine is linked to the top-layer features, yielding max-margin discrimination. Deep deconvolutional inference is employed when testing, to infer the latent features, and the top-layer features are connected with the max-margin classifier for discrimination tasks. The algorithm is efficiently trained via Monte Carlo expectation-maximization (MCEM), with implementation on graphical processor units (GPUs) for efficient large-scale learning, and fast testing. Excellent results are obtained on several benchmark datasets, including ImageNet, demonstrating that the proposed model achieves results that are highly competitive with similarly sized convolutional neural networks. Yunchen Pu, Xin Yuan 0002, Andrew Stevens 0005, Chunyuan Li, Lawrence Carin |
AISTATS | 2 |
| 2016 | A general framework for reconstruction and classification from compressive measurements with side informationabstractWe develop a general framework for compressive linear-projection measurements with side information. Side information is an additional signal correlated with the signal of interest. We investigate the impact of side information on classification and signal recovery from low-dimensional measurements. Motivated by real applications, two special cases of the general model are studied. In the first, a joint Gaussian mixture model is manifested on the signal and side information. The second example again employs a Gaussian mixture model for the signal, with side information drawn from a mixture in the exponential family. Theoretical results on recovery and classification accuracy are derived. The presence of side information is shown to yield improved performance, both theoretically and experimentally. Liming Wang 0004, Francesco Renna, Xin Yuan 0002, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
ICASSP | 3 |
| 2016 | A new array geometry for DOA estimation with enhanced degrees of freedomabstractThis work presents a new array geometry, which is capable of providing O(M2N2) degrees of freedom (DOF) using only MN physical sensors via utilizing the second-order statistics of the received data. This new array is composed of multiple, identical minimum redundancy subarrays, whose positions follow a minimum redundancy configuration. Thus the new array is a minimum redundancy array (MRA) of MRA subarrays, and is termed as nested MRA. The sensor positions, aperture length, and the number of DOF of the new array can be predicted if these parameters of MRA subarrays are given. Numerical simulations demonstrate the superiorities of the proposed array geometry in resolving more sources than sensors and DOA estimation. Minglei Yang 0001, Alexander M. Haimovich, Baixiao Chen, Xin Yuan 0002 |
ICASSP | 4 |
| 2016 | Generalized alternating projection based total variation minimization for compressive sensingabstractWe consider the total variation (TV) minimization problem used for compressive sensing and solve it using the generalized alternating projection (GAP) algorithm. Extensive results demonstrate the high performance of proposed algorithm on compressive sensing, including two dimensional images, hyperspectral images and videos. We further derive the Alternating Direction Method of Multipliers (ADMM) framework with TV minimization for video and hyperspectral image compressive sensing under the CACTI and CASSI framework, respectively. Connections between GAP and ADMM are also provided. Xin Yuan 0002 |
ICIP | 1 |
| 2016 | Compressive video microscope via structured illuminationabstractA compressive video microscope based on structured illumination is built. The source-side illumination coding scheme allows the emission photons being collected by the full aperture of the microscope objective, and thus is suitable for the fluorescence readout mode. A block-wise total variation algorithm has been proposed to address the mismatch between the illumination pattern size and the detector pixel size. Image sequences with a temporal compression ratio of 4:1 are demonstrated. Xin Yuan 0002 |
ICIP | 1 |
| 2016 | Variational Autoencoder for Deep Learning of Images, Labels and CaptionsabstractA novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is used as an image encoder; the CNN is used to approximate a distribution for the latent DGDN features/code. The latent code is also linked to generative models for labels (Bayesian support vector machine) or captions (recurrent neural network). When predicting a label/caption for a new image at test, averaging is performed across the distribution of latent codes; this is computationally efficient as a consequence of the learned CNN-based encoder. Since the framework is capable of modeling the image in the presence/absence of associated labels/captions, a new semi-supervised setting is manifested for CNN learning with images; the framework even allows unsupervised CNN learning, based on images alone. Yunchen Pu, Zhe Gan, Ricardo Henao, Xin Yuan 0002, Chunyuan Li, Andrew Stevens 0005, Lawrence Carin |
NIPS | 4 |
| 2016 | Classification and Reconstruction of High-Dimensional Signals From Low-Dimensional Features in the Presence of Side InformationabstractThis paper offers a characterization of fundamental limits on the classification and reconstruction of high-dimensional signals from low-dimensional features, in the presence of side information. We consider a scenario where a decoder has access both to linear features of the signal of interest and to linear features of the side information signal; while the side information may be in a compressed form, the objective is recovery or classification of the primary signal, not the side information. The signal of interest and the side information are each assumed to have (distinct) latent discrete labels; conditioned on these two labels, the signal of interest and side information are drawn from a multivariate Gaussian distribution that correlates the two. With joint probabilities on the latent labels, the overall signal-(side information) representation is defined by a Gaussian mixture model. By considering bounds to the misclassification probability associated with the recovery of the underlying signal label, and bounds to the reconstruction error associated with the recovery of the signal of interest itself, we then provide sharp sufficient and/or necessary conditions for these quantities to approach zero when the covariance matrices of the Gaussians are nearly low rank. These conditions, which are reminiscent of the well-known Slepian-Wolf and Wyner-Ziv conditions, are the function of the number of linear features extracted from signal of interest, the number of linear features extracted from the side information signal, and the geometry of these signals and their interplay. Moreover, on assuming that the signal of interest and the side information obey such an approximately low-rank model, we derive the expansions of the reconstruction error as a function of the deviation from an exactly low-rank model; such expansions also allow the identification of operational regimes, where the impact of side information on signal reconstruction is most relevant. Our framework, which offers a principled mechanism to integrate side information in high-dimensional data problems, is also tested in the context of imaging applications. In particular, we report state-of-theart results in compressive hyperspectral imaging applications, where the accompanying side information is a conventional digital photograph. Francesco Renna, Liming Wang 0004, Xin Yuan 0002, Jianbo Yang, Galen Reeves, A. Robert Calderbank, Lawrence Carin, Miguel R. D. Rodrigues |
IEEE Trans. Inf. Theory | 3 |
| 2015 | Multi-scale Bayesian reconstruction of compressive X-ray imageabstractA novel multi-scale dictionary based Bayesian reconstruction algorithm is proposed for compressive X-ray imaging, which encodes the material's spectrum by Poisson measurements. Inspired by recently developed compressive X-ray imaging systems [1], this work aims to recover the material's spectrum from the compressive coded image by leveraging a reference spectrum library. Instead of directly using the huge and redundant library as a dictionary, which is cumbersome in computation and difficult for selecting those active dictionary atoms, a multi-scale tree structured dictionary is refined from the spectrum library, and following this a Bayesian reconstruction algorithm is developed. Experimental results on real data demonstrate superior performance in comparison with traditional methods. Jiaji Huang, Xin Yuan 0002, A. Robert Calderbank |
ICASSP | 2 |
| 2015 | Collaborative compressive X-ray image reconstructionabstractThe Poisson Factor Analysis (PFA) is applied to recover signals from a Poisson compressive sensing system. Motivated by the recently developed compressive X-ray imaging system, Coded Aperture Coherent Scatter Spectral Imaging (CACSSI) [1], we propose a new Bayesian reconstruction algorithm. The proposed Poisson-Gamma (PG) approach uses multiple measurements to refine our knowledge on both sensing matrix and background noise to overcome the uncertainties and inaccuracy of the hardware system. Therefore, a collaborative compressive X-ray image reconstruction algorithm is proposed under a Bayesian framework. Experimental results on real data show competitive performance in comparison with point estimation based methods. Jiaji Huang, Xin Yuan 0002, A. Robert Calderbank |
ICASSP | 2 |
| 2015 | Polynomial-phase signal direction-finding and source-tracking with a single acoustic vector sensorabstractThis paper introduces a new ESPRIT-based algorithm to estimate the direction-of-arrival of an arbitrary degree polynomial-phase signal with a single acoustic vector-sensor. The proposed time-invariant ESPRIT algorithm is based on a matrix-pencil pair derived from the time-delayed data-sets collected by a single acoustic vector-sensor. This approach requires neither a prior knowledge of the polynomial-phase signal's coefficients nor a prior knowledge of the polynomial-phase signal's frequency-spectrum. Furthermore, a preprocessing technique is proposed to incorporate the single-forgetting-factor algorithm and multiple-forgetting-factor adaptive tracking algorithm to track a polynomial-phase signal using one acoustic vector sensor. Simulation results verify the efficacy of the proposed direction finding and source tracking algorithms. Xin Yuan 0002, Jiaji Huang, A. Robert Calderbank |
ICASSP | 1 |
| 2015 | Non-Gaussian Discriminative Factor Models via the Max-Margin Rank-LikelihoodabstractWe consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a max-margin version of the rank-likelihood. A discriminative factor model is then developed, integrating the new max-margin rank-likelihood and (linear) Bayesian support vector machines, which are also built on the max-margin principle. The discriminative factor model is further extended to the nonlinear case through mixtures of local linear classifiers, via Dirichlet processes. Fully local conjugacy of the model yields efficient inference with both Markov Chain Monte Carlo and variational Bayes approaches. Extensive experiments on benchmark and real data demonstrate superior performance of the proposed model and its potential for applications in computational biology. Xin Yuan 0002, Ricardo Henao, Ephraim Tsalik, Raymond Langley, Lawrence Carin |
ICML | 1 |
| 2015 | Classification and reconstruction of compressed GMM signals with side informationabstractThis paper offers a characterization of performance limits for classification and reconstruction of high-dimensional signals from noisy compressive measurements, in the presence of side information. We assume the signal of interest and the side information signal are drawn from a correlated mixture of distributions/components, where each component associated with a specific class label follows a Gaussian mixture model (GMM). We provide sharp sufficient and/or necessary conditions for the phase transition of the misclassification probability and the reconstruction error in the low-noise regime. These conditions, which are reminiscent of the well-known Slepian-Wolf and Wyner-Ziv conditions, are a function of the number of measurements taken from the signal of interest, the number of measurements taken from the side information signal, and the geometry of these signals and their interplay. Francesco Renna, Liming Wang 0004, Xin Yuan 0002, Jianbo Yang, Galen Reeves, A. Robert Calderbank, Lawrence Carin, Miguel R. D. Rodrigues |
ISIT | 3 |
| 2015 | A concentration-of-measure inequality for multiple-measurement modelsabstractClassical compressive sensing typically assumes a single measurement, and theoretical analysis often relies on corresponding concentration-of-measure results. There are many real-world applications involving multiple compressive measurements, from which the underlying signals may be estimated. In this paper, we establish a new concentration-of-measure inequality for a block-diagonal structured random compressive sensing matrix with Rademacher-ensembles. We discuss applications of this newly-derived inequality to two appealing compressive multiple-measurement models: for Gaussian and Poisson systems. In particular, Johnson-Lindenstrauss-type results and a compressed-domain classification result are derived for a Gaussian multiple-measurement model. We also propose, as another contribution, theoretical performance guarantees for signal recovery for multi-measurement Poisson systems, via the inequality. Liming Wang 0004, Jiaji Huang, Xin Yuan 0002, Volkan Cevher, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
ISIT | 3 |
| 2015 | Signal Recovery and System Calibration from Multiple Compressive Poisson MeasurementsabstractThe measurement matrix employed in compressive sensing typically cannot be known precisely a priori and must be estimated via calibration. One may take multiple compressive measurements, from which the measurement matrix and underlying signals may be estimated jointly. This is of interest as well when the measurement matrix may change as a function of the details of what is measured. This problem has been considered recently for Gaussian measurement noise, and here we develop this idea with application to Poisson systems. A collaborative maximum likelihood algorithm and alternating proximal gradient algorithm are proposed, and associated theoretical performance guarantees are established based on newly derived concentration-of-measure results. A Bayesian model is then introduced, to improve flexibility and generality. Connections between the maximum likelihood methods and the Bayesian model are developed, and example results are presented for a real compressive X-ray imaging system. Liming Wang 0004, Jiaji Huang, Xin Yuan 0002, Kalyani Krishnamurthy, Joel A. Greenberg, Volkan Cevher, Miguel R. D. Rodrigues, David J. Brady, A. Robert Calderbank, Lawrence Carin |
SIAM J. Imaging Sci. | 3 |
| 2015 | Compressive Sensing by Learning a Gaussian Mixture Model From MeasurementsabstractCompressive sensing of signals drawn from a Gaussian mixture model (GMM) admits closed-form minimum mean squared error reconstruction from incomplete linear measurements. An accurate GMM signal model is usually not available a priori, because it is difficult to obtain training signals that match the statistics of the signals being sensed. We propose to solve that problem by learning the signal model in situ, based directly on the compressive measurements of the signals, without resorting to other signals to train a model. A key feature of our method is that the signals being sensed are treated as random variables and are integrated out in the likelihood. We derive a maximum marginal likelihood estimator (MMLE) that maximizes the likelihood of the GMM of the underlying signals given only their linear compressive measurements. We extend the MMLE to a GMM with dominantly low-rank covariance matrices, to gain computational speedup. We report extensive experimental results on image inpainting, compressive sensing of high-speed video, and compressive hyperspectral imaging (the latter two based on real compressive cameras). The results demonstrate that the proposed methods outperform state-of-the-art methods by significant margins. Jianbo Yang, Xuejun Liao, Xin Yuan 0002, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin |
IEEE Trans. Image Process. | 3 |
| 2014 | Low-Cost Compressive Sensing for Color Video and DepthabstractA simple and inexpensive (low-power and low-bandwidth) modification is made to a conventional off-the-shelf color video camera, from which we recover multiple color frames for each of the original measured frames, and each of the recovered frames can be focused at a different depth. The recovery of multiple frames for each measured frame is made possible via high-speed coding, manifested via translation of a single coded aperture, the inexpensive translation is constituted by mounting the binary code on a piezoelectric device. To simultaneously recover depth information, a liquid lens is modulated at high speed, via a variable voltage. Consequently, during the aforementioned coding process, the liquid lens allows the camera to sweep the focus through multiple depths. In addition to designing and implementing the camera, fast recovery is achieved by an anytime algorithm exploiting the group-sparsity of wavelet/DCT coefficients. Xin Yuan 0002, Patrick Llull, Xuejun Liao, Jianbo Yang, David J. Brady, Guillermo Sapiro, Lawrence Carin |
CVPR | 1 |
| 2014 | Bayesian Nonlinear Support Vector Machines and Discriminative Factor Modeling
Ricardo Henao, Xin Yuan 0002, Lawrence Carin |
NIPS | 2 |
| 2014 | Real-time air quality estimation based on color image processingabstractThis paper address the problem of efficient, realtime estimation of the particulate mass concentration, exactly PM2.5 (particles with aerodynamic diameters less than 2.5 μm) from a superb view image. And the proposed method is to achieve high degree of accuracy at the cost of only modest user's effort by analyzing the relationship between the PM2.5 and the degradation of the observed image. With the fitting algorithm with experimental data, the PM2.5 could be real-time estimated by a general camera with little artificial participation, and the correlation coefficient produced by our data set and the standard observation will be as high as 0.8219, as the MSE (Mean Squared Error) value 51.2324 μg/m3. Haoqian Wang, Xin Yuan 0002, Xingzheng Wang, Yongbing Zhang 0002, Qionghai Dai |
VCIP | 2 |
| 2014 | Coherent sources direction finding and polarization estimation with various compositions of spatially spread polarized antenna arrays
Xin Yuan 0002 |
Signal Process. | 1 |
| 2014 | Video Compressive Sensing Using Gaussian Mixture ModelsabstractA Gaussian mixture model (GMM)-based algorithm is proposed for video reconstruction from temporally compressed video measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The GMM-based inversion method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed inversion method with videos reconstructed from simulated compressive video measurements, and from a real compressive video camera. We also use the GMM as a tool to investigate adaptive video compressive sensing, i.e., adaptive rate of temporal compression. Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin |
IEEE Trans. Image Process. | 2 |
| 2013 | Gaussian mixture model for video compressive sensingabstractA Gaussian Mixture Model (GMM)-based algorithm is proposed for video reconstruction from temporal compressed measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The developed GMM reconstruction method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed GMM with videos reconstructed from simulated compressive video measurements and from a real compressive video camera. Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, Guillermo Sapiro, David J. Brady, Lawrence Carin |
ICIP | 2 |
| 2013 | Adaptive temporal compressive sensing for videoabstractThis paper introduces the concept of adaptive temporal compressive sensing (CS) for video. We propose a CS algorithm to adapt the compression ratio based on the scene's temporal complexity, computed from the compressed data, without compromising the quality of the reconstructed video. The temporal adaptivity is manifested by manipulating the integration time of the camera, opening the possibility to realtime implementation. The proposed algorithm is a generalized temporal CS approach that can be incorporated with a diverse set of existing hardware systems. Xin Yuan 0002, Jianbo Yang, Patrick Llull, Xuejun Liao, Guillermo Sapiro, David J. Brady, Lawrence Carin |
ICIP | 1 |
| 2012 | Polynomial-phase signal source tracking using an electromagnetic vector-sensorabstractA pre-processing technique is developed to track a polynomial-phase signal using an electromagnetic vector-sensor, which can be collocated or spatially-spread. The performance of the single-forgetting-factor algorithm incorporating the proposed pre-processing approach is improved significantly, and it even surpasses the performance of the multiple-forgetting-factor algorithm in a polynomial-phase source scenario. Simulation results verify the efficacy of the proposed technique. Xin Yuan 0002 |
ICASSP | 1 |