Xiaodong Wang 0026

dblp:07/1021-26 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-0331-362XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging
Ping Wang 0029, Lishun Wang, Gang Qu 0005, Xiaodong Wang 0026, Yulun Zhang 0001, Xin Yuan 0002
CVPR4
2025 SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive Image
abstract
Snapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality limitations are general issues that have limited wider adoption. This paper introduces SCI-Gaussian, the first 3D-aware SCI reconstruction based on 3D Gaussian splatting (3D-GS). This method utilizes an explicit 3D representation to achieve efficient and high-quality scene reconstruction. The motivation stems from the highly efficient representation and surprising quality of 3D-GS, despite when applied to SCI system, it encounters difficulties in generating point initialization for explicit Gaussians and accurate pose recovery from a single SCI measured image. Specifically, we effectively initialize these Gaussians through sampling a coarsely trained NeRF at various hash structures, then model the physical formation of the SCI measurement and jointly optimize Gaussians and camera trajectories with a bundle adjustment formulation during exposure time. Extensive experiments on synthetic and real-world datasets demonstrate that SCI-Gaussian outperforms the state-of-the-art (SOTA) methods, achieving comparable or better results with significantly 10× faster training and 1000× faster rendering speed than the most recent NeRF-based method.
Xiaodong Wang 0026, Xin Yuan 0002, Mark D. Butala, Gaoang Wang
ICASSP3
2025 Texture-aware Intrinsic Image Decomposition with Model- and Learning-based Priors
abstract
This paper aims to recover the intrinsic reflectance layer and shading layer given a single image. Though this intrinsic image decomposition problem has been studied for decades, it remains a significant challenge in cases of complex scenes, i.e. spatially-varying lighting effect and rich textures. In this paper, we propose a novel method for handling severe lighting and rich textures in intrinsic image decomposition, which enables to produce high-quality intrinsic images for real-world images. Specifically, we observe that previous learning-based methods tend to produce texture-less and over-smoothing intrinsic images, which can be used to infer the lighting and texture information given a RGB image. In this way, we design a texture-guided regularization term and formulate the decomposition problem into an optimization framework, to separate the material textures and lighting effect. We demonstrate that combining the novel texture-aware prior can produce superior results to existing approaches. Code is available at https://github.com/xiaodongwo/Efficient IID.
Xiaodong Wang 0026, Zijun He, Xin Yuan 0002
ICME1
2025 Unfolding Framework with Complex-Valued Deformable Attention for High-Quality Computer-Generated Hologram Generation
abstract
Computer-generated holography (CGH) has gained wide attention with deep learning-based algorithms. However, due to its nonlinear and ill-posed nature, challenges remain in achieving accurate and stable reconstruction. Specifically, (i) the widely used end-to-end networks treat the reconstruction model as a black box, ignoring underlying physical relationships, which reduces interpretability and flexibility. (ii) CNN-based CGH algorithms have limited receptive fields, hindering their ability to capture long-range dependencies and global context. (iii) Angular spectrum method (ASM)-based models are constrained to finite near-fields. In this paper, we propose a Deep Unfolding Network (DUN) that decomposes gradient descent into two modules: an adaptive bandwidth-preserving model (ABPM) and a phase-domain complex-valued denoiser (PCD), providing more flexibility. ABPM allows for wider working distances compared to ASM-based methods. At the same time, PCD leverages its complex-valued deformable self-attention module to capture global features and enhance performance, achieving a PSNR over 35 dB. Experiments on simulated and real data show state-of-the-art results. Code is available at https://github.com/HannahZhang1926/Complex-Valued-Deformable-Transformer-for-CGH.
Haomiao Zhang, Zhangyuan Li, Yanling Piao, Xiaodong Wang 0026, Xiongfei Su, Xin Yuan 0002
ICME5
2025 Spectral Compressive Imaging via Chromaticity-Intensity Decomposition
abstract
In coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it difficult to recover the intrinsic spectral reflectance that remains invariant to lighting conditions. To address these challenges, we propose a chromaticity-intensity decomposition framework, which disentangles an HSI into a spatially smooth intensity map and a spectrally variant chromaticity cube. The chromaticity encodes lighting-invariant reflectance, enriched with high-frequency spatial details and local spectral sparsity. Building on this decomposition, we develop CIDNet—a Chromaticity-Intensity Decomposition unfolding network within a dual-camera CASSI system. CIDNet integrates a hybrid spatial-spectral Transformer tailored to reconstruct fine-grained and sparse spectral chromaticity and a degradation-aware, spatially-adaptive noise estimation module that captures anisotropic noise across iterative stages. Extensive experiments on both synthetic and real-world CASSI datasets demonstrate that our method achieves superior performance in both spectral and chromaticity fidelity. Code is released at: \url{https://github.com/xiaodongwo/CIDNet}.
Xiaodong Wang 0026, Zijun He, Ping Wang 0029, Lishun Wang, Xin Yuan 0002
NeurIPS1
2025 LCTC: Lightweight Convolutional Thresholding Sparse Coding Network Prior for Compressive Hyperspectral Imaging
abstract
Compressive spectral imaging has garnered significant attention for its ability to effectively enhance the captured spatial and spectral information. Predominant methods, based on compressive sensing, typically formulate the imaging task as a constrained optimization problem and rely on hand-crafted priors to model the sparsity of spectral images. However, these approaches often suffer from suboptimal performance due to the inherent difficulty of identifying an appropriate transform space where spectral images exhibit sparsity. To overcome this limitation, we propose a novel convolutional sparse coding-inspired untrained network prior for fast and adaptive identification of the sparse transform domain and compressible signal. Specifically, a Lightweight Convolutional Thresholding sparse Coding (LCTC) network is designed as the sparse transform domain, with its inputs interpreted as sparse coefficients. Crucially, both the transform domain and its coefficients are solved in a self-supervised learning manner. Furthermore, we demonstrate that LCTC prior can be seamlessly incorporated into the iterative optimization algorithm as a Plug-and-Play (PnP) regularization. Both the LCTC and PnP-LCTC exhibit superior performance compared to previous methods. Experiments under various scenarios validate the effectiveness and efficiency of our approach.
Yurong Chen 0003, Yaonan Wang 0001, Xiaodong Wang 0026, Xin Yuan 0002, Hui Zhang 0023
IEEE Trans. Image Process.3
2024 SCINeRF: Neural Radiance Fields from a Snapshot Compressive Image
abstract
In this paper, we explore the potential of Snapshot Compressive Imaging (SCI) technique for recovering the underlying 3D scene representation from a single temporal compressed image. SCI is a cost-effective method that enables the recording of high-dimensional data, such as hyperspectral or temporal information, into a single image using low- cost 2D imaging sensors. To achieve this, a series of specially designed 2D masks are usually employed, which not only reduces storage requirements but also offers potential privacy protection. Inspired by this, to take one step further, our approach builds upon the powerful 3D scene representation capabilities of neural radiance fields (NeRF). Specifically, we formulate the physical imaging process of SCI as part of the training of NeRF, allowing us to exploit its impressive performance in capturing complex scene structures. To assess the effectiveness of our method, we conduct extensive evaluations using both synthetic data and real data captured by our SCI system. Extensive experimental results demonstrate that our proposed approach sur- passes the state-of-the-art methods in terms of image re- construction and novel view image synthesis. Moreover, our method also exhibits the ability to restore high frame- rate multi-view consistent images by leveraging SCI and the rendering capabilities of NeRF. The code is available at https://github.com/WU-CVGL/SCINeRF.
Xiaodong Wang 0026, Ping Wang 0029, Xin Yuan 0002, Peidong Liu 0001
CVPR2