Ping Wang 0029

dblp:37/1304-29 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0001-2746-5102ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Breaking Measurement Barriers: From Compressed Sensing to Deep Reconstruction
abstract
Deep learning methods have achieved remarkable success in image compressed sensing (CS) task, namely reconstructing a high-fidelity image from its compressed measurement. However, existing methods are deficient in incoherent compressed measurement at sensing phase and implicit measurement representations at reconstruction phase, limiting the overall performance. In this work, we answer two questions: (i) how to improve the measurement incoherence for decreasing the ill-posedness; (ii) how to learn informative representations from measurements. To this end, we propose a novel asymmetric Kronecker CS (AKCS) model and theoretically present its better incoherence than previous Kronecker CS with minimal increase of complexity. Moreover, apart from the explicit measurement representations in gradient descent projection in unfolding networks, we further propose a measurement-aware cross attention (MACA) mechanism to learn implicit measurement representations. We integrate AKCS and MACA into a widely-used unfolding architecture to get a measurement-enhanced unfolding network (MEUNet). Extensive experiments demonstrate that the proposed MEUNet achieves state-of-the-art (SOTA) performance in reconstruction accuracy with high efficiency.
Gang Qu 0005, Ping Wang 0029, Siming Zheng, Xin Yuan 0002
AAAI2
2026 High-Speed FHD Full-Color Video Computer-Generated Holography
abstract
Computer-generated holography (CGH) is a promising technology for next-generation displays. However, generating high-speed, high-quality holographic video requires both high frame rate display and efficient computation, but is constrained by two key limitations: (i) Learning-based models often produce over-smoothed phases with narrow angular spectra, causing severe color crosstalk in high frame rate full-color displays such as depth-division multiplexing and thus resulting in a trade-off between frame rate and color fidelity. (ii) Existing frame-by-frame optimization methods typically optimize frames independently, neglecting spatial-temporal correlations between consecutive frames and leading to computationally inefficient solutions. To overcome these challenges, in this paper, we propose a novel high-speed full-color video CGH generation scheme. First, we introduce Spectrum-Guided Depth Division Multiplexing (SGDDM), which optimizes phase distributions via frequency modulation, enabling high-fidelity full-color display at high frame rates. Second, we present HoloMamba, a lightweight asymmetric Mamba-Unet architecture that explicitly models spatial-temporal correlations across video sequences to enhance reconstruction quality and computational efficiency. Extensive simulated and real-world experiments demonstrate that SGDDM achieves high-fidelity full-color display without compromise in frame rate, while HoloMamba generates FHD (1080p) full-color holographic video at over 260 FPS, more than 2.6 times faster than the prior state-of-the-art Divide-Conquer-and-Merge Strategy.
Haomiao Zhang, Yanling Piao, Zhangyuan Li, Ping Wang 0029, Xin Yuan 0002
AAAI8
2025 Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging
Ping Wang 0029, Lishun Wang, Gang Qu 0005, Xiaodong Wang 0026, Yulun Zhang 0001, Xin Yuan 0002
CVPR1
2025 Motion-Aware Reconstruction for Video Snapshot Compressive Imaging
abstract
Video Snapshot Compressive Imaging (SCI) provides an elegant solution for recording fast-dynamics motion through optical compression and computational reconstruction. However, existing video reconstruction algorithms are computationally intensive and typically require substantial computing and storage resources to large areas of static background with minimal information, leading to significant resource waste. To this end, this paper proposes a motion-aware reconstruction paradigm, in which moving objects are separated from the static background to identify Regions Of Interest (ROI) in the measurement domain and then video reconstruction is confined to ROI to save reconstruction costs. Specifically, a Motion Decomposition Network (MoDeNet) is proposed to get ROI by deep unfolding of Half-Quadratic Splitting (HQS). Subsequently, high-performance reconstruction network is used to recover ROI videos. Extensive experiments demonstrate that the proposed paradigm reduces the reconstruction costs of video SCI effectively and efficiently, and the proposed MoDeNet outperforms previous decomposition algorithms in terms of accuracy and speed.
Zhangyuan Li, Ping Wang 0029, Haomiao Zhang, Xin Yuan 0002
ICIP2
2025 Spectral Compressive Imaging via Chromaticity-Intensity Decomposition
abstract
In coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it difficult to recover the intrinsic spectral reflectance that remains invariant to lighting conditions. To address these challenges, we propose a chromaticity-intensity decomposition framework, which disentangles an HSI into a spatially smooth intensity map and a spectrally variant chromaticity cube. The chromaticity encodes lighting-invariant reflectance, enriched with high-frequency spatial details and local spectral sparsity. Building on this decomposition, we develop CIDNet—a Chromaticity-Intensity Decomposition unfolding network within a dual-camera CASSI system. CIDNet integrates a hybrid spatial-spectral Transformer tailored to reconstruct fine-grained and sparse spectral chromaticity and a degradation-aware, spatially-adaptive noise estimation module that captures anisotropic noise across iterative stages. Extensive experiments on both synthetic and real-world CASSI datasets demonstrate that our method achieves superior performance in both spectral and chromaticity fidelity. Code is released at: \url{https://github.com/xiaodongwo/CIDNet}.
Xiaodong Wang 0026, Zijun He, Ping Wang 0029, Lishun Wang, Xin Yuan 0002
NeurIPS3
2025 Hard EXIF: Protecting Image Authorship Through Metadata, Hardware, and Content
abstract
With the rapid proliferation of digital image content and advancements in image editing technologies, the protection of digital image authorship has become an increasingly important issue. Traditional methods for authorship protection include registering authorship through certification organization, utilizing image metadata such as Exchangeable Image File Format (EXIF) data, and employing watermarking techniques to prove ownership. In recent years, blockchain-based technologies have also been introduced to enhance authorship protection further. However, these approaches face challenges in balancing four key attributes: strong legal validity, high security, low cost, and high usability. Authorship registration is often cumbersome, EXIF metadata can be easily extracted and tampered with, watermarking techniques are vulnerable to various forms of attack, and blockchain technology is complex to implement and requires long-term maintenance. In response to these challenges, this paper introduces a new framework Hard EXIF, designed to balance these multiple attributes while delivering improved performance. The proposed method integrates metadata with physically unclonable functions (PUFs) for the first time, creating unique device fingerprints and embedding them into images using watermarking techniques. By leveraging the security and simplicity of hash functions and PUFs, this method enhances EXIF security while minimizing costs. Experimental results demonstrate that the Hard EXIF framework achieves an average peak signal-to-noise ratio (PSNR) of 42.89 dB, with a similarity of 99.46% between the original and watermarked images, and the extraction error rate is only 0.0017. These results show that the Hard EXIF framework balances legal validity, security, cost, and usability, promising authorship protection with great potential for wider application.
Yushu Zhang 0001, Xiangli Xiao, Ping Wang 0029, Wenying Wen
IEEE Trans. Image Process.5
2024 SCINeRF: Neural Radiance Fields from a Snapshot Compressive Image
abstract
In this paper, we explore the potential of Snapshot Compressive Imaging (SCI) technique for recovering the underlying 3D scene representation from a single temporal compressed image. SCI is a cost-effective method that enables the recording of high-dimensional data, such as hyperspectral or temporal information, into a single image using low- cost 2D imaging sensors. To achieve this, a series of specially designed 2D masks are usually employed, which not only reduces storage requirements but also offers potential privacy protection. Inspired by this, to take one step further, our approach builds upon the powerful 3D scene representation capabilities of neural radiance fields (NeRF). Specifically, we formulate the physical imaging process of SCI as part of the training of NeRF, allowing us to exploit its impressive performance in capturing complex scene structures. To assess the effectiveness of our method, we conduct extensive evaluations using both synthetic data and real data captured by our SCI system. Extensive experimental results demonstrate that our proposed approach sur- passes the state-of-the-art methods in terms of image re- construction and novel view image synthesis. Moreover, our method also exhibits the ability to restore high frame- rate multi-view consistent images by leveraging SCI and the rendering capabilities of NeRF. The code is available at https://github.com/WU-CVGL/SCINeRF.
Xiaodong Wang 0026, Ping Wang 0029, Xin Yuan 0002, Peidong Liu 0001
CVPR3
2024 Dual-Scale Transformer for Large-Scale Single-Pixel Imaging
abstract
Single-pixel imaging (SPI) is a potential computational imaging technique which produces image by solving an ill-posed reconstruction problem from few measurements captured by a single-pixel detector. Deep learning has achieved impressive success on SPI reconstruction. However, previ-ous poor reconstruction performance and impractical imaging model limit its real-world applications. In this paper, we propose a deep unfolding network with hybrid-attention Transformer on Kronecker SPI model, dubbed HATNet, to im-prove the imaging quality of real SPI cameras. Specifically, we unfold the computation graph of the iterative shrinkage-thresholding algorithm (ISTA) into two alternative modules: efficient tensor gradient descent and hybrid-attention multi-scale denoising. By virtue of Kronecker SPI, the gradient descent module can avoid high computational overheads rooted in previous gradient descent modules based on vector-ized SPI. The denoising module is an encoder-decoder archi-tecture powered by dual-scale spatial attention for high- and low-frequency aggregation and channel attention for global information recalibration. Moreover, we build a SPI proto-type to verify the effectiveness of the proposed method. Ex-tensive experiments on synthetic and real data demonstrate that our method achieves the state-of-the-art performance. The source code and pre-trained models are available at https://github.com/Gang-Qu/HATNet-SPI.
Gang Qu 0005, Ping Wang 0029, Xin Yuan 0002
CVPR2
2024 Hierarchical Separable Video Transformer for Snapshot Compressive Imaging
Ping Wang 0029, Yulun Zhang 0001, Lishun Wang, Xin Yuan 0002
ECCV (81)1
2023 Deep Optics for Video Snapshot Compressive Imaging
abstract
Video snapshot compressive imaging (SCI) aims to capture a sequence of video frames with only a single shot of a 2D detector, whose backbones rest in optical modulation patterns (also known as masks) and a computational reconstruction algorithm. Advanced deep learning algorithms and mature hardware are putting video SCI into practical applications. Yet, there are two clouds in the sunshine of SCI: i) low dynamic range as a victim of high temporal multiplexing, and ii) existing deep learning algorithms’ degradation on real system. To address these challenges, this paper presents a deep optics framework to jointly optimize masks and a reconstruction network. Specifically, we first propose a new type of structural mask to realize motionaware and full-dynamic-range measurement. Considering the motion awareness property in measurement domain, we develop an efficient network for video SCI reconstruction using Transformer to capture long-term temporal dependencies, dubbed Res2former. Moreover, sensor response is introduced into the forward model of video SCI to guarantee end-to-end model training close to real system. Finally, we implement the learned structural masks on a digital micro-mirror device. Experimental results on synthetic and real data validate the effectiveness of the proposed frame-work. We believe this is a milestone for real-world video SCI. The source code and data are available at https://github.com/pwangcs/DeepOpticsSCI.
Ping Wang 0029, Lishun Wang, Xin Yuan 0002
ICCV1
2023 SAUNet: Spatial-Attention Unfolding Network for Image Compressive Sensing
abstract
Image Compressive Sensing (CS) enables compressed capture of natural images via a spatial multiplexing camera and accurate reconstruction from few measurements via an advanced algorithm. Deep learning, especially deep unfolding, has recently achieved impressive success in image CS reconstruction. However, existing learning-based methods have been developed for block (usually with 33 X 33 pixels) CS instead of full image CS. Apart from the difficulties in hardware implementation, block CS breaks the global pixel interactions, limiting the overall performance. In this paper, we propose the first two-dimensional deep unfolding framework, and further develop a Spatial-Attention Unfolding Network (SAUNet) for full image CS reconstruction by alternately performing a spatially-adaptive gradient descent module and a cross-stage multi-scale denoising module. The gradient descent module has the spatial self-adaptation to the degradation of in-process image. The denoising module is a three-level U-shaped structure powered by Convolutional Self-Attention (CSA) mechanism. Inspired by Transformer, CSA is designed to adaptively aggregate spatially local information and adaptively recalibrate channel-wise global information with only normal convolutional operator. Extensive experiments demonstrate that SAUNet outperforms the state-of-the-art methods by a large margin. The source code and pre-trained models are available at https://github.com/pwangcs/SAUNet.
Ping Wang 0029, Xin Yuan 0002
ACM Multimedia1
2021 Privacy-Assured FogCS: Chaotic Compressive Sensing for Secure Industrial Big Image Data Processing in Fog Computing
abstract
In the age of the industrial big data, there are several significant problems such as high-overhead data acquisition, data privacy leakage, and data tampering. Fog computing capability is rapidly expanding to address not only network congestion issues but data security issues. This article presents a chaotic compressive sensing (CS) scheme for securely processing industrial big image data in the fog computing paradigm, called privacy-assured FogCS. Specially, the sine logistic modulation map is used to drive the privacy-assured, authenticated, and block CS for secure image data collection in the sensor nodes. After sampling, the measurements are normalized in the fog nodes. The normalized measurements can achieve the perfect secrecy and their energy values are further masked through the proposed permutation-diffusion architecture. Finally, these relevant data are transmitted to the clouds (data centers) for storage, reconstruction, and authentication if required. In addition, a hardware implementation reference on a field programmable gate array is designed. Simulation analyses show the feasibility and efficiency of the privacy-assured FogCS scheme.
Yushu Zhang 0001, Ping Wang 0029, Hui Huang 0008, Youwen Zhu, Di Xiao 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics2
2020 Secure Transmission of Compressed Sampling Data Using Edge Clouds
abstract
Cloud capability is considered to be extended to the edge of the Internet for improving the security of data transmission. Compressive sensing (CS) has been widely studied as a built-in privacy-preserving layer to provide some cryptographic features while sampling and compressing, including data confidentiality guarantees and data integrity guarantees. Unfortunately, most existing CS-based ciphers are too lightweight or highly complex to meet the requirements of both high security of transmitting the captured data over the Internet and low energy consumption of sensing devices in the Internet of Things (IoT). In this article, a secure transmission framework for CS data by combining CS-based cipher and edge computing is proposed. From the perspective of security, the double-layer encryption mechanism and double-layer authentication mechanism are rooted in it by performing some privacy-preserving operations, including CS-based encryption, CS-based hash, information splitting, strong encryption, and feature extraction. Most significantly, the proposed framework is very useful for resource-limited IoT applications.
Yushu Zhang 0001, Ping Wang 0029, Liming Fang 0001, Xing He 0001, Bing Chen 0002
IEEE Trans. Ind. Informatics2
2019 A robust and secure image sharing scheme with personal identity information embedded
Ping Wang 0029, Xing He 0001, Yushu Zhang 0001, Wenying Wen, Ming Li 0029
Comput. Secur.1