Hao Wei 0005

dblp:96/133-5 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-1747-1218ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A lightweight model for perceptual image compression via implicit priors
Hao Wei 0005, Yiwen Jia, Chenyang Ge, Saeed Anwar, Ajmal Mian
Neural Networks1
2025 One-Step Diffusion for Perceptual Image Compression
abstract
Diffusion-based image compression methods have achieved notable progress, delivering high perceptual quality at low bitrates. However, their practical deployment is hindered by significant inference latency and heavy computational overhead, primarily due to the large number of denoising steps required during decoding. To address this problem, we propose a diffusion-based image compression method that requires only a single-step diffusion process, significantly improving inference speed. To enhance the perceptual quality of reconstructed images, we introduce a discriminator that operates on compact feature representations instead of raw pixels, leveraging the fact that features better capture high-level texture and structural details. Experimental results show that our method delivers comparable compression performance while offering a 46× faster inference speed compared to recent diffusion-based approaches. The source code and models are available at https://github.com/cheesejiang/OSDiff.
Yiwen Jia, Hao Wei 0005, Chenyang Ge
VCIP2
2025 Toward Extreme Image Compression With Latent Feature Guidance and Diffusion Prior
abstract
Image compression at extremely low bitrates (below 0.1 bits per pixel (bpp)) is a significant challenge due to substantial information loss. In this work, we propose a novel two-stage extreme image compression framework that exploits the powerful generative capability of pre-trained diffusion models to achieve realistic image reconstruction at extremely low bitrates. In the first stage, we treat the latent representation of images in the diffusion space as guidance, employing a VAE-based compression approach to compress images and initially decode the compressed information into content variables. The second stage leverages pre-trained stable diffusion to reconstruct images under the guidance of content variables. Specifically, we introduce a small control module to inject content information while keeping the stable diffusion model fixed to maintain its generative capability. Furthermore, we design a space alignment loss to force the content variables to align with the diffusion space and provide the necessary constraints for optimization. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches in terms of visual performance at extremely low bitrates. The source code and trained models are available athttps://github.com/huai-chang/DiffEIC.
Hao Wei 0005, Chenyang Ge
IEEE Trans. Circuits Syst. Video Technol.3
2025 RDEIC: Accelerating Diffusion-Based Extreme Image Compression With Relay Residual Diffusion
abstract
Diffusion-based extreme image compression methods have achieved impressive performance at extremely low bitrates. However, constrained by the iterative denoising process that starts from pure noise, these methods are limited in both fidelity and efficiency. To address these two issues, we present Relay Residual Diffusion Extreme Image Compression (RDEIC), which leverages compressed feature initialization and residual diffusion. Specifically, we first use the compressed latent features of the image with added noise, instead of pure noise, as the starting point to eliminate the unnecessary initial stages of the denoising process. Second, we directly derive a novel residual diffusion equation from Stable Diffusion’s original diffusion equation that reconstructs the raw image by iteratively removing the added noise and the residual between the compressed and target latent features. In this way, we effectively combine the efficiency of residual diffusion with the powerful generative capability of Stable Diffusion. Third, we propose a fixed-step fine-tuning strategy to eliminate the discrepancy between the training and inference phases, thereby further improving the reconstruction quality. Extensive experiments demonstrate that the proposed RDEIC achieves state-of-the-art visual quality and outperforms existing diffusion-based extreme image compression methods in both fidelity and efficiency. The source code and pre-trained models are available at https://github.com/huai-chang/RDEIC.
Hao Wei 0005, Chenyang Ge, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multimodal contrastive learning for face anti-spoofing
Pengchao Deng, Chenyang Ge, Hao Wei 0005
Eng. Appl. Artif. Intell.3
2024 Real-world image deblurring using data synthesis and feature complementary network
abstract
Abstract Many learning‐based approaches to image deblurring have received increasing attention in recent years. However, the models trained on existing synthetic datasets do not generalize well to real‐world blur, resulting in undesirable artifacts and residual blur. This work attempts to address this problem from two aspects: training data synthesis and network architecture. To narrow the domain gap between synthetic and real domains, a realistic blur synthesis pipeline to generate high‐quality blurred data is proposed. Since the blur is non‐uniform and has different scales and degrees, a parallel feature complementary module to fully exploit the local and non‐local information, which improves the feature representation and helps the network to perceive the non‐uniform blur, is developed. In addition, a spatial Fourier reconstruction block to facilitate correct detail recovery in the spatial and Fourier domains is introduced. Based on these two designs, an effective encoder–decoder network for deblurring is designed. Extensive experiments demonstrate the validity and superiority of the proposed blur synthesis method and deblurring network. In particular, the proposed deblurring network can achieve superior or comparable performance to Restormer, while saving 70% of network parameters and 53% of floating point operations (FLOPs).
Hao Wei 0005, Chenyang Ge, Pengchao Deng
IET Image Process.1
2024 RGB Guided ToF Imaging System: A Survey of Deep Learning-Based Methods
Matteo Poggi, Pengchao Deng, Hao Wei 0005, Chenyang Ge, Stefano Mattoccia
Int. J. Comput. Vis.4
2024 Toward Extreme Image Rescaling With Generative Prior and Invertible Prior
abstract
The goal of image rescaling is to embed the information from high-resolution images into low-resolution images and then reconstruct the high-resolution images in reverse. Existing methods either focus on small scaling factors or do not generalize well to natural images with diverse content in extreme settings, i.e., using extreme scaling factors (e.g., 16× and 32×). When performing extreme rescaling, previous methods often fail to produce plausible high-quality results due to insufficient cues in low-resolution images. In this work, we propose an extreme natural image rescaling framework that exploits the rich generative prior integrated into the GAN model trained on large-scale natural images to reduce the ambiguity of extreme upscaling. Considering the invertible bijective transformation between quantized features and low-resolution image, we develop an invertible feature recovery module that generates semantically sound low-resolution image while maximizing the preservation of useful features for the subsequent upscaling. Furthermore, we propose a multi-scale refinement module that explicitly introduces the supervised ground truth information to mitigate unpleasant artifacts and distortions. Extensive experiments show that the proposed rescaling framework formulated by the above components achieves significantly better visual performance than state-of-the-art methods.
Hao Wei 0005, Chenyang Ge, Pengchao Deng
IEEE Trans. Circuits Syst. Video Technol.1
2023 Under-Display ToF Imaging with Efficient Transformer
abstract
Demand for full screen in consumer electronics has propelled the development of under-display image processing. For time-of-flight cameras, they are usually placed under TOLED display. Due to the presence of pixels and display patterns in the TOLED panel, depth maps from under-display time-of-flight (UD-ToF) are noisy, blurry, and inaccurate. We propose a non-local method based on Vision Transformer for UD-ToF depth restoration to address these issues. Specifically, a novel feature attention block is designed to incorporate non-local depth features. Additionally, we take ToF raw measurements rather than the depth map as input, allowing the network to extract informative features from the raw domain. We conduct comprehensive experiments on real RUD-TOF and synthetic SUD-TOF benchmark datasets, and the results indicate that the proposed transformer-based method achieves better performance than the state-of-the-art algorithms.
Hao Wei 0005, Pengchao Deng, Chenyang Ge
VCIP2
2023 Non-uniform Deblurring by Deep Sharpness Edge Guided Model
abstract
In this paper, we propose a two-branch deblurring framework. Given a blurred image, we first extract the edge map and employ an edge refinement network to recover the structure. Then the refined edge map is utilized to guide the subsequent deblurring process for correct structure recovery. Specifically, we develop a lightweight omni-dimensional attention module for long-range dependencies modeling and plug it into the edge refinement network, which effectively handles blur patterns with high variation. Furthermore, we propose a dynamic feature upsample module, which integrates dynamic convolution with upsampling and adaptively deals with the non-uniform blur. Extensive experiments show that our method outperforms state-of-the-art methods.
Hao Wei 0005, Chenyang Ge, Pengchao Deng
VCIP1
2023 Depth Restoration in Under-Display Time-of-Flight Imaging
abstract
Under-display imaging has recently received considerable attention in both academia and industry. As a variation of this technique, under-display ToF (UD-ToF) cameras enable depth sensing for full-screen devices. However, it also brings problems of image blurring, signal-to-noise ratio and ranging accuracy reduction. To address these issues, we propose a cascaded deep network to improve the quality of UD-ToF depth maps. The network comprises two subnets, with the first using a complex-valued network in raw domain to perform denoising, deblurring and raw measurements enhancement jointly, while the second refining depth maps in depth domain based on the proposed multi-scale depth enhancement block (MSDEB). To enable training, we establish a data acquisition device and construct a real UD-ToF dataset by collecting real paired ToF raw data. Besides, we also build a large-scale synthetic UD-ToF dataset through noise analysis. The quantitative and qualitative evaluation results on public datasets and ours demonstrate that the presented network outperforms state-of-the-art algorithms and can further promote full-screen devices in practical applications.
Chenyang Ge, Pengchao Deng, Hao Wei 0005, Matteo Poggi, Stefano Mattoccia
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Attention-Aware Dual-Stream Network for Multimodal Face Anti-Spoofing
abstract
Since the rapid development of face recognition systems using 3D cameras, the public has demanded great safety regulations for these devices. As a closely related topic, multimodal face anti-spoofing (FAS) has become an indispensable part of face recognition systems. However, existing multimodal FAS tools suffer from performance degradation under external low-lighting conditions and insufficient representation capabilities of fusion features. To address these issues, we present an attention-aware dual-stream fusion method using 3D cameras (i.e., IR+Depth) and considering both fine-grained and global features. Specifically, we introduce a surface normal generator using depth maps to obtain robust and discriminative representations. Then, we leverage the attention mechanism to split each stream into two branches. The first branch explores complementary global information between different modalities, while the second branch captures subtle and local features from each modality. The system regards multimodal FAS as a fine-grained classification problem. Moreover, to ensure that local areas in the image do not overlap and belong to the same class, a joint loss function is developed and proven to further boost the performance of FAS. We extensively evaluate our proposed strategies on various multimodal databases, and the results show that when compared with current state-of-the-art multimodal methods, our framework achieves superior performance.
Pengchao Deng, Chenyang Ge, Hao Wei 0005
IEEE Trans. Inf. Forensics Secur.4
2020 Deep Video Deblurring Using Sharpness Features From Exemplars
abstract
Video deblurring is a challenging problem as the blur in videos is usually caused by camera shake, object motion, depth variation, etc. Existing methods usually impose handcrafted image priors or use end-to-end trainable networks to solve this problem. However, using image priors usually leads to highly non-convex problems while directly using end-to-end trainable networks in a regression generates over-smoothes details in the restored images. In this paper, we explore the sharpness features from exemplars to help the blur removal and details restoration. We first estimate optical flow to explore the temporal information which can help to make full use of neighboring information. Then, we develop an encoder and decoder network and explore the sharpness features from exemplars to guide the network for better image restoration. We train the proposed algorithm in an end-to-end manner and show that using sharpness features from exemplars can help blur removal and details restoration. Both quantitative and qualitative evaluations demonstrate that our method performs favorably against state-of-the-art approaches on the benchmark video deblurring datasets and real-world images.
Xinguang Xiang, Hao Wei 0005, Jinshan Pan
IEEE Trans. Image Process.2