Pengchao Deng

dblp:277/3367 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MS-YOLO: a multi-scale model for accurate and efficient blood cell detection
Shengqi Chen 0003, Pengchao Deng, Wenting Yu
Pattern Anal. Appl.2
2025 Learnable Fractional Reaction-Diffusion Dynamics for Under-Display ToF Imaging and Beyond
abstract
Under-display ToF imaging aims to achieve accurate depth sensing through a ToF camera placed beneath a screen panel. However, transparent OLED (TOLED) layers introduce severe degradations-such as signal attenuation, multi-path interference (MPI), and temporal noise-that significantly compromise depth quality. To alleviate this drawback, we propose Learnable Fractional Reaction-Diffusion Dynamics (LFRD2), a hybrid framework that combines the expressive power of neural networks with the interpretability of physical modeling. Specifically, we implement a time-fractional reaction-diffusion module that enables iterative depth refinement with dynamically generated differential orders, capturing long-term dependencies. In addition, we introduce an efficient continuous convolution operator via coefficient prediction and repeated differentiation to further improve restoration quality. Experiments on four benchmark datasets demonstrate the effectiveness of our approach. The code is publicly available at https://github.com/wudiqx106/LFRD2.
Matteo Poggi, Pengchao Deng, Stefano Mattoccia
ICCV4
2025 Memory guided representation learning for cross-domain face anti-spoofing
Pengchao Deng, Zhiheng Fu, Shengjun Xu, Chenyang Ge, Farid Boussaïd, Mohammed Bennamoun
Eng. Appl. Artif. Intell.1
2025 Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose Estimation
abstract
Nine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization to intra-class unknown objects. However, these methods require manual collection and labeling of large-scale real-world training data. To address this problem, we introduce a diffusion-based paradigm for domain-generalized category-level 9-DoF object pose estimation. Our motivation is to leverage the latent generalization ability of the diffusion model to address the domain generalization challenge in object pose estimation. This entails training the model exclusively on rendered synthetic data to achieve generalization to real-world scenes. We propose an effective diffusion model to redefine 9-DoF object pose estimation from a generative perspective. Our model does not require any 3D shape priors during training or inference. By employing the Denoising Diffusion Implicit Model, we demonstrate that the reverse diffusion process can be executed in as few as 3 steps, achieving near real-time performance. Finally, we design a robotic grasping system comprising both hardware and software components. Through comprehensive experiments on two benchmark datasets and the real-world robotic system, we show that our method achieves state-of-the-art domain generalization performance.
Jian Liu 0014, Wei Sun 0028, Pengchao Deng, Chongpei Liu, Nicu Sebe, Hossein Rahmani 0001, Ajmal Mian
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Multimodal contrastive learning for face anti-spoofing
Pengchao Deng, Chenyang Ge, Hao Wei 0005
Eng. Appl. Artif. Intell.1
2024 Real-world image deblurring using data synthesis and feature complementary network
abstract
Abstract Many learning‐based approaches to image deblurring have received increasing attention in recent years. However, the models trained on existing synthetic datasets do not generalize well to real‐world blur, resulting in undesirable artifacts and residual blur. This work attempts to address this problem from two aspects: training data synthesis and network architecture. To narrow the domain gap between synthetic and real domains, a realistic blur synthesis pipeline to generate high‐quality blurred data is proposed. Since the blur is non‐uniform and has different scales and degrees, a parallel feature complementary module to fully exploit the local and non‐local information, which improves the feature representation and helps the network to perceive the non‐uniform blur, is developed. In addition, a spatial Fourier reconstruction block to facilitate correct detail recovery in the spatial and Fourier domains is introduced. Based on these two designs, an effective encoder–decoder network for deblurring is designed. Extensive experiments demonstrate the validity and superiority of the proposed blur synthesis method and deblurring network. In particular, the proposed deblurring network can achieve superior or comparable performance to Restormer, while saving 70% of network parameters and 53% of floating point operations (FLOPs).
Hao Wei 0005, Chenyang Ge, Pengchao Deng
IET Image Process.4
2024 RGB Guided ToF Imaging System: A Survey of Deep Learning-Based Methods
Matteo Poggi, Pengchao Deng, Hao Wei 0005, Chenyang Ge, Stefano Mattoccia
Int. J. Comput. Vis.3
2024 Toward Extreme Image Rescaling With Generative Prior and Invertible Prior
abstract
The goal of image rescaling is to embed the information from high-resolution images into low-resolution images and then reconstruct the high-resolution images in reverse. Existing methods either focus on small scaling factors or do not generalize well to natural images with diverse content in extreme settings, i.e., using extreme scaling factors (e.g., 16× and 32×). When performing extreme rescaling, previous methods often fail to produce plausible high-quality results due to insufficient cues in low-resolution images. In this work, we propose an extreme natural image rescaling framework that exploits the rich generative prior integrated into the GAN model trained on large-scale natural images to reduce the ambiguity of extreme upscaling. Considering the invertible bijective transformation between quantized features and low-resolution image, we develop an invertible feature recovery module that generates semantically sound low-resolution image while maximizing the preservation of useful features for the subsequent upscaling. Furthermore, we propose a multi-scale refinement module that explicitly introduces the supervised ground truth information to mitigate unpleasant artifacts and distortions. Extensive experiments show that the proposed rescaling framework formulated by the above components achieves significantly better visual performance than state-of-the-art methods.
Hao Wei 0005, Chenyang Ge, Pengchao Deng
IEEE Trans. Circuits Syst. Video Technol.5
2023 Under-Display ToF Imaging with Efficient Transformer
abstract
Demand for full screen in consumer electronics has propelled the development of under-display image processing. For time-of-flight cameras, they are usually placed under TOLED display. Due to the presence of pixels and display patterns in the TOLED panel, depth maps from under-display time-of-flight (UD-ToF) are noisy, blurry, and inaccurate. We propose a non-local method based on Vision Transformer for UD-ToF depth restoration to address these issues. Specifically, a novel feature attention block is designed to incorporate non-local depth features. Additionally, we take ToF raw measurements rather than the depth map as input, allowing the network to extract informative features from the raw domain. We conduct comprehensive experiments on real RUD-TOF and synthetic SUD-TOF benchmark datasets, and the results indicate that the proposed transformer-based method achieves better performance than the state-of-the-art algorithms.
Hao Wei 0005, Pengchao Deng, Chenyang Ge
VCIP4
2023 Non-uniform Deblurring by Deep Sharpness Edge Guided Model
abstract
In this paper, we propose a two-branch deblurring framework. Given a blurred image, we first extract the edge map and employ an edge refinement network to recover the structure. Then the refined edge map is utilized to guide the subsequent deblurring process for correct structure recovery. Specifically, we develop a lightweight omni-dimensional attention module for long-range dependencies modeling and plug it into the edge refinement network, which effectively handles blur patterns with high variation. Furthermore, we propose a dynamic feature upsample module, which integrates dynamic convolution with upsampling and adaptively deals with the non-uniform blur. Extensive experiments show that our method outperforms state-of-the-art methods.
Hao Wei 0005, Chenyang Ge, Pengchao Deng
VCIP4
2023 Depth Restoration in Under-Display Time-of-Flight Imaging
abstract
Under-display imaging has recently received considerable attention in both academia and industry. As a variation of this technique, under-display ToF (UD-ToF) cameras enable depth sensing for full-screen devices. However, it also brings problems of image blurring, signal-to-noise ratio and ranging accuracy reduction. To address these issues, we propose a cascaded deep network to improve the quality of UD-ToF depth maps. The network comprises two subnets, with the first using a complex-valued network in raw domain to perform denoising, deblurring and raw measurements enhancement jointly, while the second refining depth maps in depth domain based on the proposed multi-scale depth enhancement block (MSDEB). To enable training, we establish a data acquisition device and construct a real UD-ToF dataset by collecting real paired ToF raw data. Besides, we also build a large-scale synthetic UD-ToF dataset through noise analysis. The quantitative and qualitative evaluation results on public datasets and ours demonstrate that the presented network outperforms state-of-the-art algorithms and can further promote full-screen devices in practical applications.
Chenyang Ge, Pengchao Deng, Hao Wei 0005, Matteo Poggi, Stefano Mattoccia
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Attention-Aware Dual-Stream Network for Multimodal Face Anti-Spoofing
abstract
Since the rapid development of face recognition systems using 3D cameras, the public has demanded great safety regulations for these devices. As a closely related topic, multimodal face anti-spoofing (FAS) has become an indispensable part of face recognition systems. However, existing multimodal FAS tools suffer from performance degradation under external low-lighting conditions and insufficient representation capabilities of fusion features. To address these issues, we present an attention-aware dual-stream fusion method using 3D cameras (i.e., IR+Depth) and considering both fine-grained and global features. Specifically, we introduce a surface normal generator using depth maps to obtain robust and discriminative representations. Then, we leverage the attention mechanism to split each stream into two branches. The first branch explores complementary global information between different modalities, while the second branch captures subtle and local features from each modality. The system regards multimodal FAS as a fine-grained classification problem. Moreover, to ensure that local areas in the image do not overlap and belong to the same class, a joint loss function is developed and proven to further boost the performance of FAS. We extensively evaluate our proposed strategies on various multimodal databases, and the results show that when compared with current state-of-the-art multimodal methods, our framework achieves superior performance.
Pengchao Deng, Chenyang Ge, Hao Wei 0005
IEEE Trans. Inf. Forensics Secur.1
2019 Valid depth data extraction and correction for time-of-flight camera
abstract
In this paper an algorithm is presented to extract the valid depth data and correct the values of flying pixels by using depth information and confidence image. An adaptive segmentation for the measured depth image is executed based on kernel density estimation and one-pass connected component labeling. Then a modified structure tensor is used to detect the invalid pixels and the flying pixels contained in the depth image. Finally these pixels are corrected with the bi-cubic interpolation method or selectively removed by voting operation. And also, the erroneous pixels are excluded with augmented confidence. Experimental results have demonstrated the effectiveness of our algorithm.
Chenyang Ge, Huimin Yao, Pengchao Deng
ICMV4