Cheng Peng 0008

dblp:82/3044-8 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction
abstract
Recent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth estimation and alignment can provide dense point cloud using few views; however, the resulting pose accuracy is suboptimal. In this work, we present SPARS3R, which combines the advantages of accurate pose estimation from Structure-from-Motion and dense point cloud from depth estimation. To this end, SPARS3R first performs a Global Fusion Alignment process that maps a prior dense point cloud to a sparse point cloud from Structure-from-Motion based on triangulated correspondences. RANSAC is applied during this process to distinguish inliers and outliers. SPARS3R then performs a second, Semantic Outlier Alignment step, which extracts semantically coherent regions around the outliers and performs local alignment in these regions. Along with several improvements in the evaluation process, we demonstrate that SPARS3R can achieve photorealistic rendering with sparse images and significantly outperforms existing approaches.
Yutao Tang, Yuxiang Guo 0001, Deming Li, Cheng Peng 0008
CVPR4
2025 METAREG: Robust Camera Parameter Estimation by Leveraging Noisy Camera Extrinsics
abstract
Novel view synthesis methods, such as Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), rely on Structure-from-Motion (SfM) pipelines like COLMAP for camera parameter estimation. However, these pipelines are prone to errors due to factors like doppelgangers and perspective distortion. While edge devices (e.g., mobile phones) capture images with embedded GPS and IMU data, this meta-data is often noisy. We introduce MetaReg, a robust pipeline that improves camera parameter estimation by leveraging noisy GPS metadata. MetaReg enhances COLMAP with pre-and post-processing steps: the preprocessing stage estimates image overlap using metadata to reduce unnecessary image matching, while the post-processing stage aligns estimated camera coordinates with the world coordinate system using noisy GPS priors. Experiments on challenging datasets demonstrate that MetaReg significantly improves camera parameter estimation, enhancing robustness and accuracy.
Chandrakanth Gudavalli, Tajuddin Manhar Mohammed, Ananth Vishnu Bhaskar, Elliot Staudt, Cheng Peng 0008, Abhay Yadav, Rama Chellappa, Shivkumar Chandrasekaran, B. S. Manjunath
ICIP5
2025 MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
abstract
In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis. Although recent adaptations of Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have improved in these areas, they tend to oversmooth and are prone to overfitting. In this paper, we present MS-GS, a novel framework designed with \textbf{M}ulti-appearance capabilities in \textbf{S}parse-view scenarios using 3D\textbf{GS}. To address the lack of support due to sparse initializations, our approach is built on the geometric priors elicited from monocular depth estimations. The key lies in extracting and utilizing local semantic regions with a Structure-from-Motion (SfM) points anchored algorithm for reliable alignment and geometry cues. Then, to introduce multi-view constraints, we propose a series of geometry-guided supervision steps at virtual views in pixel and feature levels to encourage 3D consistency and reduce overfitting. We also introduce a dataset and an in-the-wild experiment setting to set up more realistic benchmarks. We demonstrate that MS-GS achieves photorealistic renderings under various challenging sparse-view and multi-appearance conditions, and outperforms existing approaches significantly across different datasets.
Deming Li, Yutao Tang, Ravi Ramamoorthi, Rama Chellappa, Cheng Peng 0008
NeurIPS6
2025 GaitContour: Efficient Gait Recognition Based on a Contour-Pose Representation
abstract
Gait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by learning methods based on two input formats: silhouette images and sparse keypoints. Compared to image-based approaches, keypoint-based methods can achieve significantly higher efficiency due to their sparsity. However, sparsity also results in information loss, thereby reducing performance. In this work, we propose a novel, keypoint-based Contour-Pose representation, which compactly encodes both body shape and parts information. We further propose a local-to-global architecture, called GaitContour, to leverage this novel representation and efficiently compute subject embedding in two stages. The first stage consists of a local transformer that extracts features from five different body regions. The second stage then aggregates the regional features to estimate a global human gait representation. Such a design significantly reduces the complexity of the attention operation and improves both efficiency and performance. Through large scale experiments, Gait-Contour is shown to perform significantly better than previous keypoint-based methods. Furthermore, the ContourPose representation also achieves new SoTA performances on fusion-based gait recognition methods.
Yuxiang Guo 0001, Anshul Shah 0001, Jiang Liu 0014, Ayush Gupta 0001, Rama Chellappa, Cheng Peng 0008
WACV6
2025 VILLS: Video-Image Learning to Learn Semantics for Person Re-Identification
abstract
Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and inaccurate feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods.
Siyuan Huang 0005, Ram Prabhakar, Yuxiang Guo 0001, Rama Chellappa, Cheng Peng 0008
WACV5
2024 BAGS: Blur Agnostic Gaussian Splatting Through Multi-scale Kernel Modeling
Cheng Peng 0008, Yutao Tang, Nengyu Wang, Xijun Liu, Deming Li, Rama Chellappa
ECCV (80)1
2024 Distillation-guided Representation Learning for Unconstrained Gait Recognition
abstract
Gait recognition holds the promise of robustly identifying subjects based on walking patterns instead of appearance information. While previous approaches have performed well for curated indoor data, they tend to underperform in unconstrained situations, e.g. in outdoor, long distance scenes, etc. We propose a framework, termed GAit DEtection and Recognition (GADER), for human authentication in challenging outdoor scenarios. Specifically, GADER leverages a Double Helical Signature to detect segments that contain human movement and builds discriminative features through a novel gait recognition method, where only frames containing gait information are used. To further enhance robustness, GADER encodes viewpoint information in its architecture, and distills representation from an auxiliary RGB recognition model, which enables GADER to learn from silhouette and RGB data at training time. At test time, GADER only infers from the silhouette modality. We evaluate our method on multiple State-of-The-Arts(SoTA) gait baselines and demonstrate consistent improvements on indoor and outdoor datasets, especially with a significant 25.2% improvement on unconstrained, remote gait data.
Yuxiang Guo 0001, Siyuan Huang 0005, Ram Prabhakar, Chun Pong Lau 0001, Rama Chellappa, Cheng Peng 0008
IJCB6
2024 LP-3DGS: Learning to Prune 3D Gaussian Splatting
abstract
Recently, 3D Gaussian Splatting (3DGS) has become one of the mainstream methodologies for novel view synthesis (NVS) due to its high quality and fast rendering speed. However, as a point-based scene representation, 3DGS potentially generates a large number of Gaussians to fit the scene, leading to high memory usage. Improvements that have been proposed require either an empirical pre-set pruning ratio or importance score threshold to prune the point cloud. Such hyperparameters require multiple rounds of training to optimize and achieve the maximum pruning ratio while maintaining the rendering quality for each scene. In this work, we propose learning-to-prune 3DGS (LP-3DGS), where a trainable binary mask is applied to the importance score to automatically find a favorable pruning ratio. Instead of using the traditional straight-through estimator (STE) method to approximate the binary mask gradient, we redesign the masking function to leverage the Gumbel-Sigmoid method, making it differentiable and compatible with the existing training process of 3DGS. Extensive experiments have shown that LP-3DGS consistently achieves a good balance between efficiency and high quality.
Zhaoliang Zhang, Tianchen Song, Li Yang 0009, Cheng Peng 0008, Rama Chellappa, Deliang Fan
NeurIPS5
2023 PDRF: Progressively Deblurring Radiance Field for Fast Scene Reconstruction from Blurry Images
abstract
We present Progressively Deblurring Radiance Field (PDRF), a novel approach to efficiently reconstruct high quality radiance fields from blurry images. While current State-of-The-Art (SoTA) scene reconstruction methods achieve photo-realistic renderings from clean source views, their performances suffer when the source views are affected by blur, which is commonly observed in the wild. Previous deblurring methods either do not account for 3D geometry, or are computationally intense. To addresses these issues, PDRF uses a progressively deblurring scheme for radiance field modeling, which can accurately model blur with 3D scene context. PDRF further uses an efficient importance sampling scheme that results in fast scene optimization. We perform extensive experiments and show that PDRF is 15X faster than previous SoTA while achieving better performance on both synthetic and real scenes.
Cheng Peng 0008, Rama Chellappa
AAAI1
2023 Multi-Modal Human Authentication Using Silhouettes, Gait and RGB
abstract
Whole-body-based human authentication is a promising approach for remote biometrics scenarios. Current literature focuses on either body recognition based on RGB images or gait recognition based on body shapes and walking patterns; both have their advantages and drawbacks. In this work, we propose Dual-Modal Ensemble (DME), which combines both RGB and silhouette data to achieve more robust performances for indoor and outdoor whole-body based recognition. Within DME, we propose GaitPattern, which is inspired by the double helical gait pattern used in traditional gait analysis. The GaitPattern contributes to robust identification performance over a large range of viewing angles. Extensive experimental results on the CASIA-B dataset demonstrate that the proposed method outperforms state-of-the-art recognition systems. We also provide experimental results using the newly collected BRIAR dataset.
Yuxiang Guo 0001, Cheng Peng 0008, Chun Pong Lau 0001, Rama Chellappa
FG2
2022 HyperSegNAS: Bridging One-Shot Neural Architecture Search with 3D Medical Image Segmentation using HyperNet
abstract
Semantic segmentation of 3D medical images is a challenging task due to the high variability of the shape and pattern of objects (such as organs or tumors). Given the recent success of deep learning in medical image segmentation, Neural Architecture Search (NAS) has been introduced to find high-performance 3D segmentation network architectures. However, because of the massive computational requirements of 3D data and the discrete optimization nature of architecture search, previous NAS methods require a long search time or necessary continuous relaxation, and commonly lead to sub-optimal network architectures. While one-shot NAS can potentially address these disadvantages, its application in the segmentation domain has not been well studied in the expansive multi-scale multi-path search space. To enable one-shot NAS for medical image segmentation, our method, named HyperSegNAS, introduces a HyperNet to assist super-net training by incorporating architecture topology information. Such a HyperNet can be removed once the super-net is trained and introduces no overhead during architecture search. We show that HyperSegNAS yields better performing and more intuitive architectures compared to the previous state-of-the-art (SOTA) segmentation networks; furthermore, it can quickly and accurately find good architecture candidates under different computing constraints. Our method is evaluated on public datasets from the Medical Segmentation Decathlon (MSD) challenge, and achieves SOTA performances.
Cheng Peng 0008, Andriy Myronenko, Ali Hatamizadeh, Vishwesh Nath, Md Mahfuzur Rahman Siddiquee, Yufan He, Daguang Xu, Rama Chellappa, Dong Yang 0005
CVPR1
2022 Undersampled MRI Reconstruction with Side Information-Guided Normalisation
Xinwen Liu 0003, Jing Wang 0062, Cheng Peng 0008, Shekhar Chandra, Feng Liu 0005, Shaohua Kevin Zhou
MICCAI (6)3
2022 Towards Performant and Reliable Undersampled MR Reconstruction via Diffusion Model Sampling
Cheng Peng 0008, Shaohua Kevin Zhou, Vishal M. Patel, Rama Chellappa
MICCAI (6)1
2022 GAN-based disentanglement learning for chest X-ray rib suppression
Luyi Han, Yuanyuan Lyu, Cheng Peng 0008, Shaohua Kevin Zhou
Medical Image Anal.3
2021 XraySyn: Realistic View Synthesis From a Single Radiograph Through CT Priors
abstract
A radiograph visualizes the internal anatomy of a patient through the use of X-ray, which projects 3D information onto a 2D plane. Hence, radiograph analysis naturally requires physicians to relate their prior knowledge about 3D human anatomy to 2D radiographs. Synthesizing novel radiographic views in a small range can assist physicians in interpreting anatomy more reliably; however, radiograph view synthesis is heavily ill-posed, lacking in paired data, and lacking in differentiable operations to leverage learning-based approaches. To address these problems, we use Computed Tomography (CT) for radiograph simulation and design a differentiable projection algorithm, which enables us to achieve geometrically consistent transformations between the radiography and CT domains. Our method, XraySyn, can synthesize novel views on real radiographs through a combination of realistic simulation and finetuning on real radiographs. To the best of our knowledge, this is the first work on radiograph view synthesis. We show that by gaining an understanding of radiography in 3D space, our method can be applied to radiograph bone extraction and suppression without requiring groundtruth bone labels.
Cheng Peng 0008, Haofu Liao, Gina Wong, Jiebo Luo 0001, Shaohua Kevin Zhou, Rama Chellappa
AAAI1
2021 U-DuDoNet: Unpaired Dual-Domain Network for CT Metal Artifact Reduction
Yuanyuan Lyu, Jiajun Fu, Cheng Peng 0008, Shaohua Kevin Zhou
MICCAI (6)3
2021 DA-VSR: Domain Adaptable Volumetric Super-Resolution for Medical Images
Cheng Peng 0008, Shaohua Kevin Zhou, Rama Chellappa
MICCAI (6)1
2020 SAINT: Spatially Aware Interpolation NeTwork for Medical Slice Synthesis
abstract
Deep learning-based single image super-resolution (SISR) methods face various challenges when applied to 3D medical volumetric data (i.e., CT and MR images) due to the high memory cost and anisotropic resolution, which adversely affect their performance. Furthermore, mainstream SISR methods are designed to work over specific upsampling factors, which makes them ineffective in clinical practice. In this paper, we introduce a Spatially Aware Interpolation NeTwork (SAINT) for medical slice synthesis to alleviate the memory constraint that volumetric data poses. Compared to other super-resolution methods, SAINT utilizes voxel spacing information to provide desirable levels of details, and allows for the upsampling factor to be determined on the fly. Our evaluations based on 853 CT scans from four datasets that contain liver, colon, hepatic vessels, and kidneys show that SAINT consistently outperforms other SISR methods in terms of medical slice synthesis quality, while using only a single model to deal with different upsampling factors.
Cheng Peng 0008, Wei-An Lin, Haofu Liao, Rama Chellappa, Shaohua Kevin Zhou
CVPR1
2019 DuDoNet: Dual Domain Network for CT Metal Artifact Reduction
abstract
Computed tomography (CT) is an imaging modality widely used for medical diagnosis and treatment. CT images are often corrupted by undesirable artifacts when metallic implants are carried by patients, which creates the problem of metal artifact reduction (MAR). Existing methods for reducing the artifacts due to metallic implants are inadequate for two main reasons. First, metal artifacts are structured and non-local so that simple image domain enhancement approaches would not suffice. Second, the MAR approaches which attempt to reduce metal artifacts in the X-ray projection (sinogram) domain inevitably lead to severe secondary artifact due to sinogram inconsistency. To overcome these difficulties, we propose an end-to-end trainable Dual Domain Network (DuDoNet) to simultaneously restore sinogram consistency and enhance CT images. The linkage between the sigogram and image domains is a novel Radon inversion layer that allows the gradients to back-propagate from the image domain to the sinogram domain during training. Extensive experiments show that our method achieves significant improvements over other single domain MAR approaches. To the best of our knowledge, it is the first end-to-end dual-domain network for MAR.
Wei-An Lin, Haofu Liao, Cheng Peng 0008, Xiaohang Sun, Jingdan Zhang, Jiebo Luo 0001, Rama Chellappa, Shaohua Kevin Zhou
CVPR3