Zhaoqi Su

dblp:211/7235 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-3651-8373ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 76% Face, body and person analysis · 19% Robot navigation and mapping · 5%
Computer graphics and multimedia
5 papers
Geometric modeling and processing · 44% Computer animation and physical simulation · 29% Visual content generation and editing · 20%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.022023
CaPhy: Capturing Physical Properties for Animatable Human Avatars · ICCV 2023
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › 3D vision › 3d reconstruction › object reconstruction
3d head reconstruction
0.812024
3D Gaussian Parametric Head Model · ECCV (35) 2024
Geometric modeling and processing › 3d face modeling
parametric head model
0.812024
3D Gaussian Parametric Head Model · ECCV (35) 2024
Computer vision › 3D vision › human body modeling › 3d human modeling
animatable human avatar
0.712023
CaPhy: Capturing Physical Properties for Animatable Human Avatars · ICCV 2023
Computer vision › 3D vision › motion capture
human performance capture
0.712023
CaPhy: Capturing Physical Properties for Animatable Human Avatars · ICCV 2023
Visual content generation and editing
3d content editing
0.712023
DeepCloth: Neural Garment Representation for Shape and Style Editing · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer animation and physical simulation
performance capture
0.612022
MulayCap: Multi-Layer Human Performance Capture Using a Monocular Video Camera · IEEE Trans. Vis. Comput. Graph. 2022
Computer animation and physical simulation
cloth simulation
0.422023
CaPhy: Capturing Physical Properties for Animatable Human Avatars · ICCV 2023
MulayCap: Multi-Layer Human Performance Capture Using a Monocular Video Camera · IEEE Trans. Vis. Comput. Graph. 2022
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › Face, body and person analysis
human pose estimation
0.312018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › 3D vision › depth estimation
depth reconstruction
0.312017
BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera · ICCV 2017
Robotics › Robot navigation and mapping › sensor fusion
geometric data fusion
0.312017
BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera · ICCV 2017
Computer vision › 3D vision › 3d shape reconstruction
non-rigid surface reconstruction
0.312017
BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera · ICCV 2017
Virtual and augmented reality › telepresence
3d telepresence
0.112018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Virtual and augmented reality
immersive experience
0.112018
Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras · IEEE Trans. Vis. Comput. Graph. 2018
Computer vision › 3D vision › motion capture
human motion capture
0.112017
BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera · ICCV 2017

Methods — techniques the papers use, named apart from their topics

3d gaussian representation · 1.5unsupervised training · 1.3physics-based loss · 1.3gradient constraint · 1.33d-supervised training · 1.3latent space optimization · 0.7convolutional neural network · 0.7UV position map · 0.7gradient descent · 0.6cloth simulation · 0.6parametric model-based motion estimation · 0.3skeleton-embedded surface fusion · 0.3graph-node deformation · 0.3articulated motion prior · 0.3
YearPublicationVenuePosition
2026 Parametric Gaussian Human Model: Generalizable Prior for Efficient and Realistic Human Avatar Modeling
abstract
Photorealistic and animatable human avatars are a key enabler for virtual/augmented reality, telepresence, and digital entertainment. While recent advances in 3D Gaussian Splatting (3DGS) have greatly improved rendering quality and efficiency, existing methods still face fundamental challenges, including time-consuming per-subject optimization and poor generalization under sparse monocular inputs. In this work, we present the Parametric Gaussian Human Model (PGHM), a generalizable and efficient framework that integrates human priors into 3DGS for fast and high-fidelity avatar reconstruction from monocular videos. PGHM introduces two core components: (1) a UV-aligned latent identity map that compactly encodes subject-specific geometry and appearance into a learnable feature tensor; and (2) a disentangled Multi-Head U-Net that predicts Gaussian attributes by decomposing static, pose-dependent, and view-dependent components via conditioned decoders. This design enables robust rendering quality under challenging poses and viewpoints, while allowing efficient subject adaptation without requiring multiview capture or long optimization time. Experiments show that PGHM is significantly more efficient than optimization-from-scratch methods, requiring only approximately 20 minutes per subject to produce avatars with comparable visual quality, thereby demonstrating its practical applicability for real-world monocular avatar creation.
Jingxiang Sun, Yushuo Chen 0001, Zhaoqi Su, Zhuo Su 0006, Yebin Liu
3DV4
2026 ThermoSplat: Cross-modal 3D Gaussian splatting with feature modulation and geometry decoupling
Zhaoqi Su, Shihai Chen, Xinyan Lin, Liqin Huang, Zhipeng Su, Xiaoqiang Lu
Neurocomputing1
2024 3D Gaussian Parametric Head Model
Yuelang Xu, Lizhen Wang 0002, Zerong Zheng, Zhaoqi Su, Yebin Liu
ECCV (35)4
2023 CaPhy: Capturing Physical Properties for Animatable Human Avatars
abstract
We present CaPhy, a novel method for reconstructing animatable human avatars with realistic dynamic properties for clothing. Specifically, we aim for capturing the geometric and physical properties of the clothing from real observations. This allows us to apply novel poses to the human avatar with physically correct deformations and wrinkles of the clothing. To this end, we combine unsupervised training with physics-based losses and 3D-supervised training using scanned data to reconstruct a dynamic model of clothing that is physically realistic and conforms to the human scans. We also optimize the physical parameters of the underlying physical model from the scans by introducing gradient constraints of the physics-based losses. In contrast to previous work on 3D avatar reconstruction, our method is able to generalize to novel poses with realistic dynamic cloth deformations. Experiments on several subjects demonstrate that our method can estimate the physical properties of the garments, resulting in superior quantitative and qualitative results compared with previous methods.
Zhaoqi Su, Liangxiao Hu, Siyou Lin, Hongwen Zhang 0001, Shengping Zhang, Justus Thies, Yebin Liu
ICCV1
2023 DeepCloth: Neural Garment Representation for Shape and Style Editing
abstract
Garment representation, editing and animation are challenging topics in the area of computer vision and graphics. It remains difficult for existing garment representations to achieve smooth and plausible transitions between different shapes and topologies. In this work, we introduce, DeepCloth, a unified framework for garment representation, reconstruction, animation and editing. Our unified framework contains 3 components: First, we represent the garment geometry with a "topology-aware UV-position map", which allows for the unified description of various garments with different shapes and topologies by introducing an additional topology-aware UV-mask for the UV-position map. Second, to further enable garment reconstruction and editing, we contribute a method to embed the UV-based representations into a continuous feature space, which enables garment shape reconstruction and editing by optimization and control in the latent space, respectively. Finally, we propose a garment animation method by unifying our neural garment representation with body shape and pose, which achieves plausible garment animation results leveraging the dynamic information encoded by our shape and style representation, even under drastic garment editing operations. To conclude, with DeepCloth, we move a step forward in establishing a more flexible and general 3D garment digitization framework. Experiments demonstrate that our method can achieve state-of-the-art garment representation performance compared with previous methods.
Zhaoqi Su, Tao Yu 0007, Yangang Wang 0001, Yebin Liu
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 MulayCap: Multi-Layer Human Performance Capture Using a Monocular Video Camera
abstract
We introduce MulayCap, a novel human performance capture method using a monocular video camera without the need for pre-scanning. The method uses "multi-layer" representations for geometry reconstruction and texture rendering, respectively. For geometry reconstruction, we decompose the clothed human into multiple geometry layers, namely a body mesh layer and a garment piece layer. The key technique behind is a Garment-from-Video (GfV) method for optimizing the garment shape and reconstructing the dynamic cloth to fit the input video sequence, based on a cloth simulation model which is effectively solved with gradient descent. For texture rendering, we decompose each input image frame into a shading layer and an albedo layer, and propose a method for fusing a fixed albedo map and solving for detailed garment geometry using the shading layer. Compared with existing single view human performance capture systems, our "multi-layer" approach bypasses the tedious and time consuming scanning step for obtaining a human specific mesh template. Experimental results demonstrate that MulayCap produces realistic rendering of dynamically changing details that has not been achieved in any previous monocular video camera systems. Benefiting from its fully semantic modeling, MulayCap can be applied to various important editing applications, such as cloth editing, re-targeting, relighting, and AR applications.
Zhaoqi Su, Weilin Wan 0001, Tao Yu 0007, Lingjie Liu, Lu Fang 0001, Wenping Wang 0001, Yebin Liu
IEEE Trans. Vis. Comput. Graph.1
2020 View Synthesis from multi-view RGB data using multilayered representation and volumetric estimation
abstract
Aiming at free-view exploration of complicated scenes, this paper presents a method for interpolating views among multi RGB cameras. In this study, we combine the idea of cost volume, which represent 3D information, and 2D semantic segmentation of the scene, to accomplish view synthesis of complicated scenes. We use the idea of cost volume to estimate the depth and confidence map of the scene, and use a multi-layer representation and resolution of the data to optimize the view synthesis of the main object. /Conclusions By applying different treatment methods on different layers of the volume, we can handle complicated scenes containing multiple persons and plentiful occlusions. We also propose the view-interpolation→multi-view reconstruction→view interpolation pipeline to iteratively optimize the result. We test our method on varying data of multi-view scenes and generate decent results.
Zhaoqi Su, Tiansong Zhou, Kun Li 0001, David J. Brady, Yebin Liu
Virtual Real. Intell. Hardw.1
2018 Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
abstract
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.
Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.9
2017 BodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera
abstract
We propose BodyFusion, a novel real-time geometry fusion method that can track and reconstruct non-rigid surface motion of a human performance using a single consumer-grade depth camera. To reduce the ambiguities of the non-rigid deformation parameterization on the surface graph nodes, we take advantage of the internal articulated motion prior for human performance and contribute a skeleton-embedded surface fusion (SSF) method. The key feature of our method is that it jointly solves for both the skeleton and graph-node deformations based on information of the attachments between the skeleton and the graph nodes. The attachments are also updated frame by frame based on the fused surface geometry and the computed deformations. Overall, our method enables increasingly denoised, detailed, and complete surface reconstruction as well as the updating of the skeleton and attachments as the temporal depth frames are fused. Experimental results show that our method exhibits substantially improved nonrigid motion fusion performance and tracking robustness compared with previous state-of-the-art fusion methods. We also contribute a dataset for the quantitative evaluation of fusion-based dynamic scene reconstruction algorithms using a single depth camera.
Tao Yu 0007, Feng Xu 0005, Zhaoqi Su, Jianhui Zhao 0002, Qionghai Dai, Yebin Liu
ICCV5