VLDB 2026 Research / reviewers in the wild / expert
Zhengming Yu
dblp:344/3515
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0003-0553-8125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Geometric modeling and processing · 40% Rendering · 27% Visual content generation and editing · 19% | |
| Artificial intelligence
5 papers |
Generative modeling · 60% 3D vision · 22% Robot manipulation · 18% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
3d shape generation |
1.6 | 2 | 2025 | SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation · SIGGRAPH Asia 2025 Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models · ECCV (39) 2024 |
Machine learning › Generative modeling
diffusion model |
1.0 | 2 | 2025 | Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models · ECCV (39) 2024 SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation · SIGGRAPH Asia 2025 |
Robotics › Robot manipulation › tactile sensing › force/tactile sensing
deformation sensing |
0.9 | 1 | 2025 | DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image · ICLR 2025 |
Computer vision › 3D vision › 3d generation
image-to-3d generation |
0.9 | 1 | 2025 | SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation · SIGGRAPH Asia 2025 |
Geometric modeling and processing › shape representation
surface representation |
0.9 | 1 | 2025 | SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation · SIGGRAPH Asia 2025 |
Visual content generation and editing
3d content creation |
0.8 | 1 | 2024 | Disentangled Clothed Avatar Generation from Text Descriptions · ECCV (52) 2024 |
Visual content generation and editing › 3d content creation
text-to-3d avatar generation |
0.8 | 1 | 2024 | Disentangled Clothed Avatar Generation from Text Descriptions · ECCV (52) 2024 |
Virtual and augmented reality › avatar
animatable avatar |
0.7 | 1 | 2023 | MonoHuman: Animatable Human Neural Field from Monocular Video · CVPR 2023 |
Geometric modeling and processing › deformation modeling
deformation field |
0.7 | 1 | 2023 | MonoHuman: Animatable Human Neural Field from Monocular Video · CVPR 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | MonoHuman: Animatable Human Neural Field from Monocular Video · CVPR 2023 |
Rendering
neural rendering |
0.7 | 1 | 2023 | DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering · ICCV 2023 |
Rendering
novel view synthesis |
0.7 | 1 | 2023 | DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering · ICCV 2023 |
Machine learning › Generative modeling › diffusion model › 3d shape generation
text-to-3d generation |
0.2 | 1 | 2024 | Disentangled Clothed Avatar Generation from Text Descriptions · ECCV (52) 2024 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.2 | 1 | 2023 | DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric Rendering · ICCV 2023 |
Rendering › novel view synthesis
free-viewpoint rendering |
0.2 | 1 | 2023 | MonoHuman: Animatable Human Neural Field from Monocular Video · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 3.0weakly supervised learning · 1.7transformer · 1.7multi-layer spherical projection · 1.72d diffusion prior · 1.7disentangled representation · 1.5multi-view capture · 1.3adversarial priors · 0.9adversarial prior · 0.9neural actor rendering · 0.7keyframe correspondence search · 0.7bidirectional deformation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageabstractReconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, complex deformations, and the ambiguity of the single-view setting. The previous state-of-the-art, Decaf, employs a global fitting optimization guided by contact and deformation estimation networks trained on studio-collected data with 3D annotations. However, Decaf suffers from a time-consuming optimization process and limited generalization capability due to its reliance on 3D annotations of hand-face interaction data. To address these issues, we present DICE, the first end-to-end method for Deformation-aware hand-face Interaction reCovEry from a single image. DICE estimates the poses of hands and faces, contacts, and deformations simultaneously using a Transformer-based architecture. It features disentangling the regression of local deformation fields and global mesh vertex locations into two network branches, enhancing deformation and contact estimation for precise and robust hand-face mesh recovery. To improve generalizability, we propose a weakly-supervised training approach that augments the training set using in-the-wild images without 3D ground-truth annotations, employing the depths of 2D keypoints estimated by off-the-shelf models and adversarial priors of poses for supervision. Our experiments demonstrate that DICE achieves state-of-the-art performance on a standard benchmark and in-the- wild data in terms of accuracy and physical plausibility. Additionally, our method operates at an interactive rate (20 fps) on an Nvidia 4090 GPU, whereas Decaf requires more than 15 seconds for a single image. The code will be available at: https://github.com/Qingxuan-Wu/DICE. Qingxuan Wu, Zhiyang Dou, Sirui Xu 0002, Soshi Shimada, Chen Wang 0049, Zhengming Yu, Yuan Liu 0025, Cheng Lin 0001, Zeyu Cao, Taku Komura, Vladislav Golyanik, Christian Theobalt, Wenping Wang 0001, Lingjie Liu |
ICLR | 6 |
| 2025 | SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape GenerationabstractExisting single-view 3D generative models typically adopt multiview diffusion priors to reconstruct object surfaces, yet they remain prone to inter-view inconsistencies and are unable to faithfully represent complex internal structure or nontrivial topologies. In particular, we encode geometry information by projecting it onto a bounding sphere and unwrapping it into a compact and structural multi-layer 2D Spherical Projection (SP) representation. Operating solely in the image domain, SPGen offers three key advantages simultaneously: (1) Consistency. The injective SP mapping encodes surface geometry with a single viewpoint which naturally eliminates view inconsistency and ambiguity; (2) Flexibility. Multi-layer SP maps represent nested internal structures and support direct lifting to watertight or open 3D surfaces; (3) Efficiency. The image-domain formulation allows the direct inheritance of powerful 2D diffusion priors and enables efficient finetuning with limited computational resources. Extensive experiments demonstrate that SPGen significantly outperforms existing baselines in geometric quality and computational efficiency. Jingdong Zhang 0003, Weikai Chen 0001, Yuan Liu 0025, Jionghao Wang, Zhengming Yu, Zhuowen Shen, Bo Yang 0070, Wenping Wang 0001, Xin Li 0003 |
SIGGRAPH Asia | 5 |
| 2024 | Disentangled Clothed Avatar Generation from Text Descriptions
Jionghao Wang, Yuan Liu 0025, Zhiyang Dou, Zhengming Yu, Yongqing Liang 0001, Cheng Lin 0001, Rong Xie 0004, Li Song 0001, Xin Li 0003, Wenping Wang 0001 |
ECCV (52) | 4 |
| 2024 | Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models
Zhengming Yu, Zhiyang Dou, Xiaoxiao Long, Cheng Lin 0001, Zekun Li 0002, Yuan Liu 0025, Norman Müller, Taku Komura, Marc Habermann, Christian Theobalt, Xin Li 0003, Wenping Wang 0001 |
ECCV (39) | 1 |
| 2023 | MonoHuman: Animatable Human Neural Field from Monocular VideoabstractAnimating virtual avatars with free-view control is crucial for various applications like virtual reality and digital entertainment. Previous studies have attempted to utilize the representation power of the neural radiance field (NeRF) to reconstruct the human body from monocular videos. Recent works propose to graft a deformation network into the NeRF to further model the dynamics of the human neural field for animating vivid human motions. However, such pipelines either rely on pose-dependent representations or fall short of motion coherency due to frameindependent optimization, making it difficult to generalize to unseen pose sequences realistically. In this paper, we propose a novel framework MonoHuman, which robustly renders view-consistent and high-fidelity avatars under arbitrary novel poses. Our key insight is to model the deformation field with bi-directional constraints and explicitly leverage the off-the-peg keyframe information to reason the feature correlations for coherent results. Specifically, we first propose a Shared Bidirectional Deformation module, which creates a pose-independent generalizable deformation field by disentangling backward and forward deformation correspondences into shared skeletal motion weight and separate non-rigid motions. Then, we devise a Forward Correspondence Search module, which queries the correspondence feature of keyframes to guide the rendering network. The rendered results are thus multi-view consistent with high fidelity, even under challenging novel pose settings. Extensive experiments demonstrate the superiority of our proposed MonoHuman over state-of-the-art methods. Zhengming Yu, Wayne Wu, Kwan-Yee Lin |
CVPR | 1 |
| 2023 | DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric RenderingabstractRealistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather impoverished in terms of diversity (e.g., outfit's fabric/material, body's interaction with objects, and motion sequences), which are crucial for rendering effect. Researchers are usually constrained to explore and evaluate a small set of rendering problems on current datasets, while real-world applications require methods to be robust across different scenarios. In this work, we present DNA-Rendering, a large-scale, high-fidelity repository of human performance data for neural actor rendering. DNA-Rendering presents several appealing attributes. First, our dataset contains over 1500 human subjects, 5000 motion sequences, and 67.5M frames' data volume. Upon the massive collections, we provide human subjects with grand categories of pose actions, body shapes, clothing, accessories, hairdos, and object intersection, which ranges the geometry and appearance variances from everyday life to professional occasions. Second, we provide rich assets for each subject – 2D/3D human body keypoints, foreground masks, SMPLX models, cloth/accessory materials, multi-view images, and videos. These assets boost the current method's accuracy on downstream rendering tasks. Third, we construct a professional multi-view system to capture data, which contains 60 synchronous cameras with max 4096 × 3000 resolution, 15 fps speed, and stern camera calibration steps, ensuring high-quality resources for task training and evaluation.Along with the dataset, we provide a large-scale and quantitative benchmark in full-scale, with multiple tasks to evaluate the existing progress of novel view synthesis, novel pose animation synthesis, and novel identity rendering methods. In this manuscript, we describe our DNA-Rendering effort as a revealing of new observations, challenges, and future directions to human-centric rendering. The dataset, code, and benchmarks will be publicly available at https://dna-rendering.github.io/. Ruixiang Chen, Siming Fan, Wanqi Yin, Zhongang Cai, Jingbo Wang 0003, Yang Gao 0042, Zhengming Yu, Zhengyu Lin, Daxuan Ren, Lei Yang 0045, Ziwei Liu 0002, Chen Change Loy, Chen Qian 0006, Wayne Wu, Dahua Lin, Bo Dai 0002, Kwan-Yee Lin |
ICCV | 9 |