Ye Zhu 0003

dblp:80/3703-3 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-9842-3690ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Identity-Preserving Video Dubbing Using Motion Warping
Runzhen Liu, Qinjie Lin, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Yu Li 0003, Chuhua Xian, Fa-Ting Hong
Int. J. Comput. Vis.5
2025 HRAvatar: High-Quality and Relightable Gaussian Head Avatar
abstract
Reconstructing animatable and high-quality 3D head avatars from monocular videos, especially with realistic relighting, is a valuable task. However, the limited information from single-view input, combined with the complex head poses and facial movements, makes this challenging. Previous methods achieve real-time performance by combining 3D Gaussian Splatting with a parametric head model, but the resulting head quality suffers from inaccurate face tracking and limited expressiveness of the deformation model. These methods also fail to produce realistic effects under novel lighting conditions. To address these issues, we propose HRAvatar, a 3DGS-based method that reconstructs high-fidelity, relightable 3D head avatars. HRA-vatar reduces tracking errors through end-to-end optimization and better captures individual facial deformations using learnable blendshapes and learnable linear blend skinning. Additionally, it decomposes head appearance into several physical properties and incorporates physically-based shading to account for environmental lighting. Extensive experiments demonstrate that HRAvatar not only reconstructs superior-quality heads but also achieves realistic visual effects under varying lighting conditions. Video results and code are available at the project page.
Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Kangjie Chen, Minghan Qin, Yu Li 0003, Haoqian Wang
CVPR4
2025 Canonswap: High-Fidelity and Consistent Video Face Swapping Via Canonical Space Modulation
Ye Zhu 0003, Yunfei Liu 0001, Lijian Lin, Cong Wan, Zijian Cai, Yu Li 0003, Shao-Lun Huang
ICCV2
2025 GUAVA: Generalizable Upper Body 3D Gaussian Avatar
abstract
Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual IDs, which is both complex and time-consuming. Furthermore, limited by SMPLX's expressiveness, these methods often focus on body motion but struggle with facial expressions. To address these challenges, we first introduce an expressive human model (EHM) to enhance facial expression capabilities and develop an accurate tracking method. Based on this template model, we propose GUAVA, the first framework for fast animatable upper-body 3D Gaussian avatar reconstruction. We leverage inverse texture mapping and projection sampling techniques to infer Ubody (upper-body) Gaussians from a single image. The rendered images are refined through a neural refiner. Experimental results demonstrate that GUAVA significantly outperforms previous methods in rendering quality and offers significant speed improvements, with reconstruction times in the sub-second range (0.1s), and supports real-time animation and rendering.
Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Minghan Qin, Yu Li 0003, Haoqian Wang
ICCV4
2025 Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning
abstract
This paper presents Qffusion, a dual-frame-guided framework for portrait video editing. Specifically, we consider a design principle of "animation for editing", and train Qffusion as a general animation framework from two still reference images while we can use it for portrait video editing easily by applying modified start and end frames as references during inference. Leveraging the powerful generative power of Stable Diffusion, we propose a Quadrant-grid Arrangement (QGA) scheme for latent re-arrangement, which arranges the latent codes of two reference images and that of four facial conditions into a four-grid fashion, separately. Then, we fuse features of these two modalities and use self-attention for both appearance and temporal learning, where representations at different times are jointly modeled under QGA. Our Qffusion can achieve stable video editing without additional networks or complex training stages, where only the input format of Stable Diffusion is modified. Further, we propose a Quadrant-grid Propagation (QGP) inference strategy, which enjoys a unique advantage on stable arbitrary-length video generation by processing reference and condition frames recursively. Through extensive experiments, Qffusion consistently outperforms state-of-the-art techniques on portrait video editing.
Maomao Li, Lijian Lin, Yunfei Liu 0001, Ye Zhu 0003, Yu Li 0003
IEEE Trans. Vis. Comput. Graph.4
2022 Robust Human Matting via Semantic Guidance
Xiangguang Chen, Ye Zhu 0003, Yu Li 0003, Bingtao Fu, Lei Sun 0009, Ying Shan, Shan Liu 0001
ACCV (2)2
2022 Composite Photograph Harmonization with Complete Background Cues
abstract
Compositing portrait photographs or videos to novel backgrounds is an important application in computational photography. Seamless blending along boundaries and globally harmonic colors are two desired properties of the photo-realistic composition of foregrounds and new backgrounds. Existing works are dedicated to either foreground alpha matte generation or after-blending harmonization, leading to sub-optimal background replacement when putting foregrounds and backgrounds together. In this work, we unify the two objectives in a single framework to obtain realistic portrait image composites. Specifically, we investigate the usage of a target background and find that a complete background plays a vital role in both seamlessly blending and harmonization. We develop a network to learn the composition process given an imperfect alpha matte with appearance features extracted from the complete background to adjust color distribution. Our dedicated usage of a complete background enables realistic portrait image composition and also temporally stable results on videos. Extensive quantitative and qualitative experiments on both synthetic and real-world data demonstrate that our method achieves state-of-the-art performance.
Yazhou Xing, Yu Li 0003, Xintao Wang 0002, Ye Zhu 0003, Qifeng Chen 0001
ACM Multimedia4
2022 Hybrid Warping Fusion for Video Frame Interpolation
Yu Li 0003, Ye Zhu 0003, Ruoteng Li, Xintao Wang 0002, Ying Shan
Int. J. Comput. Vis.2
2021 Attentive deep network for blind motion deblurring on dynamic scenes
Yong Xu 0007, Ye Zhu 0003, Yuhui Quan, Hui Ji 0002
Comput. Vis. Image Underst.2