Kaifeng Zou

dblp:303/6522 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-4460-3690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A Unified Framework for Industrial Cel-Animation Colorization with Temporal-Structural Awareness
Xiaoyi Feng, Tao Huang 0022, Peng Wang 0168, Zizhou Huang, Haihang Zhang, Yuntao Zou, Dagang Li 0001, Kaifeng Zou
ICCV8
2025 4D Facial Expression Diffusion Model
abstract
Facial expression generation is one of the most challenging and long-sought aspects of character animation, with many interesting applications. The challenging task, traditionally having relied heavily on digital craftspersons, remains yet to be explored. In this article, we introduce a generative framework for generating 3D facial expression sequences (i.e., 4D faces) that can be conditioned on different inputs to animate an arbitrary 3D face mesh. It is composed of two tasks: (1) learning the generative model that is trained over a set of 3D landmark sequences and (2) generating 3D mesh sequences of an input facial mesh driven by the generated landmark sequences. The generative model is based on a Denoising Diffusion Probabilistic Model (DDPM), which has achieved remarkable success in generative tasks of other domains. While it can be trained unconditionally, its reverse process can still be conditioned by various condition signals. This allows us to efficiently develop several downstream tasks involving various conditional generations, by using expression labels, text, partial sequences, or simply a facial geometry. To obtain the full mesh deformation, we then develop a landmark-guided encoder-decoder to apply the geometrical deformation embedded in landmarks on a given facial mesh. Experiments show that our model has learned to generate realistic, quality expressions solely from the dataset of relatively small size, improving over the state-of-the-art methods. Videos and qualitative comparisons with other methods can be found at https://github.com/ZOUKaifeng/4DFM . Code and models will be made available upon acceptance.
Kaifeng Zou, Sylvain Faisan, Boyang Yu 0002, Sébastien Valette, Hyewon Seo
ACM Trans. Multim. Comput. Commun. Appl.1
2023 3D Facial Expression Generator Based on Transformer VAE
abstract
We present a generative model for the 3D facial expression mesh sequences, from onset to the termination of a desired expression. We tailor a Transformer VAE architecture: The encoder compresses a sequence of facial landmarks into an expression-aware regularized latent space, while the decoder generates a new sequence from the sampled latent variable, conditioned on a desired expression. After a landmark-guided mesh deformation, a given 3D neutral face is driven to an animated mesh sequence with the expected expression. The generated sequences are consistent, of quality, and exhibit a good level of diversity, improving over state-of-the-art methods. We validate our model by conducting extensive experiments on two representative datasets. The supplementary video and code are available on a GitHub page (https://github.com/ZOUKaifeng/FacialExpressionGeneration).
Kaifeng Zou, Boyang Yu 0002, Hyewon Seo
ICIP1
2023 Disentangling high-level factors and their features with conditional vector quantized VAEs
Kaifeng Zou, Sylvain Faisan, Fabrice Heitz, Sébastien Valette
Pattern Recognit. Lett.1
2023 Disentangled representations: towards interpretation of sex determination from hip bone
Kaifeng Zou, Sylvain Faisan, Fabrice Heitz, Marie Epain, Pierre Croisille, Laurent Fanton, Sébastien Valette
Vis. Comput.1
2022 Joint Disentanglement of Labels and Their Features with VAE
abstract
Most of previous semi-supervised methods that seek to obtain disentangled representations using variational autoencoders divide the latent representation into two components: the non-interpretable part and the disentangled part that explicitly models the factors of interest. With such models, features associated with high-level factors are not explicitly modeled, and they can either be lost, or at best entangled in the other latent variables, thus leading to bad disentanglement properties. To address this problem, we propose a novel conditional dependency structure where both the labels and their features belong to the latent space. We show using the CelebA dataset that the proposed model can learn meaningful representations, and we provide quantitative and qualitative comparisons with other approaches that show the effectiveness of the proposed method.
Kaifeng Zou, Sylvain Faisan, Fabrice Heitz, Sébastien Valette
ICIP1
2021 DSNet: Dynamic Skin Deformation Prediction by Recurrent Neural Network
Hyewon Seo, Kaifeng Zou, Frederic Cordier
CGI2