Yujin Chai

dblp:303/0544 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-5525-6527ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Face, body and person analysis · 50% Representation and self-supervised learning · 50%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
facial expression analysis
1.622025
A Semantic Talking Style Space for Speech-Driven Facial Animation · IEEE Trans. Vis. Comput. Graph. 2025
Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement · IEEE Trans. Vis. Comput. Graph. 2024
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
style-content disentanglement
1.622025
A Semantic Talking Style Space for Speech-Driven Facial Animation · IEEE Trans. Vis. Comput. Graph. 2025
Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement · IEEE Trans. Vis. Comput. Graph. 2024
Computer animation and physical simulation
facial animation
1.622025
A Semantic Talking Style Space for Speech-Driven Facial Animation · IEEE Trans. Vis. Comput. Graph. 2025
Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement · IEEE Trans. Vis. Comput. Graph. 2024
Computer animation and physical simulation › facial animation
speech-driven facial animation
1.622025
A Semantic Talking Style Space for Speech-Driven Facial Animation · IEEE Trans. Vis. Comput. Graph. 2025
Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement · IEEE Trans. Vis. Comput. Graph. 2024

Methods — techniques the papers use, named apart from their topics

self-supervision · 1.7deep neural network · 1.7two-pass style swapping · 1.5joint training with shared decoder · 1.5
YearPublicationVenuePosition
2025 A Semantic Talking Style Space for Speech-Driven Facial Animation
abstract
We present a latent talking style space with semantic meanings for speech-driven 3D facial animation. The style space is learned from 3D speech facial animations via a self-supervision paradigm without any style labeling, leading to an automatic separation of high-level attributes, i.e., different channels of the latent style code possess different semantic meanings, such as a wide/slightly open mouth, a grinning/round mouth, and frowning/raising eyebrows. The style space enables intuitive and flexible control of talking styles in speech-driven facial animation through manipulating the channels of style code. To effectively learn such a style space, we propose a two-stage approach, involving two deep neural networks, to disentangle the person identity, speech content, and talking style contained in 3D speech facial animations. The training is performed on a novel dataset of 3D talking faces of various styles, constructed from over ten hours of videos of 200 subjects collected from the Internet.
Yujin Chai, Yanlin Weng, Tianjia Shao, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Personalized Audio-Driven 3D Facial Animation via Style-Content Disentanglement
abstract
We present a learning-based approach for generating 3D facial animations with the motion style of a specific subject from arbitrary audio inputs. The subject style is learned from a video clip (1-2 minutes) either downloaded from the Internet or captured through an ordinary camera. Traditional methods often require many hours of the subject's video to learn a robust audio-driven model and are thus unsuitable for this task. Recent research efforts aim to train a model from video collections of a few subjects but ignore the discrimination between the subject style and underlying speech content within facial motions, leading to inaccurate style or articulation. To solve the problem, we propose a novel framework that disentangles subject-specific style and speech content from facial motions. The disentanglement is enabled by two novel training mechanisms. One is two-pass style swapping between two random subjects, and the other is joint training of the decomposition network and audio-to-motion network with a shared decoder. After training, the disentangled style is combined with arbitrary audio inputs to generate stylized audio-driven 3D facial animations. Compared with start-of-the-art methods, our approach achieves better results qualitatively and quantitatively, especially in difficult cases like bilabial plosive and bilabial nasal phonemes.
Yujin Chai, Tianjia Shao, Yanlin Weng, Kun Zhou 0001
IEEE Trans. Vis. Comput. Graph.1
2022 Speech-driven facial animation with spectral gathering and temporal attention
Yujin Chai, Yanlin Weng, Lvdi Wang, Kun Zhou 0001
Frontiers Comput. Sci.1