Naiming Yao

dblp:362/2262 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0002-0797-9314ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 53% Virtual and augmented reality · 47%
Artificial intelligence
1 paper
Generative modeling · 50% Question answering and dialogue systems · 50%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
1.012026
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation · IEEE Trans. Vis. Comput. Graph. 2026
Machine learning › Generative modeling › face synthesis
talking face generation
1.012026
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation · IEEE Trans. Vis. Comput. Graph. 2026
Computer animation and physical simulation › facial animation
blendshape retargeting
0.812024
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters · VR 2024
Computer animation and physical simulation
facial animation
0.812024
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters · VR 2024
Virtual and augmented reality
virtual characters
0.812024
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters · VR 2024
Virtual and augmented reality › virtual reality
social virtual reality
0.312026
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation · IEEE Trans. Vis. Comput. Graph. 2026
Virtual and augmented reality
telepresence
0.312026
RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation · IEEE Trans. Vis. Comput. Graph. 2026
Human-AI interaction
human-in-the-loop
0.212024
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters · VR 2024

Methods — techniques the papers use, named apart from their topics

mesh-based facial motion · 2.0learnable query mechanism · 2.03d gaussian splatting · 2.0motion retargeting · 1.5deep learning · 1.5
YearPublicationVenuePosition
2026 RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
abstract
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS)-based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation, with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Our method supports applications such as VR telepresence, social VR, and embodied conversational agents. The socially-aware conditioning can also be extended to other human motion generation tasks. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released.
Peng Chen 0046, Xiaobao Wei, Yi Yang 0060, Naiming Yao, Hui Chen 0020, Feng Tian 0001
IEEE Trans. Vis. Comput. Graph.4
2025 CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation
abstract
Gaze estimation methods typically encounter significant performance degradation in generalized tasks due to the domain mismatch between the source and target domains. Existing approaches attempt to utilize various domain generalization techniques. However, their generalization capabilities are limited since they are constrained to a single visual modality. Notably, large-scale contrastive language-image pre-training (CLIP) models have been widely applied to downstream visual tasks for their robust generalization capabilities, but the potential of CLIP for regression tasks has not been fully explored. To bridge this gap, we introduce a novel framework called CR-CLIP, which endows CLIP with the capability to generalize gaze estimation. Specifically, we convert gaze labels into textual descriptions and achieve alignment between images and text signals with gaze cues, thereby extracting generalized gaze-related features. To enhance the model’s understanding of the numerical relationships of gaze directions, we propose a novel regression loss function based on image-text similarity. Additionally, we fine-tune the model on the original gaze dataset, achieving high precision in generalized gaze estimation. Experimental results show that our proposed method achieves state-of-the-art performance on four generalized gaze estimation tasks.
Yitong Zhu, Xurong Xie, Naiming Yao, Hui Chen 0020, Feng Tian 0001
ICASSP3
2024 Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters
abstract
Animating virtual characters has always been a fundamental research problem in virtual reality (VR). Facial animations play a crucial role as they effectively convey emotions and attitudes of virtual humans. However, creating such facial animations can be challenging, as current methods often involve utilization of expensive motion capture devices or significant investments of time and effort from human animators in tuning animation parameters. In this paper, we propose a holistic solution to automatically animate virtual human faces. In our solution, a deep learning model was first trained to retarget the facial expression from input face images to virtual human faces by estimating the blendshape coefficients. This method offers the flexibility of generating animations with characters of different appearances and blendshape topologies. Second, a practical toolkit was developed using Unity 3D, making it compatible with the most popular VR applications. The toolkit accepts both image and video as input to animate the target virtual human faces and enables users to manipulate the animation results. Furthermore, inspired by the spirit of Human-in-the-loop (HITL), we leveraged user feedback to further improve the performance of the model and toolkit, thereby increasing the customization properties to suit user preferences. The whole solution, for which we will make the code public, has the potential to accelerate the generation of facial animations for use in VR applications. https://github.com/showlab/BYOC
Zechen Bai, Peng Chen 0046, Xiaolan Peng, Naiming Yao, Hui Chen 0020
VR5