EDBT 2026 Demo / reviewers in the wild / expert
Xuan Wang 0024
dblp:34/4799-24
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0003-2777-0179ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Computer animation and physical simulation · 64% Visual content generation and editing · 32% Virtual and augmented reality · 4% | |
| Artificial intelligence
1 paper |
Segmentation and scene understanding · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Immersive interaction · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer animation and physical simulation
facial animation |
1.6 | 2 | 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style Preservation · IEEE Trans. Vis. Comput. Graph. 2025 Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusion · IEEE Trans. Vis. Comput. Graph. 2024 |
Computer animation and physical simulation › facial animation
speech-driven facial animation |
0.9 | 1 | 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style Preservation · IEEE Trans. Vis. Comput. Graph. 2025 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.8 | 1 | 2024 | CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque Segmentation · IEEE Trans. Multim. 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque Segmentation · IEEE Trans. Multim. 2024 |
Computer animation and physical simulation › facial animation
3d facial animation |
0.8 | 1 | 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusion · IEEE Trans. Vis. Comput. Graph. 2024 |
Immersive interaction
avatar |
0.3 | 1 | 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style Preservation · IEEE Trans. Vis. Comput. Graph. 2025 |
Virtual and augmented reality
immersive media |
0.2 | 1 | 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusion · IEEE Trans. Vis. Comput. Graph. 2024 |
Methods — techniques the papers use, named apart from their topics
transformer decoder · 1.7style disentanglement · 1.7motion encoder · 1.7vector quantized variational autoencoder · 0.8transformer · 0.8pyramid fusion · 0.8latent diffusion model · 0.8cluster center attention · 0.8audio synchronization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style PreservationabstractIt is a challenging task to create realistic 3D avatars that accurately replicate individuals' speech and unique talking styles for speech-driven facial animation. Existing techniques have made remarkable progress but still struggle to achieve lifelike mimicry. This article proposes "TalkingStyle", a novel method to generate personalized talking avatars while retaining the talking style of the person. Our approach uses a set of audio and animation samples from an individual to create new facial animations that closely resemble their specific talking style, synchronized with speech. We disentangle the style codes from the motion patterns, allowing our method to associate a distinct identifier with each person. To manage each aspect effectively, we employ three separate encoders for style, speech, and motion, ensuring the preservation of the original style while maintaining consistent motion in our stylized talking avatars. Additionally, we propose a new style-conditioned transformer decoder, offering greater flexibility and control over the facial avatar styles. We comprehensively evaluate TalkingStyle through qualitative and quantitative assessments, as well as user studies demonstrating its superior realism and lip synchronization accuracy compared to current state-of-the-art methods. Wenfeng Song, Xuan Wang 0024, Shi Zheng, Shuai Li 0001, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | FusionCraft: Fusing Emotion and Identity in Cross-Modal 3D Facial Animation
Zhenyu Lv, Xuan Wang 0024, Wenfeng Song, Xia Hou |
ICIC (10) | 2 |
| 2024 | CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque SegmentationabstractDental plaque segmentation is crucial for maintaining oral health. However, accurately segmenting dental plaque in unconstrained environments can be challenging due to its low contrast and high variability in appearance. While existing transformer-based networks rely on attention mechanisms for each pixel, they do not take into account the relationships between neighboring pixels. Consequently, feature extraction is limited, making it difficult to achieve accurate segmentation of low-contrast images. To address this issue, we propose a simple yet efficient cluster center transformer that improves dental plaque segmentation by clustering image pixels based on multiple levels of feature maps' intensity and texture information. By grouping similar pixels into regions, the proposed method enables the transformers to focus on the local contour and edge around the teeth regions, adapting to the low contrast and high variability of plaque appearance, leading to more accurate and efficient segmentation of dental plaque in dental images. Additionally, we designed Multiple Granularity Perceptions using a pyramid fusion mechanism to capture multiple scales of vision features, thereby enhancing the low-contrast vision features. The proposed method can benefit the dental diagnosis and treatment planning process by improving the accuracy and efficiency of dental plaque segmentation. Our proposed method achieved state-of-the-art results on the dental plaque dataset (Li et al., 2020), with intersection over union (IoU) of 60.91% and pixel accuracy (PA) of 76.81%, all of which were the highest among all methods, demonstrating its effectiveness in plaque segmentation in unconstrained environments. Wenfeng Song, Xuan Wang 0024, Shuai Li 0001, Aimin Hao |
IEEE Trans. Multim. | 2 |
| 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusionabstract3D Facial animations, crucial to augmented and mixed reality digital media, have evolved from mere aesthetic elements to potent storytelling media. Despite considerable progress in facial animation of neutral emotions, existing methods still struggle to capture the authenticity of emotions. This paper introduces a novel approach to capture fine facial expressions and generate facial animations using audio synchronization. Our method consists of two key components: First, the Local-to-global Latent Diffusion Model (LG-LDM) tailored for authentic facial expressions, which can integrate audio, time step, facial expressions, and other conditions towards possible encoding of emotionally rich yet latent features in response to possibly noisy raw audio signals. The core of LG-LDM is our carefully designed Facial Denoiser Model (FDM) for aligning the local-to-global animation feature with audio. Second, we redesign an Emotion-centric Vector Quantized-Variational AutoEncoder framework (EVQ-VAE) to finely decode the subtle differences under different emotions and reconstruct the final 3D facial geometry. Our work significantly contributes to the key challenges of emotionally realistic 3D facial animation for audio synchronization and enhances the immersive experience and emotional depth in augmented and mixed reality applications. We provide a reproducibility kit including our code, dataset, and detailed instructions for running the experiments. This kit is available at https://github.com/wangxuanx/Face-Diffusion-Model. Wenfeng Song, Xuan Wang 0024, Yiming Jiang 0018, Shuai Li 0001, Aimin Hao, Xia Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |