VLDB 2026 Research / reviewers in the wild / expert
Ziyi Chen 0005
dblp:37/1439-5
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0009-4064-324XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementabstractCo-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS key-point movements, which often lead to artifacts like blurry hands and distorted fingers. In response to these challenges, we present the Implicit Motion-Audio Entanglement (IMAE) method for co-speech gesture video generation. IMAE strengthens audio control by entangling implicit motion parameters, including pose and expression, with audio inputs. Our method utilizes a two-branch framework that combines an audio-to-motion generation branch with a video diffusion branch, enabling realistic gesture generation without requiring additional inputs during inference. To improve training efficiency, we propose a two-stage slow-fast training strategy that balances memory constraints while facilitating the learning of meaningful gestures from long frame sequences. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple metrics. Project Page. Xinjie Li 0002, Ziyi Chen 0005, Xinlu Yu, Iek-Heng Chu, Peng Chang 0002, Jing Xiao 0006 |
CVPR | 2 |
| 2025 | Data-Free Knowledge Distillation with Diffusion ModelsabstractRecently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any access to training data. Although diffusion models are adept at synthesizing high-fidelity photorealistic images across various domains, existing methods cannot be easiliy implemented to DFKD. To bridge that gap, this paper proposes a novel approach based on diffusion models, DiffDFKD. Specifically, DiffDFKD involves targeted optimizations in two key areas. Firstly, DiffDFKD utilizes valuable information from teacher models to guide the pre-trained diffusion models’ data synthesis, generating datasets that mirror the training data distribution and effectively bridge domain gaps. Secondly, to reduce computational burdens, DiffDFKD introduces Latent CutMix Augmentation, an efficient technique, to enhance the diversity of diffusion model-generated images for DFKD while preserving key attributes for effective knowledge transfer. Extensive experiments validate the efficacy of DiffDFKD, yielding state-of-the-art results exceeding existing DFKD approaches.We release our code at https://github.com/xhqi0109/DiffDFKD. Xiaohua Qi, Renda Li, Qiang Ling 0001, Jun Yu 0001, Ziyi Chen 0005, Peng Chang 0002, Jing Xiao 0006 |
ICME | 6 |
| 2025 | Input-Aware Sparse Attention for Real-Time Co-Speech Video GenerationabstractDiffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps and costly attention mechanisms, preventing real-time deployment. In this work, we distill a many-step diffusion video model into a few-step student model. Unfortunately, directly applying recent diffusion distillation methods degrades video quality and falls short of real-time performance. To address these issues, our new video distillation method leverages input human pose conditioning for both attention and loss functions. We first propose using accurate correspondence between input human pose keypoints to guide attention to relevant regions, such as the speaker’s face, hands, and upper body. This input-aware sparse attention reduces redundant computations and strengthens temporal correspondences of body parts, improving inference efficiency and motion coherence. To further enhance visual quality, we introduce an input-aware distillation loss that improves lip synchronization and hand motion realism. By integrating our input-aware sparse attention and distillation loss, our method achieves real-time performance with improved visual quality compared to recent audio-driven and input-driven methods. We also conduct extensive experiments showing the effectiveness of our algorithmic design choices. Beijia Lu, Ziyi Chen 0005, Jing Xiao 0006, Jun-Yan Zhu |
SIGGRAPH Asia | 2 |
| 2025 | SyncDiff: Diffusion-Based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved SynchronizationabstractTalking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image fidelity. Recent studies demonstrate that GAN-based and diffusion-based models achieve state-of-the-art (SOTA) performance on this task, with diffusion-based models achieving superior image fidelity but experiencing lower synchronization compared to their GAN-based counterparts. To this end, we propose SYNcDIFF, a simple yet effective approach to improve diffusion-based models using a temporal pose frame with information bottleneck and facial-informative audio features extracted from AVHuBERT, as conditioning input into the diffusion process. We evaluate SYNcDIFF on two canonical talking head datasets, LRS2 and LRS3 for direct comparison with other SOTA models. Experiments on LRS2/LRS3 datasets show that SYNcDIFF achieves a synchronization score 27.7%/62.3% relatively higher than previous diffusion-based methods, while preserving their high-fidelity characteristics. Xulin Fan, Heting Gao, Ziyi Chen 0005, Peng Chang 0002, Mark Hasegawa-Johnson |
WACV | 3 |
| 2024 | Co-speech Gesture Video Generation with 3D Human Meshes
Aniruddha Mahapatra, Renda Li, Ziyi Chen 0005, Boyang Ding, Shoulei Wang, Jun-Yan Zhu, Peng Chang 0002, Jing Xiao 0006 |
ECCV (89) | 4 |
| 2022 | Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation AssessmentabstractAutomatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous efforts typically only model one aspect (e.g., accuracy) at one granularity (e.g., at the phoneme-level). In this work, we explore modeling multi-aspect pronunciation assessment at multiple granularities. Specifically, we train a Goodness Of Pronunciation feature-based Transformer (GOPT) with multi-task learning. Experiments show that GOPT achieves the best results on speechocean762 with a public automatic speech recognition (ASR) acoustic model trained on Librispeech. Yuan Gong 0001, Ziyi Chen 0005, Iek-Heng Chu, Peng Chang 0002, James R. Glass |
ICASSP | 2 |