Jiaran Cai

dblp:245/3322 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 85% Computer animation and physical simulation · 15%
Artificial intelligence
2 papers
3D vision · 63% Question answering and dialogue systems · 37%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
talking head generation
1.922026
Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025
Visual content generation and editing › video generation
long video generation
1.012026
Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026
Computer animation and physical simulation › character animation
multi-character animation
1.012026
Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026
Visual content generation and editing
video generation
1.012026
Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026
Visual content generation and editing › diffusion model
diffusion-based generation
0.912025
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025
Visual content generation and editing › image animation
portrait animation
0.912025
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
multi-choice reading comprehension
0.412019
Multi-Matching Network for Multiple Choice Reading Comprehension · AAAI 2019
Computer vision › 3D vision › correspondence estimation
semantic correspondence
0.412019
Multi-Matching Network for Multiple Choice Reading Comprehension · AAAI 2019
Computer vision › 3D vision › implicit neural representation
implicit 3d representation
0.312025
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025

Methods — techniques the papers use, named apart from their topics

motion disentanglement · 1.7emotion control module · 1.7diffusion model · 1.7reward feedback · 1.0diffusion transformer · 1.0classifier-free guidance · 1.0LoRA · 1.0multi-matching network · 0.4compose-match framework · 0.4
YearPublicationVenuePosition
2026 Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
abstract
Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync accuracy, temporal coherence for long video generation, and multi-character animation. In this work, we propose a diffusion transformer (DiT)-based framework for generating lifelike talking videos of arbitrary length, and introduce a training-free method for multi-character audio-driven animation. First, we employ a LoRA-based training strategy combined with a position shift inference approach, which enables efficient long video generation while preserving the capabilities of the foundation model. Moreover, we combine partial parameter updates with reward feedback to enhance both lip synchronization and natural body motion. Finally, we propose a training-free approach, Mask Classifier-Free Guidance (Mask-CFG), for multi-character animation, which requires no specialized datasets or model modifications and supports audio-driven animation for three or more characters. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving high-quality, temporally coherent, and multi-character audio-driven video generation in a simple, efficient, and cost-effective manner.
Xingpei Ma, Shenneng Huang, Jiaran Cai, Yuansheng Guan, Hanfeng Zhao, Shunsi Zhang
AAAI3
2025 Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
abstract
Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as inaccurate lip-sync, inappropriate head posture and the lack of fine-grained control over facial expressions. In order to introduce more face-guided conditions beyond speech audio clips, a novel two-stage training framework Playmate is proposed to generate more lifelike facial expressions and talking faces. In the first stage, we introduce a decoupled implicit 3D representation along with a meticulously designed motion-decoupled module to facilitate more accurate attribute disentanglement and generate expressive talking videos directly from audio cues. Then, in the second stage, we introduce an emotion-control module to encode emotion control information into the latent space, enabling fine-grained control over emotions and thereby achieving the ability to generate talking videos with desired emotion. Extensive experiments demonstrate that Playmate not only outperforms existing state-of-the-art methods in terms of video quality, but also exhibits strong competitiveness in lip synchronization while offering improved flexibility in controlling emotion and head pose. The code will be available at https://github.com/Playmate111/Playmate.
Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang, Shunsi Zhang
ICML2
2019 Multi-Matching Network for Multiple Choice Reading Comprehension
abstract
Multiple-choice machine reading comprehension is an important and challenging task where the machine is required to select the correct answer from a set of candidate answers given passage and question. Existing approaches either match extracted evidence with candidate answers shallowly or model passage, question and candidate answers with a single paradigm of matching. In this paper, we propose Multi-Matching Network (MMN) which models the semantic relationship among passage, question and candidate answers from multiple different paradigms of matching. In our MMN model, each paradigm is inspired by how human think and designed under a unified compose-match framework. To demonstrate the effectiveness of our model, we evaluate MMN on a large-scale multiple choice machine reading comprehension dataset (i.e. RACE). Empirical results show that our proposed model achieves a significant improvement compared to strong baselines and obtains state-of-the-art results.
Jiaran Cai, Hankui Zhuo
AAAI2