VLDB 2026 Research / reviewers in the wild / expert
Jiaran Cai
dblp:245/3322
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 85% Computer animation and physical simulation · 15% | |
| Artificial intelligence
2 papers |
3D vision · 63% Question answering and dialogue systems · 37% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing
talking head generation |
1.9 | 2 | 2026 | Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026 Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025 |
Visual content generation and editing › video generation
long video generation |
1.0 | 1 | 2026 | Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026 |
Computer animation and physical simulation › character animation
multi-character animation |
1.0 | 1 | 2026 | Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026 |
Visual content generation and editing
video generation |
1.0 | 1 | 2026 | Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback · AAAI 2026 |
Visual content generation and editing › diffusion model
diffusion-based generation |
0.9 | 1 | 2025 | Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025 |
Visual content generation and editing › image animation
portrait animation |
0.9 | 1 | 2025 | Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
multi-choice reading comprehension |
0.4 | 1 | 2019 | Multi-Matching Network for Multiple Choice Reading Comprehension · AAAI 2019 |
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
0.4 | 1 | 2019 | Multi-Matching Network for Multiple Choice Reading Comprehension · AAAI 2019 |
Computer vision › 3D vision › implicit neural representation
implicit 3d representation |
0.3 | 1 | 2025 | Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
motion disentanglement · 1.7emotion control module · 1.7diffusion model · 1.7reward feedback · 1.0diffusion transformer · 1.0classifier-free guidance · 1.0LoRA · 1.0multi-matching network · 0.4compose-match framework · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward FeedbackabstractRecent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync accuracy, temporal coherence for long video generation, and multi-character animation. In this work, we propose a diffusion transformer (DiT)-based framework for generating lifelike talking videos of arbitrary length, and introduce a training-free method for multi-character audio-driven animation. First, we employ a LoRA-based training strategy combined with a position shift inference approach, which enables efficient long video generation while preserving the capabilities of the foundation model. Moreover, we combine partial parameter updates with reward feedback to enhance both lip synchronization and natural body motion. Finally, we propose a training-free approach, Mask Classifier-Free Guidance (Mask-CFG), for multi-character animation, which requires no specialized datasets or model modifications and supports audio-driven animation for three or more characters. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving high-quality, temporally coherent, and multi-character audio-driven video generation in a simple, efficient, and cost-effective manner. Xingpei Ma, Shenneng Huang, Jiaran Cai, Yuansheng Guan, Hanfeng Zhao, Shunsi Zhang |
AAAI | 3 |
| 2025 | Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided DiffusionabstractRecent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as inaccurate lip-sync, inappropriate head posture and the lack of fine-grained control over facial expressions. In order to introduce more face-guided conditions beyond speech audio clips, a novel two-stage training framework Playmate is proposed to generate more lifelike facial expressions and talking faces. In the first stage, we introduce a decoupled implicit 3D representation along with a meticulously designed motion-decoupled module to facilitate more accurate attribute disentanglement and generate expressive talking videos directly from audio cues. Then, in the second stage, we introduce an emotion-control module to encode emotion control information into the latent space, enabling fine-grained control over emotions and thereby achieving the ability to generate talking videos with desired emotion. Extensive experiments demonstrate that Playmate not only outperforms existing state-of-the-art methods in terms of video quality, but also exhibits strong competitiveness in lip synchronization while offering improved flexibility in controlling emotion and head pose. The code will be available at https://github.com/Playmate111/Playmate. Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang, Shunsi Zhang |
ICML | 2 |
| 2019 | Multi-Matching Network for Multiple Choice Reading ComprehensionabstractMultiple-choice machine reading comprehension is an important and challenging task where the machine is required to select the correct answer from a set of candidate answers given passage and question. Existing approaches either match extracted evidence with candidate answers shallowly or model passage, question and candidate answers with a single paradigm of matching. In this paper, we propose Multi-Matching Network (MMN) which models the semantic relationship among passage, question and candidate answers from multiple different paradigms of matching. In our MMN model, each paradigm is inspired by how human think and designed under a unified compose-match framework. To demonstrate the effectiveness of our model, we evaluate MMN on a large-scale multiple choice machine reading comprehension dataset (i.e. RACE). Empirical results show that our proposed model achieves a significant improvement compared to strong baselines and obtains state-of-the-art results. Jiaran Cai, Hankui Zhuo |
AAAI | 2 |