Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jianzhi Long

dblp:321/1590 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 77% Computer animation and physical simulation · 23%
Artificial intelligence
1 paper
Generative modeling · 62% Efficient and distributed learning · 38%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
talking head generation
2.022026
SimDEM: Audio-Driven Emotional Talking Head Generation With Simplified and Decoupled Expression Modeling · IEEE Trans. Multim. 2026
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation · AAAI 2026
Machine learning › Generative modeling › diffusion model
diffusion model acceleration
1.012026
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation · AAAI 2026
Visual content generation and editing › talking head generation
emotional talking head generation
1.012026
SimDEM: Audio-Driven Emotional Talking Head Generation With Simplified and Decoupled Expression Modeling · IEEE Trans. Multim. 2026
Computer animation and physical simulation › facial animation
speech-driven facial animation
1.012026
SimDEM: Audio-Driven Emotional Talking Head Generation With Simplified and Decoupled Expression Modeling · IEEE Trans. Multim. 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.312026
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
token pruning
0.312026
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation · AAAI 2026
Visual content generation and editing › face editing
facial expression synthesis
0.312026
SimDEM: Audio-Driven Emotional Talking Head Generation With Simplified and Decoupled Expression Modeling · IEEE Trans. Multim. 2026

Methods — techniques the papers use, named apart from their topics

feature caching · 2.0decoupled foreground attention · 2.0key expression modeling · 1.0decoupled expression modeling · 1.0
YearPublicationVenuePosition
2026 Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
abstract
Diffusion-based talking head models generate high-quality, photorealistic videos but suffer from slow inference, limiting practical applications. Existing acceleration methods for gen- eral diffusion models fail to exploit the temporal and spatial redundancies unique to talking head generation. In this paper, we propose a task-specific framework addressing these inefficiencies through two key innovations. First, we introduce Lightning-fast Caching-based Parallel denoising prediction (LightningCP), caching static features to bypass most model layers in inference time. We also enable parallel prediction using cached features and estimated noisy latents as inputs, efficiently bypassing sequential sampling. Second, we propose Decoupled Foreground Attention (DFA) to further accelerate attention computations, exploiting the spatial decoupling in talking head videos to restrict attention to dynamic foreground regions. Additionally, we remove reference features in certain layers to bring extra speedup. Extensive experiments demonstrate that our framework significantly improves inference speed while preserving video quality.
Jianzhi Long, Rongcheng Tu, Dacheng Tao
AAAI1
2026 SimDEM: Audio-Driven Emotional Talking Head Generation With Simplified and Decoupled Expression Modeling
abstract
Emotional talking head generation has advanced significantly because of its potential to enhance naturalness and expressiveness in applications like gaming, virtual reality, and video conferencing. However, existing methods struggle with overcomplicated facial expression modeling, which introduces redundancies, reduces learning efficiency, and limits emotion accuracy. To address these limitations, we propose SimDEM, a novel framework for emotional talking head generation that leverages simplified and decoupled expression modeling. By retaining only critical dimensions of the expression representation via key expression modeling, we reduce noise in training and improve emotion accuracy. Furthermore, the decoupling of facial regions into independent dynamics for the upper and lower face enables more effective learning of localized emotional expressions, enhancing the visual fidelity of the output. Extensive evaluations demonstrate that SimDEM outperforms state-of-the-art methods, achieving superior emotion accuracy and visual quality in synthesizing expressive talking head videos.
Jianzhi Long, Rongcheng Tu, Dacheng Tao
IEEE Trans. Multim.1
2025 PointWavelet: Learning in Spectral Domain for 3-D Point Cloud Analysis
abstract
With recent success of deep learning in 2-D visual recognition, deep-learning-based 3-D point cloud analysis has received increasing attention from the community, especially due to the rapid development of autonomous driving technologies. However, most existing methods directly learn point features in the spatial domain, leaving the local structures in the spectral domain poorly investigated. In this article, we introduce a new method, PointWavelet, to explore local graphs in the spectral domain via a learnable graph wavelet transform. Specifically, we first introduce the graph wavelet transform to form multiscale spectral graph convolution to learn effective local structural representations. To avoid the time-consuming spectral decomposition, we then devise a learnable graph wavelet transform, which significantly accelerates the overall training process. Extensive experiments on four popular point cloud datasets, ModelNet40, ScanObjectNN, ShapeNet-Part, and S3DIS, demonstrate the effectiveness of the proposed method on point cloud classification and segmentation.
Cheng Wen 0001, Jianzhi Long, Baosheng Yu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.2