Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zezhao Tian

dblp:405/7605 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0003-0605-3867ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 50% Virtual and augmented reality · 50%
Artificial intelligence
2 papers
Vision and language · 79% Question answering and dialogue systems · 21%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › avatar
avatar animation
1.012026
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction · AAAI 2026
Computer animation and physical simulation
character animation
1.012026
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction · AAAI 2026
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.912025
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension · ACM Multimedia 2025
Natural language and speech › Question answering and dialogue systems
dialogue modeling
0.312026
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction · AAAI 2026
Computer vision › Vision and language › vision-language dataset
vision-language dataset construction
0.312025
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

responsive interaction module · 2.0multimodal conditioning · 2.0emotional intensity tags · 2.0text-adaptive multi-entity perceptron · 0.9large language model · 0.9entity inter-relationship reasoner · 0.9
YearPublicationVenuePosition
2026 VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
abstract
Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora incorporating head dynamics and fine-grained multi-modality annotations limits the application of dialogue modeling. Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive, and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners. Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we propose the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude. Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics.
Xingqun Qi, Bingkun Yang, Weile Chen, Zezhao Tian, Muyi Sun, Man Zhang 0005, Zhenan Sun
AAAI5
2025 ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC) aims to localize specified entities or regions from the source image according to the given natural language descriptions. While existing methods enable single-entity localization, they overlook modeling the complex inter-entity relationship in more practical multi-entity scenes, which limits their ability to produce accurate and reliable results. Moreover, the lack of high-quality multi-entity datasets incorporating fine-grained and paired image-text-relation annotations also limits addressing this challenge. To achieve this task, we first manually construct a relation-aware multi-entity REC dataset with fine-grained relation and text annotations, namely ReMeX. Additionally, we propose ReMeREC, a novel framework that effectively integrates textual and visual cues to localize multiple entities while capturing their inter-relationship. Specifically, to mitigate the semantic ambiguity arising from the absence of explicit entity boundaries in the source natural language description, we introduce a novel Text-adaptive Multi-entity Perceptron (TMP). TMP dynamically infers both the quantity and span of entities from corresponding fine-grained text cues, thus deriving representations that preserve the unique characteristics of each entity. Meanwhile, we design the Entity Inter-relationship Reasoner (EIR) to enhance semantic distinctiveness relationship modeling, leading to a more profound perception of the global scene. Furthermore, to better capture the fine-grained linguistic prompts for delineating multiple entity boundaries and inter-relationship, we leverage LLMs to generate a small-scale textual dataset, dubbed EntityText, which serves as an effective auxiliary resource and further improves the textual understanding. Extensive experiments conducted on four benchmark datasets demonstrate the superior performance of our framework. Remarkably, ReMeREC achieves outstanding results in multi-entity grounding and complex relationship prediction, outperforming other counterparts by a large margin.
Yizhi Hu, Zezhao Tian, Xingqun Qi, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang 0005, Zhenan Sun
ACM Multimedia2