Zuhao Yang

dblp:338/9785 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 31% Generative modeling · 31% Language models and text generation · 18%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › video captioning
dense video captioning
0.912025
Timeexpert: an Expert-Guided Video Llm for Video Temporal Grounding · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.912025
Versatile Transition Generation with Image-to-Video Diffusion · ICCV 2025
Machine learning › Generative modeling › diffusion model › video diffusion model
image-to-video diffusion
0.912025
Versatile Transition Generation with Image-to-Video Diffusion · ICCV 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
QAEval: Mixture of Evaluators for Question-Answering Task Evaluation · ACL (1) 2025
Computer vision › Vision and language
moment retrieval
0.912025
Timeexpert: an Expert-Guided Video Llm for Video Temporal Grounding · ICCV 2025
Natural language and speech › Question answering and dialogue systems
question answering evaluation
0.912025
QAEval: Mixture of Evaluators for Question-Answering Task Evaluation · ACL (1) 2025
Computer vision › Vision and language
temporal grounding
0.912025
Timeexpert: an Expert-Guided Video Llm for Video Temporal Grounding · ICCV 2025
Computer vision › Video understanding and tracking › video summarization
video highlight detection
0.912025
Timeexpert: an Expert-Guided Video Llm for Video Temporal Grounding · ICCV 2025
Machine learning › Generative modeling › video generation › video frame synthesis
video synthesis
0.912025
Versatile Transition Generation with Image-to-Video Diffusion · ICCV 2025
Visual content generation and editing
video generation
0.912025
Versatile Transition Generation with Image-to-Video Diffusion · ICCV 2025
Natural language and speech › Language models and text generation
text generation evaluation
0.712023
FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-Entropy · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

representation alignment regularization · 1.7interpolation-based initialization · 1.7mixture of experts · 1.7fourier analysis · 1.3cross-entropy estimation · 1.3video large language model · 0.9dynamic load balancing · 0.9
YearPublicationVenuePosition
2025 QAEval: Mixture of Evaluators for Question-Answering Task Evaluation
abstract
Question answering (QA) tasks serve as a key benchmark for evaluating generation systems.Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with openended and unstructured responses.LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to instructions, robustness issues, and high computational costs.To overcome these challenges, we introduce QAEval, a hybrid framework combining rule-based reliability with LLM-based adaptability.QAEval utilizes two high-quality datasets: QAExtract for short-answer extraction and QAScore for scoring model training.By integrating a Mixture of Evaluators model with Dynamic Load Balancing Optimization, QAEval enables accurate, cost-effective QA evaluation.Experimental results show it outperforms models like GPT-4o and Claude-3, achieving 92.3% accuracy with only 0.6B parameters.
Tan Yue, Rui Mao 0010, Xuzhao Shi, Shuo Zhan, Zuhao Yang, Dongyan Zhao 0001
ACL (1)5
2025 Timeexpert: an Expert-Guided Video Llm for Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) aims to precisely identify video event segments in response to textual queries. The outputs of VTG tasks manifest as sequences of events, each defined by precise timestamps, saliency scores, and textual descriptions. Despite recent advances, a fundamental limitation persists in existing Video Large Language Models (Video-LLMs): they process all task tokens through identical and static pathways, failing to recognize that temporal localization, saliency assessment, and textual generation represent fundamentally distinct tasks requiring specialized processing. To address this, we introduce TimeExpert, a Mixture-of-Experts (MoE)-based Video-LLM that effectively decomposes VTG tasks by dynamically routing task-specific tokens (e.g., timestamps, saliency scores) to specialized experts, with increased computational efficiency. Our design choices enable precise handling of each subtask, leading to improved event modeling across diverse VTG applications. Extensive experiments demonstrate that TimeExpert consistently achieves state-of-the-art performance on various VTG tasks such as Dense Video Captioning, Moment Retrieval, and Video Highlight Detection.
Zuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu, Song Bai 0001
ICCV1
2025 Versatile Transition Generation with Image-to-Video Diffusion
abstract
Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos given the first and last video frames as well as descriptive text prompts is far underexplored. We present VTG, a Versatile Transition video Generation framework that can generate smooth, high-fidelity, and semantically coherent video transitions. VTG introduces interpolation-based initialization that helps preserve object identity and handle abrupt content changes effectively. In addition, it incorporates dual-directional motion fine-tuning and representation alignment regularization to mitigate the limitations of pre-trained image-to-video diffusion models in motion smoothness and generation fidelity, respectively. To evaluate VTG and facilitate future studies on unified transition generation, we collected TransitBench, a comprehensive benchmark for transition generation covering two representative transition tasks: concept blending and scene transition. Extensive experiments show that VTG achieves superior transition performance consistently across all four tasks.
Zuhao Yang, Yingchen Yu, Shijian Lu, Song Bai 0001
ICCV1
2023 FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-Entropy
abstract
Measuring the distance between machine-produced and human language is a critical open problem. Inspired by empirical findings from psycholinguistics on the periodicity of entropy in language, we propose FACE, a set of metrics based on Fourier Analysis of the estimated Cross-Entropy of language, for measuring the similarity between model-generated and human-written languages. Based on an open-ended generation task and the experimental data from previous studies, we find that FACE can effectively identify the human-model gap, scales with model size, reflects the outcomes of different sampling methods for decoding, correlates well with other evaluation metrics and with human judgment scores.
Zuhao Yang, Yingfang Yuan, Shuo Zhan, Huajun Bai
NeurIPS1