Dingyi Yang

dblp:266/2264 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 49% Language models and text generation · 46% Knowledge representation and reasoning · 2%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Integrated circuit design · 100%

Topics — the 17 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.922026
HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External Knowledge · ACL (1) 2026
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation › text generation evaluation
story evaluation
0.912025
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025
Computer vision › Vision and language
temporal grounding
0.912025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Natural language and speech › Language models and text generation
text evaluation
0.912025
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025
Natural language and speech › Language models and text generation
controllable text generation
0.812024
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline · ACL (1) 2024
Computer vision › Vision and language › image captioning › low-shot image captioning
few-shot image captioning
0.712023
Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023
Computer vision › Vision and language
image captioning
0.712023
Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023
Natural language and speech › Language models and text generation › controllable text generation
style-controlled generation
0.712023
Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text · ACL (1) 2023
Computer vision › Vision and language › image captioning › controllable image captioning
stylized image captioning
0.712023
Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023
Computer vision › Vision and language › vision-language generation
visual storytelling
0.712023
Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text · ACL (1) 2023
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
0.312025
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025
Computational science and engineering
materials science
0.312025
Wafer-scale fabrication of monolayer MoS2 films for two-dimensional electronic devices via multitube atmospheric-pressure chemical vapor deposition · Sci. China Inf. Sci. 2025
Computer vision › Vision and language › video captioning
dense video captioning
0.212024
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline · ACL (1) 2024
Natural language and speech › Language models and text generation › language modeling
conditional language model
0.212023
Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

multi-agent framework · 2.0instruction tuning · 2.0chemical vapor deposition · 1.7verifiable reward · 0.9supervised fine-tuning · 0.9summary-based evaluation · 0.9reinforcement learning · 0.9incremental evaluation · 0.9aggregation-based evaluation · 0.9storyline generation · 0.8multimodal large language model · 0.8
YearPublicationVenuePosition
2026 HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External Knowledge
abstract
We present HowToNarrate, the first generaldomain benchmark for Synchronized Video Narration.The benchmark contains 3.2K videos across seven domains, segmented into 37.5K clips with aligned narrations and associated external knowledge.Effective narration requires models to understand visual scenes, incorporate relevant knowledge, and produce coherent, length-appropriate descriptions.We systematically benchmark current Multimodal LLMs (MLLMs) on these abilities.Our analysis shows that existing MLLMs overemphasize knowledge retrieval while largely neglecting prior context (receiving less than 10% attention).Moreover, they often conflate narration context with external knowledge, leading to redundancy and incoherence.To mitigate these issues, we propose VideoNarrationAgent, a multi-agent framework that combines context compression, knowledge retrieval, and narration generation.Experiments demonstrate that our method significantly improves MLLM performance.Furthermore, instruction tuning on HowToNarrate enhances both context-awareness and length control, boosting Qwen2.5-VL'sscore from 25 to 84.Our dataset and codes are released at https://github. com/wangxueyan666/HowToNarrate.
Dingyi Yang, Qin Jin
ACL (1)2
2025 What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
abstract
In this work, we conduct systematic research in a challenging area: the automatic evaluation of book-length stories (>100K tokens).Our study focuses on two key questions: (1) understanding which evaluation aspects matter most to readers, and (2) exploring effective methods for evaluating lengthy stories.We introduce the first large-scale benchmark, LongStoryEval, comprising 600 newly published books with an average length of 121K tokens (maximum 397K).Each book includes its average rating and multiple reader reviews, presented as critiques organized by evaluation aspects.By analyzing all user-mentioned aspects, we propose an evaluation criteria structure and conduct experiments to identify the most significant aspects among the 8 top-level criteria.For evaluation methods, we compare the effectiveness of three types: aggregation-based, incrementalupdated, and summary-based evaluations.Our findings reveal that aggregation-and summarybased evaluations perform better, with the former excelling in detail assessment and the latter offering greater efficiency.Building on these insights, we further propose NovelCritique, an 8B model that leverages the efficient summarybased framework to review and score stories across specified aspects.NovelCritique outperforms commercial models like GPT-4o in aligning with human evaluations.Our datasets and codes are available at https://github. com/DingyiYang/LongStoryEval.Plot Development (Progression of events): Quality of pace, twists, conflicts, and resolutions.Structure (Organization of events): Coherence and logic of elements like rising action and climax.Ending: Satisfaction and potential of the ending.
Dingyi Yang, Qin Jin
ACL (1)1
2025 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
abstract
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https://xuboshen.github.io/Time-R1/.
Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin
NeurIPS10
2025 Wafer-scale fabrication of monolayer MoS2 films for two-dimensional electronic devices via multitube atmospheric-pressure chemical vapor deposition
Zixuan Cheng, Dingyi Yang, Yizhang Wu, Jiateng Zhang, Mingwen Zhang, Xuetao Gan, Genquan Han
Sci. China Inf. Sci.2
2024 Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
abstract
Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to share a story and attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video captioning and video story generation have made some progress. However, in practical applications, we typically require synchronized narrations for ongoing visual scenes. In this work, we introduce a new task of Synchronized Video Storytelling, which aims to generate synchronous and informative narrations for videos. These narrations, associated with each video clip, should relate to the visual content, integrate relevant knowledge, and have an appropriate word count corresponding to the clip’s duration. Specifically, a structured storyline is beneficial to guide the generation process, ensuring coherence and integrity. To support the exploration of this task, we introduce a new benchmark dataset E-SyncVidStory with rich annotations. Since existing Multimodal LLMs are not effective in addressing this task in one-shot or few-shot settings, we propose a framework named VideoNarrator that can generate a storyline for input videos and simultaneously generate narrations with the guidance of the generated or predefined storyline. We further introduce a set of evaluation metrics to thoroughly assess the generation. Both automatic and human evaluations validate the effectiveness of our approach. Our dataset, codes, and evaluations will be released.
Dingyi Yang, Chunru Zhan, Tiezheng Ge, Bo Zheng 0007, Qin Jin
ACL (1)1
2023 Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text
abstract
Most research on stylized image captioning aims to generate style-specific captions using unpaired text, and has achieved impressive performance for simple styles like positive and negative.However, unlike previous singlesentence captions whose style is mostly embodied in distinctive words or phrases, realworld styles are likely to be implied at the syntactic and discourse levels.In this work, we introduce a new task of Stylized Visual Storytelling (SVST), which aims to describe a photo stream with stylized stories that are more expressive and attractive.We propose a multitasking memory-augmented framework called StyleVSG, which is jointly trained on factual visual storytelling data and unpaired style corpus, achieving a trade-off between style accuracy and visual relevance.Particularly for unpaired stylized text, StyleVSG learns to reconstruct the stylistic story from roughly parallel visual inputs mined with the CLIP 1 model, avoiding problems caused by random mapping in previous methods.Furthermore, a memory module is designed to preserve the consistency and coherence of generated stories.Experiments show that our method can generate attractive and coherent stories with different styles, such as fairy tale, romance, and humor.The overall performance of our proposed StyleVSG surpasses state-of-the-art methods on both automatic and human evaluation metrics 2 .
Dingyi Yang, Qin Jin
ACL (1)1
2023 Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences
abstract
Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized captions for visual content, so most existing works focus on unsupervised methods that do not rely on parallel datasets. However, these approaches still require training with sufficient examples that have style labels, and the generated captions are limited to predefined styles. To address these limitations, we explore the problem of Few-Shot Stylized Visual Captioning, which aims to generate captions in any desired style, using only a few examples as guidance during inference, without requiring further training. We propose a framework called FS-StyleCap for this task, which utilizes a conditional encoder-decoder language model and a visual projection module. Our two-step training scheme proceeds as follows: first, we train a style extractor to generate style representations on an unlabeled text-only corpus. Then, we freeze the extractor and enable our decoder to generate stylized descriptions based on the extracted style vector and projected visual content vectors. During inference, our model can generate desired stylized captions by deriving the style representation from user-supplied examples. Our automatic evaluation results for few-shot sentimental visual captioning outperform state-of-the-art approaches and are comparable to models that are fully trained on labeled style corpora. Human evaluations further confirm our model's ability to handle multiple styles.
Dingyi Yang, Hongyu Chen 0005, Xinglin Hou, Tiezheng Ge, Yuning Jiang 0001, Qin Jin
ACM Multimedia1
2020 Application of unsupervised TSK fuzzy algorithm in large-scale online culture courses
Dingyi Yang
Pers. Ubiquitous Comput.3