VLDB 2026 Research / reviewers in the wild / expert
Dingyi Yang
dblp:266/2264
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 49% Language models and text generation · 46% Knowledge representation and reasoning · 2% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Integrated circuit design · 100% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.9 | 2 | 2026 | HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External Knowledge · ACL (1) 2026 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
post-training |
0.9 | 1 | 2025 | Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training |
0.9 | 1 | 2025 | Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025 |
Natural language and speech › Language models and text generation › text generation evaluation
story evaluation |
0.9 | 1 | 2025 | What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025 |
Computer vision › Vision and language
temporal grounding |
0.9 | 1 | 2025 | Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025 |
Natural language and speech › Language models and text generation
text evaluation |
0.9 | 1 | 2025 | What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
controllable text generation |
0.8 | 1 | 2024 | Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline · ACL (1) 2024 |
Computer vision › Vision and language › image captioning › low-shot image captioning
few-shot image captioning |
0.7 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Computer vision › Vision and language
image captioning |
0.7 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Natural language and speech › Language models and text generation › controllable text generation
style-controlled generation |
0.7 | 1 | 2023 | Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text · ACL (1) 2023 |
Computer vision › Vision and language › image captioning › controllable image captioning
stylized image captioning |
0.7 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Computer vision › Vision and language › vision-language generation
visual storytelling |
0.7 | 1 | 2023 | Attractive Storyteller: Stylized Visual Storytelling with Unpaired Text · ACL (1) 2023 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.3 | 1 | 2025 | Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding · NeurIPS 2025 |
Computational science and engineering
materials science |
0.3 | 1 | 2025 | Wafer-scale fabrication of monolayer MoS2 films for two-dimensional electronic devices via multitube atmospheric-pressure chemical vapor deposition · Sci. China Inf. Sci. 2025 |
Computer vision › Vision and language › video captioning
dense video captioning |
0.2 | 1 | 2024 | Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline · ACL (1) 2024 |
Natural language and speech › Language models and text generation › language modeling
conditional language model |
0.2 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
multi-agent framework · 2.0instruction tuning · 2.0chemical vapor deposition · 1.7verifiable reward · 0.9supervised fine-tuning · 0.9summary-based evaluation · 0.9reinforcement learning · 0.9incremental evaluation · 0.9aggregation-based evaluation · 0.9storyline generation · 0.8multimodal large language model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External KnowledgeabstractWe present HowToNarrate, the first generaldomain benchmark for Synchronized Video Narration.The benchmark contains 3.2K videos across seven domains, segmented into 37.5K clips with aligned narrations and associated external knowledge.Effective narration requires models to understand visual scenes, incorporate relevant knowledge, and produce coherent, length-appropriate descriptions.We systematically benchmark current Multimodal LLMs (MLLMs) on these abilities.Our analysis shows that existing MLLMs overemphasize knowledge retrieval while largely neglecting prior context (receiving less than 10% attention).Moreover, they often conflate narration context with external knowledge, leading to redundancy and incoherence.To mitigate these issues, we propose VideoNarrationAgent, a multi-agent framework that combines context compression, knowledge retrieval, and narration generation.Experiments demonstrate that our method significantly improves MLLM performance.Furthermore, instruction tuning on HowToNarrate enhances both context-awareness and length control, boosting Qwen2.5-VL'sscore from 25 to 84.Our dataset and codes are released at https://github. com/wangxueyan666/HowToNarrate. Dingyi Yang, Qin Jin |
ACL (1) | 2 |
| 2025 | What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story EvaluationabstractIn this work, we conduct systematic research in a challenging area: the automatic evaluation of book-length stories (>100K tokens).Our study focuses on two key questions: (1) understanding which evaluation aspects matter most to readers, and (2) exploring effective methods for evaluating lengthy stories.We introduce the first large-scale benchmark, LongStoryEval, comprising 600 newly published books with an average length of 121K tokens (maximum 397K).Each book includes its average rating and multiple reader reviews, presented as critiques organized by evaluation aspects.By analyzing all user-mentioned aspects, we propose an evaluation criteria structure and conduct experiments to identify the most significant aspects among the 8 top-level criteria.For evaluation methods, we compare the effectiveness of three types: aggregation-based, incrementalupdated, and summary-based evaluations.Our findings reveal that aggregation-and summarybased evaluations perform better, with the former excelling in detail assessment and the latter offering greater efficiency.Building on these insights, we further propose NovelCritique, an 8B model that leverages the efficient summarybased framework to review and score stories across specified aspects.NovelCritique outperforms commercial models like GPT-4o in aligning with human evaluations.Our datasets and codes are available at https://github. com/DingyiYang/LongStoryEval.Plot Development (Progression of events): Quality of pace, twists, conflicts, and resolutions.Structure (Organization of events): Coherence and logic of elements like rising action and climax.Ending: Satisfaction and potential of the ending. Dingyi Yang, Qin Jin |
ACL (1) | 1 |
| 2025 | Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingabstractTemporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL).
Specifically, our contributions span three key directions:
(1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task.
(2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization.
(3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks.
Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities.
Project Page: https://xuboshen.github.io/Time-R1/. Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin |
NeurIPS | 10 |
| 2025 | Wafer-scale fabrication of monolayer MoS2 films for two-dimensional electronic devices via multitube atmospheric-pressure chemical vapor deposition
Zixuan Cheng, Dingyi Yang, Yizhang Wu, Jiateng Zhang, Mingwen Zhang, Xuetao Gan, Genquan Han |
Sci. China Inf. Sci. | 2 |
| 2024 | Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineabstractVideo storytelling is engaging multimedia content that utilizes video and its accompanying narration to share a story and attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video captioning and video story generation have made some progress. However, in practical applications, we typically require synchronized narrations for ongoing visual scenes. In this work, we introduce a new task of Synchronized Video Storytelling, which aims to generate synchronous and informative narrations for videos. These narrations, associated with each video clip, should relate to the visual content, integrate relevant knowledge, and have an appropriate word count corresponding to the clip’s duration. Specifically, a structured storyline is beneficial to guide the generation process, ensuring coherence and integrity. To support the exploration of this task, we introduce a new benchmark dataset E-SyncVidStory with rich annotations. Since existing Multimodal LLMs are not effective in addressing this task in one-shot or few-shot settings, we propose a framework named VideoNarrator that can generate a storyline for input videos and simultaneously generate narrations with the guidance of the generated or predefined storyline. We further introduce a set of evaluation metrics to thoroughly assess the generation. Both automatic and human evaluations validate the effectiveness of our approach. Our dataset, codes, and evaluations will be released. Dingyi Yang, Chunru Zhan, Tiezheng Ge, Bo Zheng 0007, Qin Jin |
ACL (1) | 1 |
| 2023 | Attractive Storyteller: Stylized Visual Storytelling with Unpaired TextabstractMost research on stylized image captioning aims to generate style-specific captions using unpaired text, and has achieved impressive performance for simple styles like positive and negative.However, unlike previous singlesentence captions whose style is mostly embodied in distinctive words or phrases, realworld styles are likely to be implied at the syntactic and discourse levels.In this work, we introduce a new task of Stylized Visual Storytelling (SVST), which aims to describe a photo stream with stylized stories that are more expressive and attractive.We propose a multitasking memory-augmented framework called StyleVSG, which is jointly trained on factual visual storytelling data and unpaired style corpus, achieving a trade-off between style accuracy and visual relevance.Particularly for unpaired stylized text, StyleVSG learns to reconstruct the stylistic story from roughly parallel visual inputs mined with the CLIP 1 model, avoiding problems caused by random mapping in previous methods.Furthermore, a memory module is designed to preserve the consistency and coherence of generated stories.Experiments show that our method can generate attractive and coherent stories with different styles, such as fairy tale, romance, and humor.The overall performance of our proposed StyleVSG surpasses state-of-the-art methods on both automatic and human evaluation metrics 2 . Dingyi Yang, Qin Jin |
ACL (1) | 1 |
| 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesabstractStylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized captions for visual content, so most existing works focus on unsupervised methods that do not rely on parallel datasets. However, these approaches still require training with sufficient examples that have style labels, and the generated captions are limited to predefined styles. To address these limitations, we explore the problem of Few-Shot Stylized Visual Captioning, which aims to generate captions in any desired style, using only a few examples as guidance during inference, without requiring further training. We propose a framework called FS-StyleCap for this task, which utilizes a conditional encoder-decoder language model and a visual projection module. Our two-step training scheme proceeds as follows: first, we train a style extractor to generate style representations on an unlabeled text-only corpus. Then, we freeze the extractor and enable our decoder to generate stylized descriptions based on the extracted style vector and projected visual content vectors. During inference, our model can generate desired stylized captions by deriving the style representation from user-supplied examples. Our automatic evaluation results for few-shot sentimental visual captioning outperform state-of-the-art approaches and are comparable to models that are fully trained on labeled style corpora. Human evaluations further confirm our model's ability to handle multiple styles. Dingyi Yang, Hongyu Chen 0005, Xinglin Hou, Tiezheng Ge, Yuning Jiang 0001, Qin Jin |
ACM Multimedia | 1 |
| 2020 | Application of unsupervised TSK fuzzy algorithm in large-scale online culture courses
Dingyi Yang |
Pers. Ubiquitous Comput. | 3 |