Zihao Yue

dblp:339/2864 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-3470-5442ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
abstract
Chart editing reduces manual effort in visualization design. Typical benchmarks assume access to complete chart code, which is unrealistic for real-world applications. In this paper, we present ChartEditVista, a comprehensive benchmark consisting of 7,964 samples spanning 31 chart categories. It encompasses diverse editing instruction types and covers nearly all editable chart elements. The inputs in ChartEditVista include only the original chart image and natural language editing instructions, without the original chart codes. ChartEditVista is generated through a fully automated pipeline that produces, edits, and verifies charts, ensuring high-quality data. Besides, we introduce two novel fine-grained, rule-based evaluation metrics: the layout metric, which evaluates the position, size; and color of graphical components, and the text metric, which jointly assesses textual content and font styling. Building on top of ChartEditVista, we present ChartEditor, a model trained using a reinforcement learning framework that incorporates a novel rendering reward to simultaneously enforce code executability and visual fidelity. Through extensive experiments and human evaluations, we demonstrate that ChartEditVista provides a robust evaluation, while ChartEditor consistently outperforms models with similar-scale and larger-scale on chart editing tasks.
Liangyu Chen 0008, Yichen Xu 0003, Jianzhe Ma, Yuqi Liu 0003, Donglu Yang, Zihao Yue, Wenxuan Wang 0001, Qin Jin
AAAI7
2026 Exploring Attention Attractors in Large Language Models
abstract
This paper explores attention attractorstokens that draw significantly high attentionin large language models.We analyze them from three perspectives: (1) Functionality: We demonstrate their role in aggregating information from preceding contexts to facilitate future predictions.(2) Distribution: Through layer-wise and token-wise analysis, we reveal that attention attractors are widely distributed across layers but predominantly originate from low-semantic words like "_the".(3) Mechanism: We demonstrate the correlation between attention weights allocated to tokens with their specific activation dimension values.We hope these findings provide new insights into the attention mechanisms of large language models and inspire further exploration.Code will be released at https://github.com/luyouqi233/ AttentionAttractor.
Zihao Yue, Wenxuan Wang 0001, Qin Jin
ACL (1)2
2026 POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
abstract
Yichen Xu, Liangyu Chen, Liang Zhang, Zihao Yue, Jianzhe Ma, Wenxuan Wang, Qin Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Xu 0003, Liangyu Chen 0008, Zihao Yue, Jianzhe Ma, Wenxuan Wang 0001, Qin Jin
ACL (1)4
2025 Movie101v2: Improved Movie Narration Benchmark
abstract
Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.Unlike standard video captioning, it involves not only describing key visual details but also inferring plots that unfold across multiple movie shots, presenting distinct and complex challenges.To advance this field, we introduce Movie101v2, a largescale, bilingual dataset with enhanced data quality specifically designed for movie narration.Revisiting the task, we propose breaking down the ultimate goal of automatic movie narration into three progressive stages, offering a clear roadmap with corresponding evaluation metrics.Based on our new benchmark, we baseline a range of large vision-language models and conduct an in-depth analysis of the challenges in movie narration generation.Our findings highlight that achieving applicable movie narration generation is a fascinating goal that requires significant research.
Zihao Yue, Yepeng Zhang, Qin Jin
ACL (1)1
2025 VideoOrion: Tokenizing Object Dynamics in Videos
abstract
We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating spatial-temporal object features. Our method addresses the persistent challenge in Video-LLMs of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs. Compared to prior methods which resort to downsampling the original video or aggregating visual tokens using resamplers, leading to information loss and entangled semantics, VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations but also enables explicit object modeling of video content with minimal computational cost. Moreover, the introduced object tokens naturally allow VideoOrion to accomplish video-based referring tasks. Experimental results show that VideoOrion can learn to make good use of the object tokens, and achieves competitive results on both general video question answering and video-based referring benchmarks.
Yicheng Feng, Yijiang Li, Wanpeng Zhang 0002, Sipeng Zheng, Hao Luo 0011, Zihao Yue, Zongqing Lu 0002
ICCV6
2025 Unified Multimodal Understanding via Byte-Pair Visual Encoding
abstract
Multimodal large language models (MLLMs) have made significant progress in vision-language understanding, yet effectively aligning different modalities remains a fundamental challenge. We present a framework that unifies multimodal understanding by applying byte-pair encoding to visual tokens. Unlike conventional approaches that rely on modality-specific encoders, our method directly incorporates structural information into visual tokens, mirroring successful tokenization strategies in text-only language models. We introduce a priority-guided encoding scheme that considers both frequency and spatial consistency, coupled with a multi-stage training procedure based on curriculum-driven data composition. These enhancements enable the transformer model to better capture cross-modal relationships and reason with visual information. Comprehensive experiments demonstrate improved performance across diverse vision-language tasks. By bridging the gap between visual and textual representations, our approach contributes to the advancement of more capable and efficient multimodal foundation models.
Wanpeng Zhang 0002, Yicheng Feng, Hao Luo 0011, Yijiang Li, Zihao Yue, Sipeng Zheng, Zongqing Lu 0002
ICCV5
2025 ChartM3: Benchmarking Chart Editing with Multimodal Instructions
abstract
Charts are a fundamental visualization format widely used in data analysis across research and industry. While enabling users to edit charts based on high-level intentions is of great practical value, existing methods primarily rely on natural language instructions, which are often too ambiguous to support fine-grained editing. In this work, we introduce a novel paradigm for multimodal chart editing, where user intent is expressed through a combination of natural language and visual indicators that explicitly highlight the elements to be modified. To support this paradigm, we present ChartM3, a new benchmark for Multimodal chart editing with Multi-level complexity and Multi-perspective evaluation. ChartM3 contains 1,000 samples spanning four levels of editing difficulty. Each sample includes triplets in the form of (chart, code, multimodal instructions). To comprehensively evaluate chart editing models, ChartM3 provides metrics that assess both visual appearance and code correctness. Our benchmark reveals significant limitations in current multimodal large language models (MLLMs), including GPT-4o, particularly in their ability to interpret and act on visual indicators. To address this, we construct ChartM3-Train, a large-scale training set with 24,000 multimodal chart editing samples. Fine-tuning MLLMs on this dataset leads to substantial improvements, demonstrating the importance of multimodal supervision in building practical chart editing systems. Our datasets, codes, and evaluation tools are available at https://github.com/MLrollIT/ChartM3.
Donglu Yang, Zihao Yue, Liangyu Chen 0008, Yichen Xu 0003, Wenxuan Wang 0001, Qin Jin
ACM Multimedia3
2025 OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data
abstract
Recent advances in large multimodal models have significantly advanced video comprehension, yet their performance remains limited in first-person scenarios. The interactive nature of egocentric videos is critical for applications like embodied intelligence, but introduces complex visual contexts that conventional models struggle to capture. To bridge this gap, we introduce OpenMMEgo with innovations across three dimensions: data, model, and training strategy. To provide rich spatiotemporal visual knowledge, we curate a large-scale, high-quality dataset named OME10M, comprising over 8.2M egocentric video QA pairs synthesized from Ego4D series. We also establish OMEBench, a comprehensive benchmark for rigorous egocentric understanding assessment. To alleviate the frequent viewpoint shifts inherent in egocentric videos, we implement semantic-aware visual token compression. Further, a curriculum learning strategy is complemented to foster stable learning across various data complexities. OpenMMEgo consistently improves the performance of LMMs on egocentric benchmarks without sacrificing general video understanding performance. Notably, Qwen2.5-VL tuned with OpenMMEgo substantially outperforms other models of the same size in egocentric video understanding. The data, weights and training code will be put at https://github.com/BeingBeyond/OpenMMEgo.
Hao Luo 0011, Zihao Yue, Wanpeng Zhang 0002, Yicheng Feng, Sipeng Zheng, Deheng Ye, Zongqing Lu 0002
NeurIPS2
2025 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
abstract
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https://xuboshen.github.io/Time-R1/.
Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin
NeurIPS7
2024 Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
abstract
Large Vision-Language Models (LVLMs) often suffer from multimodal hallucinations, wherein they may create content that is not present in the visual inputs.In this paper, we explore a new angle of this issue: overly detailed training data hinders the model's ability to timely terminate generation, leading to continued outputs beyond visual perception limits.By investigating how the model decides to terminate generation with EOS, the special end-of-sentence token, we find that the model assesses the completeness of the entire sequence by comparing the generated text with the image.This observation suggests that the model possesses an inherent potential of making proper EOS decisions based on its visual perception to avoid overly lengthy outputs.To take advantage of such potential, we explore two methods to mitigate multimodal hallucinations: a training objective that enables the model to reduce hallucinations by learning from regular instruction data, and a data filtering strategy to prevent harmful training data from exacerbating model hallucinations.Both methods significantly improve the hallucination performance of LVLMs, without requiring any additional data or knowledge.1
Zihao Yue, Qin Jin
ACL (1)1
2023 Movie101: A New Movie Understanding Benchmark
abstract
To help the visually impaired enjoy movies, automatic movie narrating systems are expected to narrate accurate, coherent, and role-aware plots when there are no speaking lines of actors.Existing works benchmark this challenge as a normal video captioning task via some simplifications, such as removing role names and evaluating narrations with ngram-based metrics, which makes it difficult for automatic systems to meet the needs of real application scenarios.To narrow this gap, we construct a large-scale Chinese movie benchmark, named Movie101.Closer to real scenarios, the Movie Clip Narrating (MCN) task in our benchmark asks models to generate role-aware narration paragraphs for complete movie clips where no actors are speaking.External knowledge, such as role information and movie genres, is also provided for better movie understanding.Besides, we propose a new metric called Movie Narration Score (MNScore) for movie narrating evaluation, which achieves the best correlation with human evaluation.Our benchmark also supports the Temporal Narration Grounding (TNG) task to investigate clip localization given text descriptions.For both two tasks, our proposed methods well leverage external knowledge and outperform carefully designed baselines.The dataset and codes are released at https://github.com/yuezih/Movie101.
Zihao Yue, Anwen Hu, Qin Jin
ACL (1)1
2023 Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood Estimation
abstract
Image captioning aims to describe visual content in natural language. As 'a picture is worth a thousand words', there could be various correct descriptions for an image. However, with maximum likelihood estimation as the training objective, the captioning model is penalized whenever its prediction mismatches with the label. For instance, when the model predicts a word expressing richer semantics than the label, it will be penalized and optimized to prefer more concise expressions, referred to as *conciseness optimization*. In contrast, predictions that are more concise than labels lead to *richness optimization*. Such conflicting optimization directions could eventually result in the model generating general descriptions. In this work, we introduce Semipermeable MaxImum Likelihood Estimation (SMILE), which allows richness optimization while blocking conciseness optimization, thus encouraging the model to generate longer captions with more details. Extensive experiments on two mainstream image captioning datasets MSCOCO and Flickr30K demonstrate that SMILE significantly enhances the descriptiveness of generated captions. We further provide in-depth investigations to facilitate a better understanding of how SMILE works.
Zihao Yue, Anwen Hu, Qin Jin
NeurIPS1