Linran Xu

dblp:351/4305 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal reasoning
0.912025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025
Program synthesis and code generation › code generation with language models
chart-to-code generation
0.912025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.312025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025

Methods — techniques the papers use, named apart from their topics

benchmark evaluation · 1.7
YearPublicationVenuePosition
2025 ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
abstract
We introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs, requiring LMMs to generate the corresponding code for chart rendering. ChartMimic includes $4,800$ human-curated (figure, instruction, code) triplets, which represent the authentic chart use cases found in scientific papers across various domains (e.g., Physics, Computer Science, Economics, etc). These charts span $18$ regular types and $4$ advanced types, diversifying into $201$ subcategories. Furthermore, we propose multi-level evaluation metrics to provide an automatic and thorough assessment of the output code and the rendered charts. Unlike existing code generation benchmarks, ChartMimic places emphasis on evaluating LMMs' capacity to harmonize a blend of cognitive capabilities, encompassing visual understanding, code generation, and cross-modal reasoning. The evaluation of $3$ proprietary models and $14$ open-weight models highlights the substantial challenges posed by ChartMimic. Even the advanced GPT-4o, InternVL2-Llama3-76B only achieved an average score across Direct Mimic and Customized Mimic tasks of $82.2$ and $61.6$, respectively, indicating significant room for improvement. We anticipate that ChartMimic will inspire the development of LMMs, advancing the pursuit of artificial general intelligence.
Cheng Yang 0002, Chufan Shi, Bo Shui, Junjie Wang 0011, Mohan Jing, Linran Xu, Siheng Li, Gongye Liu, Xiaomei Nie, Deng Cai 0002, Yujiu Yang 0001
ICLR7