Tan Yue

dblp:31/2384 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-5798-6304ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 41% Language models and text generation · 35% Question answering and dialogue systems · 10%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model reasoning
adaptive reasoning
1.012026
MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking · AAAI 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking · AAAI 2026
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking · AAAI 2026
Computer vision › Vision and language
visual question answering
1.012026
We May Not Need Much Visual Encoding of Web Data for Question Answering · WWW 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
QAEval: Mixture of Evaluators for Question-Answering Task Evaluation · ACL (1) 2025
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.912025
AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis · ACM Multimedia 2025
Natural language and speech › Question answering and dialogue systems
question answering evaluation
0.912025
QAEval: Mixture of Evaluators for Question-Answering Task Evaluation · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
Revisit Finetuning strategy for Few-Shot Learning to Transfer the Emdeddings · ICLR 2023
Natural language and speech › Question answering and dialogue systems
multimodal question answering
0.312026
We May Not Need Much Visual Encoding of Web Data for Question Answering · WWW 2026

Methods — techniques the papers use, named apart from their topics

mixture of experts · 1.7visual encoding · 1.0reinforcement learning · 1.0information bottleneck · 1.0chain-of-thought · 1.0multimodal large language model · 0.9human-aligned scoring · 0.9hilbert-schmidt independence criterion · 0.9dynamic load balancing · 0.9embedding transfer · 0.7
YearPublicationVenuePosition
2026 MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking
abstract
Multimodal Large Language Models (MLLMs) have shown advanced performance in vision-language tasks. However, existing multimodal reasoning models often suffer from excessive reasoning steps, leading to high computational costs and inefficiency. In this paper, we propose the Multimodal Adaptive Reasoning Model (MARS), which enables adaptive adjustment of the reasoning strategy based on question difficulty. Specifically, MARS adopts a three-stage training framework based on our constructed training dataset (MART): 1) CoT Masking Learning to enhance reasoning logicality by predicting masked reasoning steps. 2) Adaptive Reasoning Instruction Learning to train the model to skip or keep reasoning steps according to difficulty levels. 3) CoT Lightweight Reinforcement Learning with the Information Bottleneck Principle based GRPO algorithm to reduce CoT length while maintaining performance and generalizability. Results on both in-domain and out-of-domain datasets show that MARS significantly reduces the CoT length (90.2% decrease) while improving accuracy (0.54%), outperforming existing SOTA open-source and proprietary MLLMs.
Tan Yue, Dongyan Zhao 0001
AAAI1
2026 We May Not Need Much Visual Encoding of Web Data for Question Answering
Tan Yue, Dongyan Zhao 0001
WWW1
2025 QAEval: Mixture of Evaluators for Question-Answering Task Evaluation
abstract
Question answering (QA) tasks serve as a key benchmark for evaluating generation systems.Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with openended and unstructured responses.LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to instructions, robustness issues, and high computational costs.To overcome these challenges, we introduce QAEval, a hybrid framework combining rule-based reliability with LLM-based adaptability.QAEval utilizes two high-quality datasets: QAExtract for short-answer extraction and QAScore for scoring model training.By integrating a Mixture of Evaluators model with Dynamic Load Balancing Optimization, QAEval enables accurate, cost-effective QA evaluation.Experimental results show it outperforms models like GPT-4o and Claude-3, achieving 92.3% accuracy with only 0.6B parameters.
Tan Yue, Rui Mao 0010, Xuzhao Shi, Shuo Zhan, Zuhao Yang, Dongyan Zhao 0001
ACL (1)1
2025 F2TEval: Human-Aligned Multi-Dimensional Evaluation for Figure-to-Text Task
abstract
Figure -to-Text (F2T) tasks aim to convert structured figure information into natural language text, serving as a bridge between visual perception and language understanding.However, existing evaluation methods remain limited: 1) Reference-based methods can only capture shallow semantic similarities and rely on costly labeled reference text; 2) Reference-free methods depend on multimodal large language models, which suffer from low efficiency and instruction sensitivity; 3) Existing methods provide only sample-level evaluations, lacking interpretability and alignment with expert-level multi-dimensional evaluation criteria.Accordingly, we propose F2TEval, a five-dimensional reference-free evaluation method aligned with expert criteria, covering faithfulness, completeness, conciseness, logicality, and analysis, to support fine-grained evaluation.We design a lightweight mixture-of-experts model that incorporates independent scoring heads and applies the Hilbert-Schmidt Independence Criterion to optimize the disentanglement of scoring representations across dimensions.Furthermore, we construct F2TBenchmark, a humanannotated benchmark dataset covering 21 chart types and 35 application domains, to support research on F2T evaluation.Experimental results demonstrate our model's superior performance and efficiency, outperforming Gemini-2.0and Claude-3.5 with only 0.9B parameters.
Tan Yue, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001
EMNLP1
2025 AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis
abstract
Scientific Figure Analysis (SFA) aims to derive analytical insights from figures while incorporating background instructions. Unlike conventional tasks such as figure captioning or description generation, which focus on extracting surface-level information from the sole visual modality, SFA requires an intelligent system to summarize key patterns, infer implications, and contextualize scientific findings from visual and textual inputs. It demands not only visual recognition but also the integration of scientific knowledge, multimodal understanding, and contextual reasoning. In this work, we introduce an SFA dataset, AnaFig, comprising 2,000 high-quality samples across 56 domains. All samples are evaluated by using human-aligned five-dimensional scoring criteria, resulting 10,000 human-annotated score labels. The AnaFig dataset facilitates the assessment of three critical capabilities of multimodal large language models (MLLMs): adherence to complex instructions, multimodal perception, and analytical summarization. By building a new benchmark with widely used MLLMs, this study contributes to scientific knowledge discovery and reasoning, fostering the alignment of MLLMs and human experts in scientific analysis.
Tan Yue, Xuzhao Shi, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001
ACM Multimedia1
2024 SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset
abstract
Sarcasm poses a challenge in linguistic analysis due to its implicit nature, involving an intended meaning that contradicts the literal expression. The advent of social networks has propelled the utilization of multimodal data to enhance sarcasm detection performance. In prior multimodal sarcasm detection datasets, a single label is assigned to a multimodal instance. Subsequent experiments often highlight the superiority of multimodal models by demonstrating their improvements compared to unimodal models based on these unified labels across multiple modalities. However, our investigation revealed that numerous instances of sarcasm cannot be identified using a single modality. Humans employ the conflict between a statement and factual information as a cue to detect sarcasm, and these cues can stem from different modalities. Then, a unified label for a multimodal instance may be not suitable for the associated text or image. In this work, we introduce SarcNet, a multilingual and multimodal sarcasm detection dataset in English and Chinese, consisting of 3,335 image-text pair samples. We provide annotations for sarcasm in visual, textual, and multimodal data, respectively, resulting in over 10,000 labeled instances. The separated annotation schema for unimodal and multimodal data facilitates a more accurate and reasonable assessment of unimodal and multimodal models.
Tan Yue, Xuzhao Shi, Rui Mao 0010, Zonghai Hu, Erik Cambria
LREC/COLING1
2024 Multi-modal hierarchical fusion network for fine-grained paper classification
Tan Yue, Jiedong Qin, Zonghai Hu
Multim. Tools Appl.1
2023 Revisit Finetuning strategy for Few-Shot Learning to Transfer the Emdeddings
Tan Yue, Xiang Ye, Bohan Li 0013, Yong Li 0025
ICLR2
2023 CLDM: convolutional layer dropout module
Jiafeng Zhao, Xiang Ye, Tan Yue, Yong Li 0025
Mach. Vis. Appl.3