Jiayuan Su

dblp:368/0243 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 32% Machine translation · 21% Question answering and dialogue systems · 11%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
model routing
1.012026
CP-Router: An Uncertainty-Aware Router Between LLM and LRM · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
Natural language and speech › Language models and text generation › natural language understanding › question answering
multiple-choice question answering
1.012026
CP-Router: An Uncertainty-Aware Router Between LLM and LRM · AAAI 2026
Machine learning › Reinforcement learning
multi-task reinforcement learning
1.012026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
table question answering
1.012026
SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets · AAAI 2026
Natural language and speech › Machine translation › multimodal machine translation
text image machine translation
1.012026
MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation · ACL (1) 2025
Natural language and speech › Machine translation
machine translation evaluation
0.912025
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation · ACL (1) 2025
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
0.912025
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning
0.312026
SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets · AAAI 2026

Methods — techniques the papers use, named apart from their topics

large language model · 1.9reinforcement learning · 1.0python sandbox · 1.0neuro-symbolic agent · 1.0multi-task learning · 1.0entropy-based thresholding · 1.0conformal prediction · 1.0multi-agent debate · 0.9
YearPublicationVenuePosition
2026 CP-Router: An Uncertainty-Aware Router Between LLM and LRM
abstract
Recent advances in large reasoning models (LRMs) have significantly enhanced long-chain reasoning capabilities over standard large language models (LLMs). However, LRMs often produce unnecessarily lengthy outputs even for simple queries, leading to inefficiencies or even accuracy degradation compared to LLMs. To address this, we propose CP-Router, a training-free, model-agnostic routing framework that dynamically selects between an LLM and an LRM, demonstrated with multiple-choice question answering (MCQA) prompts. The routing decision is guided by the prediction uncertainty estimates derived via Conformal Prediction (CP), which provides rigorous coverage guarantees. To improve uncertainty differentiation across inputs, we introduce Full and Binary Entropy (FBE), a novel entropy-based criterion that adaptively selects the appropriate CP threshold. Experiments across MCQA and QA benchmarks—including mathematics, logical reasoning, and Chinese chemistry—demonstrate that CP-Router efficiently reduces token usage while maintaining or even improving accuracy compared to using LRM alone. We further demonstrate the generality and robustness of CP-Router by extending it to diverse model pairings beyond the LLM–LRM setting.
Jiayuan Su, Fulin Lin, Zhaopeng Feng, Zhenyu Xiao, Xinlong Zhao, Zuozhu Liu, Hongwei Wang 0001
AAAI1
2026 SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets
abstract
Understanding and reasoning over complex spreadsheets remain fundamental challenges for large language models (LLMs), which often struggle with intricate structures and rely solely on neural computation. In this work, we propose SheetBrain, a neuro-symbolic dual-workflow agent framework for precise and interpretable reasoning over tabular data. SheetBrain consists of an understanding module that produces a comprehensive overview of the spreadsheet, including structural summaries and query-specific analyses to guide execution; an execution module that integrates a Python sandbox with preloaded table-processing libraries and an Excel helper toolkit for effective data manipulation; and a validation module that verifies the correctness of reasoning and answers, triggering re-execution if necessary. We evaluate SheetBrain on multiple public QA and manipulation benchmarks, and introduce SheetBench, a new benchmark targeting large, multi-table, and structurally complex spreadsheets. Experimental results show that SheetBrain significantly improves reasoning performance on both existing benchmarks and the more challenging scenarios presented in SheetBench.
Jiayuan Su, Mengyu Zhou, Huaxing Zeng, Mengni Jia, Haoyu Dong 0001, Xiaojun Ma 0001, Shi Han, Dongmei Zhang 0001
AAAI2
2026 MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation
abstract
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Zhijie Zhou, Wenxuan Huang, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Wenxuan Huang 0001, Jian Wu 0001, Zuozhu Liu
ACL (1)4
2026 Improving few-shot named entity recognition with distilled knowledge from large language model
Qi Li 0042, Tingyu Xie, Jiayuan Su, Jian Zhang 0083, Hongwei Wang 0001
Neurocomputing3
2025 M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation
abstract
Recent advancements in large language models (LLMs) have given rise to the LLM-as-a-judge paradigm, showcasing their potential to deliver human-like judgments. However, in the field of machine translation (MT) evaluation, current LLM-as-a-judge methods fall short of learned automatic metrics. In this paper, we propose Multidimensional Multi-Agent Debate (M-MAD), a systematic LLM-based multi-agent framework for advanced LLM-as-a-judge MT evaluation. Our findings demonstrate that M-MAD achieves significant advancements by (1) decoupling heuristic MQM criteria into distinct evaluation dimensions for fine-grained assessments; (2) employing multi-agent debates to harness the collaborative reasoning capabilities of LLMs; (3) synthesizing dimension-specific results into a final evaluation judgment to ensure robust and reliable outcomes. Comprehensive experiments show that M-MAD not only outperforms all existing LLM-as-a-judge methods but also competes with state-of-the-art reference-based automatic metrics, even when powered by a suboptimal model like GPT-4o mini. Detailed ablations and analysis highlight the superiority of our framework design, offering a fresh perspective for LLM-as-a-judge paradigm. Our code and data are publicly available at https://github.com/SU-JIAYUAN/M-MAD.
Zhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren, Yan Zhang 0004, Jian Wu 0001, Hongwei Wang 0001, Zuozhu Liu
ACL (1)2
2025 Enhancing named entity recognition with external knowledge from large language model
Qi Li 0042, Tingyu Xie, Jian Zhang 0083, Jiayuan Su, Kaixiang Yang 0001, Hongwei Wang 0001
Knowl. Based Syst.5
2025 Class Incremental Fault Diagnosis Under Limited Fault Data via Supervised Contrastive Knowledge Distillation
abstract
Class-incremental fault diagnosis requires a model to adapt to new fault classes while retaining previous knowledge. However, limited research exists for imbalanced and long-tailed data. Extracting discriminative features from few-shot fault data is challenging, and adding new fault classes often demands costly model retraining. Moreover, incremental training of existing methods risks catastrophic forgetting, and severe class imbalance can bias the model's decisions toward normal classes. To tackle these issues, we introduce a supervised contrastive knowledge distillation for class incremental fault diagnosis (SCLIFD) framework proposing supervised contrastive knowledge distillation for improved representation learning capability and less forgetting, a novel prioritized exemplar selection method for sample replay to alleviate catastrophic forgetting, and the random forest classifier to address the class imbalance. Extensive experimentation on simulated and real-world industrial datasets across various imbalance ratios demonstrates the superiority of SCLIFD over existing approaches.
Hanrong Zhang, Yifei Yao, Zixuan Wang 0028, Jiayuan Su, Mengxuan Li 0003, Peng Peng 0006, Hongwei Wang 0001
IEEE Trans. Ind. Informatics4
2023 EGDE: A Framework for Bridging the Gap in Medical Zero-shot Relation Triplet Extraction
abstract
Medical zero-shot relation triplet extraction, referred to as Med-ZeroRTE, requires the model to extract triplets comprising entities and relations from medical sentences. Importantly, the sentences include relations that were unseen during the model’s training phase. While Med-ZeroRTE had not been formally explored before this work, the limited availability of medical datasets, influenced by privacy concerns and annotation costs, emphasizes the necessity of exploring Med-ZeroRTE. This exploration faces two main challenges: Firstly, there is a gap of work specifically focused on triplet extraction from medical text in a zero-shot setting. Secondly, while a few approaches tackle the general zero-shot problems by employing generative models to produce synthetic data for unseen classes, the quality of some synthetic data remains suboptimal. Therefore, we propose a novel Enhanced Generator - Discriminator - Extractor framework (EGDE), which consists of three core modules, a prompt-tuned generator for generating synthetic samples given unseen relations, a fine-tuned discriminator for filtering qualified synthetic samples, a prompt-tuned extractor for extracting predicted medical triplets, to resolve Med-ZeroRTE and mitigate issues related to poor synthetic samples. The proposed framework is shown to be effective and superior compared to several robust baselines in experiments conducted on two distinct dataset settings.
Jiayuan Su, Jian Zhang 0083, Peng Peng 0006, Hongwei Wang 0001
BIBM1