VLDB 2026 Research / reviewers in the wild / expert
Xunjian Yin
dblp:320/5519
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-2742-2225ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEDOM: Reverse Language ModelabstractXunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu 0001, Li Lin 0014, Xinyi Wang 0003, Liangming Pan, William Yang Wang, Xiaojun Wan 0001 |
ACL (1) | 1 |
| 2026 | EAMA: Entity-Aware Multimodal Alignment Based Approach for News Image CaptioningabstractNews image captioning requires model to generate an informative caption rich in entities, with the news image and the associated news article. Current MLLMs still bear limitations in handling entity information in news image captioning tasks. Besides, generating high-quality news image captions requires a tradeoff between sufficiency and conciseness of textual input information. To explore the potential of MLLMs, we propose an Entity-Aware Multimodal Alignment (EAMA) based approach for News Image Captioning. Our approach first aligns the MLLM with two extra alignment tasks: Entity-Aware Sentence Selection task and Entity Selection task, together with News Image Captioning task. The aligned MLLM will utilize the additional entity-related information extracted by itself to supplement the textual input while generating news image captions. Our approach achieves better results than all previous models on two mainstream news image captioning datasets. Junzhe Zhang 0004, Huixuan Zhang, Xunjian Yin, Xiaojun Wan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsabstractEvaluating the performance of Grammatical Error Correction (GEC) models has become increasingly challenging, as large language model (LLM)-based GEC systems often produce corrections that diverge from provided gold references. This discrepancy undermines the reliability of traditional reference-based evaluation metrics. In this study, we propose a novel evaluation framework for GEC models, DSGram, integrating Semantic Coherence, Edit Level, and Fluency, and utilizing a dynamic weighting mechanism. Our framework employs the Analytic Hierarchy Process (AHP) in conjunction with large language models to ascertain the relative importance of various evaluation criteria. Additionally, we develop a dataset incorporating human annotations and LLM-simulated sentences to validate our algorithms and fine-tune more cost-effective models. Experimental results indicate that our proposed approach enhances the effectiveness of GEC model evaluations. Jinxiang Xie, Xunjian Yin, Xiaojun Wan 0001 |
AAAI | 3 |
| 2025 | Gödel Agent: A Self-Referential Agent Framework for Recursively Self-ImprovementabstractThe rapid advancement of large language models (LLMs) has significantly enhanced the capabilities of agents across various tasks.However, existing agentic systems, whether based on fixed pipeline algorithms or pre-defined meta-learning frameworks, cannot search the whole agent design space due to the restriction of human-designed components, and thus might miss the more optimal agent design.In this paper, we introduce Gödel Agent, a selfevolving framework inspired by the Gödel machine, enabling agents to recursively improve themselves without relying on predefined routines or fixed optimization algorithms.Gödel Agent leverages LLMs to dynamically modify its own logic and behavior, guided solely by high-level objectives through prompting.Experimental results on multiple domains demonstrate that implementation of Gödel Agent can achieve continuous self-improvement, surpassing manually crafted agents in performance, efficiency, and generalizability. Xunjian Yin, Xinyi Wang 0003, Liangming Pan, Li Lin 0014, Xiaojun Wan 0001, William Yang Wang |
ACL (1) | 1 |
| 2025 | DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsabstractWhile large language models (LLMs) demonstrate remarkable capabilities across a wide range of tasks, they remain vulnerable to generating outputs that are potentially harmful.Red teaming, which involves crafting adversarial inputs to expose vulnerabilities, is a widely adopted approach for evaluating the robustness of these models.Prior studies have indicated that LLMs are susceptible to vulnerabilities exposed through multi-turn interactions as opposed to single-turn scenarios.Nevertheless, existing methods for multi-turn attacks mainly utilize a predefined dialogue pattern, limiting their effectiveness in realistic situations.Effective attacks require adaptive dialogue strategies that respond dynamically to the initial user prompt and the evolving context of the conversation.To address these limitations, we propose DAMON, a novel multi-turn jailbreak attack method.DAMON leverages Monte Carlo Tree Search (MCTS) to systematically explore multiturn conversational spaces, efficiently identifying sub-instruction sequences that induce harmful responses.We evaluate DAMON's efficacy across five LLMs and three datasets.Our experimental results show that DAMON can effectively induce undesired behaviors. Xu Zhang 0077, Xunjian Yin, Dinghao Jing, Huixuan Zhang, Xinyu Hu 0001, Xiaojun Wan 0001 |
EMNLP | 2 |
| 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningabstractChemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large language models (LLMs) encounter difficulties handling domain-specific formulas, executing reasoning steps accurately, and integrating code ef- effectively when tackling chemical reasoning tasks. To address these challenges, we present ChemAgent, a novel framework designed to improve the performance of LLMs through a dynamic, self-updating library. This library is developed by decomposing chemical tasks into sub-tasks and compiling these sub-tasks into a structured collection that can be referenced for future queries. Then, when presented with a new problem, ChemAgent retrieves and refines pertinent information from the library, which we call memory, facilitating effective task decomposition and the generation of solutions. Our method designs three types of memory and a library-enhanced reasoning component, enabling LLMs to improve over time through experience. Experimental results on four chemical reasoning datasets from SciBench demonstrate that ChemAgent achieves performance gains of up to 46% (GPT-4), significantly outperforming existing methods. Our findings suggest substantial potential for future applications, including tasks such as drug discovery and materials science. Our code can be found at https://github.com/gersteinlab/ChemAgent. Xiangru Tang, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang 0001, Yilun Zhao 0001, Arman Cohan, Mark Gerstein |
ICLR | 5 |
| 2024 | History Matters: Temporal Knowledge Editing in Large Language ModelabstractThe imperative task of revising or updating the knowledge stored within large language models arises from two distinct sources: intrinsic errors inherent in the model which should be corrected and outdated knowledge due to external shifts in the real world which should be updated. Prevailing efforts in model editing conflate these two distinct categories of edits arising from distinct reasons and directly modify the original knowledge in models into new knowledge. However, we argue that preserving the model's original knowledge remains pertinent. Specifically, if a model's knowledge becomes outdated due to evolving worldly dynamics, it should retain recollection of the historical knowledge while integrating the newfound knowledge. In this work, we introduce the task of Temporal Knowledge Editing (TKE) and establish a benchmark AToKe (Assessment of TempOral Knowledge Editing) to evaluate current model editing methods. We find that while existing model editing methods are effective at making models remember new knowledge, the edited model catastrophically forgets historical knowledge. To address this gap, we propose a simple and general framework termed Multi-Editing with Time Objective (METO) for enhancing existing editing models, which edits both historical and new knowledge concurrently and optimizes the model's prediction for the time of each fact. Our assessments demonstrate that while AToKe is still difficult, METO maintains the effectiveness of learning new knowledge and meanwhile substantially improves the performance of edited models on utilizing historical knowledge. Xunjian Yin, Xiaojun Wan 0001 |
AAAI | 1 |
| 2024 | Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model EvaluationabstractIn recent years, substantial advancements have been made in the development of large language models, achieving remarkable performance across diverse tasks.To evaluate the knowledge ability of language models, previous studies have proposed lots of benchmarks based on question-answering pairs.We argue that it is not reliable and comprehensive to evaluate language models with a fixed question or limited paraphrases as the query, since language models are sensitive to prompt.Therefore, we introduce a novel concept named knowledge boundary to encompass both prompt-agnostic and promptsensitive knowledge within language models.Knowledge boundary avoids prompt sensitivity in language model evaluations, rendering them more dependable and robust.To explore the knowledge boundary for a given model, we propose a projected gradient descent method with semantic constraints, a new algorithm designed to identify the optimal prompt for each piece of knowledge.Experiments demonstrate a superior performance of our algorithm in computing the knowledge boundary compared to existing methods.Furthermore, we evaluate the ability of multiple language models in several domains with knowledge boundary. Xunjian Yin, Xu Zhang 0077, Jie Ruan, Xiaojun Wan 0001 |
ACL (1) | 1 |
| 2024 | Contextual Modeling for Document-level ASR Error CorrectionabstractContextual information, including the sentences in the same document and in other documents of the dataset, plays a crucial role in improving the accuracy of document-level ASR Error Correction (AEC), while most previous works ignore this. In this paper, we propose a context-aware method that utilizes a k-Nearest Neighbors (kNN) approach to enhance the AEC model by retrieving a datastore containing contextual information. We conduct experiments on two English and two Chinese datasets, and the results demonstrate that our proposed model can effectively utilize contextual information to improve document-level AEC. Furthermore, the context information from the whole dataset provides even better results. Xunjian Yin, Xiaojun Wan 0001, Wei Peng 0011, Rongjun Li, Jingyuan Yang 0008, Yanquan Zhou |
LREC/COLING | 2 |
| 2024 | Error-Robust Retrieval for Chinese Spelling CheckabstractChinese Spelling Check (CSC) aims to detect and correct error tokens in Chinese contexts, which has a wide range of applications. However, it is confronted with the challenges of insufficient annotated data and the issue that previous methods may actually not fully leverage the existing datasets. In this paper, we introduce our plug-and-play retrieval method with error-robust information for Chinese Spelling Check (RERIC), which can be directly applied to existing CSC models. The datastore for retrieval is built completely based on the training data, with elaborate designs according to the characteristics of CSC. Specifically, we employ multimodal representations that fuse phonetic, morphologic, and contextual information in the calculation of query and key during retrieval to enhance robustness against potential errors. Furthermore, in order to better judge the retrieved candidates, the n-gram surrounding the token to be checked is regarded as the value and utilized for specific reranking. The experiment results on the SIGHAN benchmarks demonstrate that our proposed method achieves substantial improvements over existing work. Xunjian Yin, Xinyu Hu 0001, Xiaojun Wan 0001 |
LREC/COLING | 1 |
| 2024 | Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityabstractThe evaluation of natural language generation (NLG) tasks is a significant and longstanding research area.With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automatic evaluation methods, which demonstrate great potential to become a new evaluation paradigm following traditional string-based and modelbased metrics.However, despite the improved performance of existing methods, they still possess some deficiencies, such as dependency on references and limited evaluation flexibility.Therefore, in this paper, we meticulously construct a large-scale NLG evaluation corpus NLG-Eval with annotations from both human and GPT-4 to alleviate the lack of relevant data in this field.Furthermore, we propose Themis, an LLM dedicated to NLG evaluation, which has been trained with our designed multi-perspective consistency verification and rating-oriented preference alignment methods.Themis can conduct flexible and interpretable evaluations without references, and it exhibits superior evaluation performance on various NLG tasks, simultaneously generalizing well to unseen tasks and surpassing other evaluation models, including GPT-4. Xinyu Hu 0001, Li Lin 0014, Mingqi Gao 0002, Xunjian Yin, Xiaojun Wan 0001 |
EMNLP | 4 |
| 2023 | ALCUNA: Large Language Models Meet New KnowledgeabstractWith the rapid development of NLP, large-scale language models (LLMs) excel in various tasks across multiple domains now.However, existing benchmarks may not adequately measure these models' capabilities, especially when faced with new knowledge.In this paper, we address the lack of benchmarks to evaluate LLMs' ability to handle new knowledge, an important and challenging aspect in the rapidly evolving world.We propose an approach called Know-Gen that generates new knowledge by altering existing entity attributes and relationships, resulting in artificial entities that are distinct from real-world entities.With KnowGen, we introduce a benchmark named ALCUNA to assess LLMs' abilities in knowledge understanding, differentiation, and association.We benchmark several LLMs, reveals that their performance in face of new knowledge is not satisfactory, particularly in reasoning between new and internal knowledge.We also explore the impact of entity similarity on the model's understanding of entity knowledge and the influence of contextual entities.We appeal to the need for caution when using LLMs in new scenarios or with new knowledge, and hope that our benchmarks can help drive the development of LLMs in face of new knowledge. Xunjian Yin, Baizhou Huang, Xiaojun Wan 0001 |
EMNLP | 1 |
| 2023 | Overview of the NLPCC 2023 Shared Task: Chinese Spelling Check
Xunjian Yin, Xiaojun Wan 0001, Linlin Yu |
NLPCC (3) | 1 |
| 2022 | How Do Seq2Seq Models Perform on End-to-End Data-to-Text Generation?abstractWith the rapid development of deep learning, Seq2Seq paradigm has become prevalent for end-to-end data-to-text generation, and the BLEU scores have been increasing in recent years.However, it is widely recognized that there is still a gap between the quality of the texts generated by models and the texts written by human.In order to better understand the ability of Seq2Seq models, evaluate their performance and analyze the results, we choose to use Multidimensional Quality Metric(MQM) to evaluate several representative Seq2Seq models on end-to-end data-to-text generation.We annotate the outputs of five models on four datasets with eight error types and find that 1) copy mechanism is helpful for the improvement in Omission and Inaccuracy Extrinsic errors but it increases other types of errors such as Addition; 2) pre-training techniques are highly effective, and pre-training strategy and model size are very significant; 3) the structure of the dataset also influences the model's performance greatly; 4) some specific types of errors are generally challenging for seq2seq models. Xunjian Yin, Xiaojun Wan 0001 |
ACL (1) | 1 |