Xianjie Wu

dblp:263/6428 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0003-7548-9824ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
abstract
Existing reinforcement learning methods for Chain-of-Thought reasoning suffer from two critical limitations. First, they operate as monolithic black boxes that provide undifferentiated reward signals, obscuring individual step contributions and hindering error diagnosis. Second, sequential decoding has O(n) time complexity. This makes real-time deployment impractical for complex reasoning tasks. We present DeCoRL (Decoupled Reasoning Chains via Coordinated Reinforcement Learning), a novel framework that transforms reasoning from sequential processing into collaborative modular orchestration. DeCoRL trains lightweight specialized models to generate reasoning sub-steps concurrently, eliminating sequential bottlenecks through parallel processing. To enable precise error attribution, the framework designs modular reward functions that score each sub-step independently. Cascaded DRPO optimization then coordinates these rewards while preserving inter-step dependencies. Comprehensive evaluation demonstrates state-of-the-art results across RM-Bench, RMB, and RewardBench, outperforming existing methods including large-scale models. DeCoRL delivers 3.8 times faster inference while maintaining superior solution quality and offers a 22.7% improvement in interpretability through explicit reward attribution. These advancements, combined with a 72.4% reduction in energy consumption and a 68% increase in throughput, make real-time deployment of complex reasoning systems a reality.
Ziyuan Gao, Di Liang, Xianjie Wu, Philippe Morel, Minlong Peng
AAAI3
2026 Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
abstract
Zekai Lin, Chao Xue, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Lei Jiang, Yu Lu, Bob Simons, Shuang Liang, Minlong Peng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zekai Lin, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Bob Simons, Minlong Peng
ACL (1)6
2026 Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
abstract
Chao Xue, Yao Wang, Mengqiao Liu, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Chenyao Lu, Lei Jiang, Yu Lu, Haibo Shi, Shuang Liang, Minlong Peng, Flora D. Salim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mengqiao Liu, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Chenyao Lu, Haibo Shi, Minlong Peng, Flora D. Salim
ACL (1)7
2026 MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA
abstract
Tables serve as a core format for representing structured data on the web, as their two-dimensional layouts effectively encode complex inter-entity relationships. However, real-world web tables often feature heterogeneous structures and rich semantics. Accurately interpreting such tables requires not only spatial layout perception but also multi-step reasoning across rows and columns, posing substantial challenges to web intelligence systems. Multimodal large language models (MLLMs) show promise in table question answering (TableQA) by leveraging visual layouts. However, their performance on complex web tables remains uneven, as existing benchmarks often blur the impact of individual difficulty factors, hindering precise capability analysis. To advance TableQA beyond superficial task difficulty and toward interpretable capability modeling, we introduce MMTableBench, a multi-level benchmark that systematically evaluates MLLMs along two fine-grained dimensions: layout complexity and reasoning complexity. By organizing table-question pairs along these axes, MMTableBench facilitates a detailed evaluation of model performance under varying structural and reasoning challenges, while revealing the respective strengths and limitations of multimodal inputs. Our comprehensive analysis shows that state-of-the-art MLLMs continue to exhibit notable limitations when confronted with complex layouts and deep reasoning tasks, underscoring persistent gaps despite the structural advantages offered by visual inputs. MMTableBench thus provides not only a rigorous evaluation framework but also a diagnostic tool for analyzing and interpreting model behaviors, enabling more transparent and explainable progress in multimodal TableQA development.
Xianjie Wu, Xiaohang Xu 0002, Tingyu Jiang, Jian Yang 0030, Di Liang, Xianfu Cheng, Zhenhe Wu, Linzheng Chai, Wei Zhang 0384, Ge Zhang 0009, Bob Simons, Tongliang Li, Zhoujun Li 0001
WWW1
2025 TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
abstract
Recent advancements in Large Language Models (LLMs) have markedly enhanced the interpretation and processing of tabular data, introducing previously unimaginable capabilities. Despite these achievements, LLMs still encounter significant challenges when applied in industrial scenarios, particularly due to the increased complexity of reasoning required with real-world tabular data, underscoring a notable disparity between academic benchmarks and practical applications. To address this discrepancy, we conduct a detailed investigation into the application of tabular data in industrial scenarios and propose a comprehensive and complex benchmark TableBench, including 18 fields within four major categories of table question answering (TableQA) capabilities. Furthermore, we introduce TableLLM, trained on our meticulously constructed training set TableInstruct, achieving comparable performance with GPT-3.5. Massive experiments conducted on TableBench indicate that both open-source and proprietary LLMs still have significant room for improvement to meet real-world demands, where the most advanced model, GPT-4, achieves only a modest score compared to humans.
Xianjie Wu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Xeron Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Tongliang Li, Zhoujun Li 0001, Guanglin Niu
AAAI1
2025 XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
abstract
In the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task. The advent of pre-trained multimodal models significantly empowers Document AI frameworks to extract key information from form documents in different formats such as PDF, Word, and images. Nonetheless, form parsing is still encumbered by notable challenges like subpar capabilities in multilingual parsing and diminished recall in industrial contexts in rich text and rich visuals. In this work, we introduce a simple but effective Multimodal and Multilingual semi-structured FORM PARSER (XFormParser), which is anchored on a comprehensive Transformer-based pre-trained language model and innovatively amalgamates semantic entity recognition (SER) and relation extraction (RE) into a unified framework. Combined with Bi-LSTM, the performance of multilingual parsing is significantly improved. Furthermore, we develop InDFormSFT, a pioneering supervised fine-tuning (SFT) industrial dataset that specifically addresses the parsing needs of forms in a variety of industrial contexts. Through rigorous testing on established benchmarks, XFormParser has demonstrated its unparalleled effectiveness and robustness. Compared to existing state-of-the-art (SOTA) models, XFormParser notably achieves up to 1.79% F1 score improvement on RE tasks in language-specific settings. It also exhibits exceptional improvements in cross-task performance in both multilingual and zero-shot settings.
Xianfu Cheng, Jian Yang 0030, Xiang Li 0117, Weixiao Zhou, Kui Wu 0007, Xiangyuan Guan, Tao Sun 0016, Xianjie Wu, Tongliang Li, Zhoujun Li 0001
COLING10
2025 Breaking Size Barrier: Enhancing Reasoning for Large-Size Table Question Answering
Xianjie Wu, Di Liang, Jian Yang 0037, Xianfu Cheng, Linzheng Chai, Tongliang Li, Liqun Yang, Zhoujun Li 0001
DASFAA (2)1
2025 Unleashing Potential of Evidence in Knowledge-Intensive Dialogue Generation
abstract
Incorporating external knowledge into dialogue generation (DG) is crucial for enhancing response accuracy, where evidence fragments serve as effective knowledgeable snippets that support factual dialogue replies. However, introducing irrelevant content beyond valid knowledge fragments can adversely affect reply quality and lead to hallucinated responses. Prior work relies on manual annotations to develop models for identifying evidence within external knowledge. However, these annotations often cover only a limited portion of the valid evidence, restricting the ability of models to mine useful evidence from retrieved knowledge. To fully Unleash the potential of evidence, we propose a framework to effectively incorporate Evidence in knowledge-Intensive Dialogue Generation (U-EIDG). Specifically, we develop an evidence miner (Evid-M) that harnesses the power of large language models (LLMs) to mine reliable evidence labels from external knowledge. Subsequently, we propose an evidence indicator (Evid-I) to effectively identify valid evidence from retrieved knowledge by utilizing these evidence labels. Furthermore, we introduce an evidence-augmented generator (EAG) incorporating an evidence-attention mechanism that enables the model to focus on segments supported by evidence. Experimental results on the MultiDoc2Dial and WoW benchmarks indicate that the proposed method significantly outperforms other baselines, with a +3∼5 points improvement in Rouge-L. Further analysis confirms the effectiveness of fully mining valid evidence fragments for knowledge-intensive dialogue generation.
Xianjie Wu, Jian Yang 0030, Tongliang Li, Yiyang Du, Linzheng Chai, Di Liang, Zhoujun Li 0001
ICASSP1
2025 SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
abstract
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work, we introduce SimpleVQA, the first comprehensive multi-modal benchmark to evaluate the factuality ability of MLLMs to answer natural language short questions. SimpleVQA is characterized by six key features: it covers multiple tasks and multiple scenarios, ensures high quality and challenging queries, maintains static and timeless reference answers, and is straightforward to evaluate. Our approach involves categorizing visual question-answering items into 9 different tasks around objective events or common knowledge and situating these within 9 topics. Rigorous quality control processes are implemented to guarantee high-quality, concise, and clear answers, facilitating evaluation with minimal variance via an LLM-as-a-judge scoring system. Using SimpleVQA, we perform a comprehensive assessment of leading 18 MLLMs and 8 text-only LLMs, delving into their image comprehension and text generation abilities by identifying and analyzing error cases.
Xianfu Cheng, Wei Zhang 0384, Jian Yang 0030, Xiangyuan Guan, Xianjie Wu, Xiang Li 0117, Ge Zhang 0009, Yuying Mai, Yutao Zeng, Zhoufutu Wen, Baorui Wang, Weixiao Zhou, Yunhong Lu, Hangyuan Ji, Tongliang Li, Wenhao Huang 0001, Zhoujun Li 0001
ICCV6
2025 McEval: Massively Multilingual Code Evaluation
abstract
Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval.
Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001
ICLR13
2025 Informative Memory Mechanism for Enhancing Crisis Response in Large Language Models
abstract
In recent years, the advancements of Large Language Models (LLMs) have enhanced human-AI interactions and memory systems, but their event results tend to be general and weak in differentiation, which will bring high-impact event results. These factors are crucial for risk assessment and decision-making. Therefore, we introduce the Informativeness-Driven Memory Mechanism (IDMM), which assesses results based on their complexity, uniqueness, and potential future impact to prioritize important memories. The IDMM continuously evaluates and adjusts the strength of memories, reinforcing critical events while allowing less informative ones to decay. This dynamic memory management system is integrated into the ShannonGuard agent, which demonstrates a 20% improvement in event recall accuracy, a 15% increase in memory persistence, and a 10-18% boost in decision-making effectiveness across multilingual environments. As a result, the approach of IDMM holds promise for domains that require handling rare, high-stakes events, offering broad applicability in adaptive memory retention.
Boyang Wang 0006, Xianjie Wu, Zhoujun Li 0001
IJCNN2
2023 Read Then Respond: Multi-granularity Grounding Prediction for Knowledge-Grounded Dialogue Generation
Yiyang Du, Shi-Wei Zhang, Xianjie Wu, Yunbo Cao, Zhoujun Li 0001
ADMA (2)3