Linzheng Chai

dblp:320/5967 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
0009-0001-2129-3207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QA
abstract
Tables serve as a core format for representing structured data on the web, as their two-dimensional layouts effectively encode complex inter-entity relationships. However, real-world web tables often feature heterogeneous structures and rich semantics. Accurately interpreting such tables requires not only spatial layout perception but also multi-step reasoning across rows and columns, posing substantial challenges to web intelligence systems. Multimodal large language models (MLLMs) show promise in table question answering (TableQA) by leveraging visual layouts. However, their performance on complex web tables remains uneven, as existing benchmarks often blur the impact of individual difficulty factors, hindering precise capability analysis. To advance TableQA beyond superficial task difficulty and toward interpretable capability modeling, we introduce MMTableBench, a multi-level benchmark that systematically evaluates MLLMs along two fine-grained dimensions: layout complexity and reasoning complexity. By organizing table-question pairs along these axes, MMTableBench facilitates a detailed evaluation of model performance under varying structural and reasoning challenges, while revealing the respective strengths and limitations of multimodal inputs. Our comprehensive analysis shows that state-of-the-art MLLMs continue to exhibit notable limitations when confronted with complex layouts and deep reasoning tasks, underscoring persistent gaps despite the structural advantages offered by visual inputs. MMTableBench thus provides not only a rigorous evaluation framework but also a diagnostic tool for analyzing and interpreting model behaviors, enabling more transparent and explainable progress in multimodal TableQA development.
Xianjie Wu, Xiaohang Xu 0002, Tingyu Jiang, Jian Yang 0030, Di Liang, Xianfu Cheng, Zhenhe Wu, Linzheng Chai, Wei Zhang 0384, Ge Zhang 0009, Bob Simons, Tongliang Li, Zhoujun Li 0001
WWW8
2025 XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
abstract
Chain-of-thought (CoT) has emerged as a powerful technique to elicit reasoning in large language models and improve a variety of downstream tasks. CoT mainly demonstrates excellent performance in English, but its usage in low-resource languages is constrained due to poor language generalization. To bridge the gap among different languages, we propose a cross-lingual instruction fine-tuning framework (xCoT) to transfer knowledge from high-resource languages to low-resource languages. Specifically, the multilingual instruction training data (xCoT-Instruct) is created to encourage the semantic alignment of multiple languages. We introduce cross-lingual in-context few-shot learning (xICL) to accelerate multilingual agreement in instruction tuning, where some fragments of source languages in examples are randomly substituted by their counterpart translations of target languages. During multilingual instruction tuning, we adopt the randomly online CoT strategy to enhance the multilingual reasoning ability of the large language model by first translating the query to another language and then answering in English. To further facilitate the language transfer, we leverage the high-resource CoT to supervise the training of low-resource languages with cross-lingual distillation. Experimental results demonstrate the superior performance of xCoT in reducing the gap among different languages, highlighting its potential to reduce the cross-lingual gap.
Linzheng Chai, Jian Yang 0030, Tao Sun 0016, Hongcheng Guo, Xinnian Liang, Jiaqi Bai 0001, Tongliang Li, Qiyao Peng 0001, Zhoujun Li 0001
AAAI1
2025 TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
abstract
Recent advancements in Large Language Models (LLMs) have markedly enhanced the interpretation and processing of tabular data, introducing previously unimaginable capabilities. Despite these achievements, LLMs still encounter significant challenges when applied in industrial scenarios, particularly due to the increased complexity of reasoning required with real-world tabular data, underscoring a notable disparity between academic benchmarks and practical applications. To address this discrepancy, we conduct a detailed investigation into the application of tabular data in industrial scenarios and propose a comprehensive and complex benchmark TableBench, including 18 fields within four major categories of table question answering (TableQA) capabilities. Furthermore, we introduce TableLLM, trained on our meticulously constructed training set TableInstruct, achieving comparable performance with GPT-3.5. Massive experiments conducted on TableBench indicate that both open-source and proprietary LLMs still have significant room for improvement to meet real-world demands, where the most advanced model, GPT-4, achieves only a modest score compared to humans.
Xianjie Wu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Xeron Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Tongliang Li, Zhoujun Li 0001, Guanglin Niu
AAAI3
2025 OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
abstract
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001
ACL (1)11
2025 M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation
abstract
Jiaheng Liu, Ken Deng, Congnan Liu, Jian Yang, Shukai Liu, He Zhu, Peng Zhao, Linzheng Chai, Yanan Wu, JinKe JinKe, Ge Zhang, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ken Deng, Congnan Liu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007
ACL (1)8
2025 MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL
abstract
Recent LLM-based Text-to-SQL methods usually suffer from significant performance degradation on “huge” databases and complex user questions that require multi-step reasoning. Moreover, most existing methods neglect the crucial significance of LLMs utilizing external tools and model collaboration. To address these challenges, we introduce MAC-SQL, a novel LLM-based multi-agent collaborative framework. Our framework comprises a core decomposer agent for Text-to-SQL generation with few-shot chain-of-thought reasoning, accompanied by two auxiliary agents that utilize external tools or models to acquire smaller sub-databases and refine erroneous SQL queries. The decomposer agent collaborates with auxiliary agents, which are activated as needed and can be expanded to accommodate new features or tools for effective Text-to-SQL parsing. In our framework, We initially leverage GPT-4 as the strong backbone LLM for all agent tasks to determine the upper bound of our framework. We then fine-tune an open-sourced instruction-followed model, SQL-Llama, by leveraging Code Llama 7B, to accomplish all tasks as GPT-4 does. Experiments show that SQL-Llama achieves a comparable execution accuracy of 43.94, compared to the baseline accuracy of 46.35 for vanilla GPT-4. At the time of writing, MAC-SQL+GPT-4 achieves an execution accuracy of 59.59 when evaluated on the BIRD benchmark, establishing a new state-of-the-art (SOTA) on its holdout test set.
Changyu Ren, Jian Yang 0030, Xinnian Liang, Jiaqi Bai 0001, Linzheng Chai, Qian-Wen Zhang, Xing Sun 0001, Zhoujun Li 0001
COLING6
2025 Breaking Size Barrier: Enhancing Reasoning for Large-Size Table Question Answering
Xianjie Wu, Di Liang, Jian Yang 0037, Xianfu Cheng, Linzheng Chai, Tongliang Li, Liqun Yang, Zhoujun Li 0001
DASFAA (2)5
2025 Unleashing Potential of Evidence in Knowledge-Intensive Dialogue Generation
abstract
Incorporating external knowledge into dialogue generation (DG) is crucial for enhancing response accuracy, where evidence fragments serve as effective knowledgeable snippets that support factual dialogue replies. However, introducing irrelevant content beyond valid knowledge fragments can adversely affect reply quality and lead to hallucinated responses. Prior work relies on manual annotations to develop models for identifying evidence within external knowledge. However, these annotations often cover only a limited portion of the valid evidence, restricting the ability of models to mine useful evidence from retrieved knowledge. To fully Unleash the potential of evidence, we propose a framework to effectively incorporate Evidence in knowledge-Intensive Dialogue Generation (U-EIDG). Specifically, we develop an evidence miner (Evid-M) that harnesses the power of large language models (LLMs) to mine reliable evidence labels from external knowledge. Subsequently, we propose an evidence indicator (Evid-I) to effectively identify valid evidence from retrieved knowledge by utilizing these evidence labels. Furthermore, we introduce an evidence-augmented generator (EAG) incorporating an evidence-attention mechanism that enables the model to focus on segments supported by evidence. Experimental results on the MultiDoc2Dial and WoW benchmarks indicate that the proposed method significantly outperforms other baselines, with a +3∼5 points improvement in Rouge-L. Further analysis confirms the effectiveness of fully mining valid evidence fragments for knowledge-intensive dialogue generation.
Xianjie Wu, Jian Yang 0030, Tongliang Li, Yiyang Du, Linzheng Chai, Di Liang, Zhoujun Li 0001
ICASSP6
2025 McEval: Massively Multilingual Code Evaluation
abstract
Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval.
Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001
ICLR1
2024 UniCoder: Scaling Code Large Language Model via Universal Code
abstract
Tao Sun, Linzheng Chai, Jian Yang, Yuwei Yin, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, Zhoujun Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tao Sun 0016, Linzheng Chai, Jian Yang 0030, Yuwei Yin, Hongcheng Guo, Liqun Yang, Zhoujun Li 0001
ACL (1)2
2024 m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt
abstract
Multilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario.
Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001
LREC/COLING8
2024 OWL: A Large Language Model for IT Operations
abstract
With the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition, machine translation, and dialogue systems. Recently, Large Language Models (LLMs) have achieved significant improvements across various domain-specific areas. However, there is a noticeable gap in the development of specialized Large Language Models (LLMs) tailored for IT operations. In this paper, we introduce the OWL, a large language model trained on our constructed Owl-Instruct with a wide range of IT-related information. Specifically, limited by the maximum input length, we propose the \textbf{H}omogeneous \textbf{M}arkov \textbf{C}ontext \textbf{E}xtension method (HMCE). The mixture-of-adapter strategy is leveraged to improve the parameter-efficient tuning across different domains or tasks. Further, we evaluate the performance of OWL on the Owl-Bench established by us and open IT-related benchmarks. OWL demonstrates superior performance results on IT tasks, which outperforms existing models by significant margins. Moreover, we hope that the findings of our work will provide more insights to revolutionize the techniques of IT operations with specialized LLMs.
Hongcheng Guo, Jian Yang 0030, Liqun Yang, Linzheng Chai, Jiaqi Bai 0001, Junran Peng, Xiaorong Hu, Dongfeng Zhang, Xu Shi 0005, Tieqiao Zheng, Liangfan Zheng, Bo Zhang 0096, Ke Xu 0001, Zhoujun Li 0001
ICLR5
2024 mt4CrossOIE: Multi-stage tuning for cross-lingual open information extraction
Tongliang Li, Linzheng Chai, Jian Yang 0030, Jiaqi Bai 0001, Yuwei Yin, Hongcheng Guo, Liqun Yang, Hebboul Zine El Abidine, Zhoujun Li 0001
Expert Syst. Appl.3
2023 QURG: Question Rewriting Guided Context-Dependent Text-to-SQL Semantic Parsing
Linzheng Chai, Dongling Xiao, Jian Yang 0030, Liqun Yang, Qian-Wen Zhang, Yunbo Cao, Zhoujun Li 0001
PRICAI (2)1