Liangyou Li

dblp:78/7942 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-0279-003XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learning
abstract
Tool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning.
Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001
AAAI6
2026 Graph-based dual-attention model for multi-bend tube forming quality prediction with basis spline cross-sectional fitting
Zheyi Li, Zili Wang 0001, Shuyou Zhang 0001, Yaochen Lin, Liangyou Li, Jianrong Tan, Yonglin Tao
Eng. Appl. Artif. Intell.5
2026 Spatial spiral tube multi-roller bending: Accurate axial prediction utilizing AWPSO-FECAM-LSTM framework
Zili Wang 0001, Yonglin Tao, Shuyou Zhang 0001, Xiaojian Liu 0002, Yaochen Lin, Liangyou Li, Jianrong Tan, Zheyi Li
Expert Syst. Appl.6
2025 Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
abstract
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002
ACL (1)6
2025 Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
abstract
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qiyuan Zhang 0001, Yufei Wang 0005, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ACL (1)4
2025 NILE: Internal Consistency Alignment in Large Language Models
abstract
Minda Hu, Qiyuan Zhang, Yufei Wang, Bowei He, Hongru Wang, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma, Irwin King. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Minda Hu, Qiyuan Zhang 0001, Yufei Wang 0005, Bowei He, Hongru Wang 0003, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma 0001, Irwin King
EMNLP7
2025 Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP5
2025 Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
abstract
Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants.
Bo Huang 0017, Yufei Wang 0005, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Wei Wang 0011
ICLR5
2025 RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
abstract
With significant efforts in recent studies, LLM-as-a-Judge has become a cost-effective alternative to human evaluation for assessing text generation quality in a wide range of tasks. However, there still remains a reliability gap between LLM-as-a-Judge and human evaluation. One important reason is the lack of guided oracles in the evaluation process. Motivated by the role of reference pervasively used in classic text evaluation, we introduce RevisEval, a novel text generation evaluation paradigm via the response-adapted references. RevisEval is driven by the key observation that an ideal reference should maintain the necessary relevance to the response to be evaluated. Specifically, RevisEval leverages the text revision capabilities of large language models (LLMs) to adaptively revise the response, then treat the revised text as the reference (response-adapted reference) for the subsequent evaluation. Extensive experiments demonstrate that RevisEval outperforms traditional reference-free and reference-based evaluation paradigms that use LLM-as-a-Judge across NLG tasks and open-ended instruction-following tasks. More importantly, our response-adapted references can further boost the classical text metrics, e.g., BLEU and BERTScore, compared to traditional references and even rival the LLM-as-a-Judge. A detailed analysis is also conducted to confirm RevisEval's effectiveness in bias reduction, the impact of inference cost, and reference relevance.
Qiyuan Zhang 0001, Yufei Wang 0005, Tiezheng Yu, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ICLR6
2025 ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL (Long Papers)4
2024 FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
abstract
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011
ACL (1)5
2024 Learning to Edit: Aligning LLMs with Knowledge Editing
abstract
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011
ACL (1)7
2024 M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun, Liangyou Li, Yuxin Jiang, Lifeng Shang, Qun Liu, Kam-Fai Wong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Yusen Sun, Liangyou Li, Lifeng Shang, Qun Liu 0001, Kam-Fai Wong
ACL (1)5
2024 MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP5
2022 Universal Conditional Masked Language Pre-training for Neural Machine Translation
abstract
Pre-trained sequence-to-sequence models have significantly improved Neural Machine Translation (NMT).Different from prior works where pre-trained models usually adopt an unidirectional decoder, this paper demonstrates that pre-training a sequenceto-sequence model but with a bidirectional decoder can produce notable performance gains for both Autoregressive and Nonautoregressive NMT.Specifically, we propose CeMAT, a conditional masked language model pre-trained on large-scale bilingual and monolingual corpora in many languages.1 We also introduce two simple but effective methods to enhance the CeMAT, aligned code-switching & masking and dynamic dual-masking.We conduct extensive experiments and show that our CeMAT can achieve significant performance improvement for all scenarios from low-to extremely highresource languages, i.e., up to +14.4 BLEU on low-resource and +7.9 BLEU on average for Autoregressive NMT.For Non-autoregressive NMT, we demonstrate it can also produce consistent performance gains, i.e., up to +5.3 BLEU.To the best of our knowledge, this is the first work to pre-train a unified model for fine-tuning on both NMT tasks.
Liangyou Li, Meng Zhang 0019, Minghao Wu, Qun Liu 0001
ACL (1)2
2021 Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter Corruption
abstract
We argue that the vulnerability of model parameters is of crucial value to the study of model robustness and generalization but little research has been devoted to understanding this matter. In this work, we propose an indicator to measure the robustness of neural network parameters by exploiting their vulnerability via parameter corruption. The proposed indicator describes the maximum loss variation in the non-trivial worst-case scenario under parameter corruption. For practical purposes, we give a gradient-based estimation, which is far more effective than random corruption trials that can hardly induce the worst accuracy degradation. Equipped with theoretical support and empirical validation, we are able to systematically investigate the robustness of different model parameters and reveal vulnerability of deep neural networks that has been rarely paid attention to before. Moreover, we can enhance the models accordingly with the proposed adversarial corruption-resistant training, which not only improves the parameter robustness but also translates into accuracy elevation.
Xu Sun 0001, Zhiyuan Zhang 0001, Xuancheng Ren, Ruixuan Luo, Liangyou Li
AAAI5
2021 Future-Guided Incremental Transformer for Simultaneous Translation
abstract
Simultaneous translation (ST) starts translations synchronously while reading source sentences, and is used in many online scenarios. The previous wait-k policy is concise and achieved good results in ST. However, wait-k policy faces two weaknesses: low training speed caused by the recalculation of hidden states and lack of future source information to guide training. For the low training speed, we propose an incremental Transformer with an average embedding layer (AEL) to accelerate the speed of calculation of the hidden states during training. For future-guided training, we propose a conventional Transformer as the teacher of the incremental Transformer, and try to invisibly embed some future information in the model through knowledge distillation. We conducted experiments on Chinese-English and German-English simultaneous translation tasks and compared with the wait-k policy to evaluate the proposed method. Our method can effectively increase the training speed by about 28 times on average at different k and implicitly embed some predictive abilities in the model, achieving better translation quality than wait-k baseline.
Shaolei Zhang 0001, Yang Feng 0004, Liangyou Li
AAAI3
2021 Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation Training
abstract
Learning multilingual and multi-domain translation model is challenging as the heterogeneous and imbalanced data make the model converge inconsistently over different corpora in real world.One common practice is to adjust the share of each corpus in the training, so that the learning process is balanced and low-resource cases can benefit from the highresource ones.However, automatic balancing methods usually depend on the intra-and interdataset characteristics, which is usually agnostic or requires human priors.In this work, we propose an approach, MULTIUAT, that dynamically adjusts the training data usage based on the model's uncertainty on a small set of trusted clean data for multi-corpus machine translation.We experiment with two classes of uncertainty measures on multilingual (16 languages with 4 settings) and multi-domain settings (4 for in-domain and 2 for out-of-domain on English-German translation) and demonstrate our approach MULTIUAT substantially outperforms its baselines, including both static and dynamic strategies.We analyze the crossdomain transfer and show the deficiency of static and similarity based methods. 1
Minghao Wu, Meng Zhang 0019, Liangyou Li, Gholamreza Haffari, Qun Liu 0001
EMNLP (1)4
2021 Document Graph for Neural Machine Translation
abstract
Previous works have shown that contextual information can improve the performance of neural machine translation (NMT).However, most existing document-level NMT methods only consider a few number of previous sentences.How to make use of the whole document as global contexts is still a challenge.To address this issue, we hypothesize that a document can be represented as a graph that connects relevant contexts regardless of their distances.We employ several types of relations, including adjacency, syntactic dependency, lexical consistency, and coreference, to construct the document graph.Then, we incorporate both source and target graphs into the conventional Transformer architecture with graph convolutional networks.Experiments on various NMT benchmarks, including IWSLT English-French, Chinese-English, WMT English-German and Opensubtitle English-Russian, demonstrate that using document graphs can significantly improve the translation quality.Extensive analysis verifies that the document graph is beneficial for capturing discourse phenomena.
Mingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu 0001, Lidia S. Chao
EMNLP (1)2
2021 Adversarial parameter defense by multi-step risk minimization
Zhiyuan Zhang 0001, Ruixuan Luo, Xuancheng Ren, Qi Su 0001, Liangyou Li, Xu Sun 0001
Neural Networks5
2016 Graph-Based Translation Via Graph Segmentation
abstract
One major drawback of phrase-based translation is that it segments an input sentence into continuous phrases.To support linguistically informed source discontinuity, in this paper we construct graphs which combine bigram and dependency relations and propose a graph-based translation model.The model segments an input graph into connected subgraphs, each of which may cover a discontinuous phrase.We use beam search to combine translations of each subgraph left-to-right to produce a complete translation.Experiments on Chinese-English and German-English tasks show that our system is significantly better than the phrase-based model by up to +1.5/+0.5 BLEU scores.By explicitly modeling the graph segmentation, our system obtains further improvement, especially on German-English.
Liangyou Li, Andy Way, Qun Liu 0001
ACL (1)1
2016 Topic-Informed Neural Machine Translation
abstract
In recent years, neural machine translation (NMT) has demonstrated state-of-the-art machine translation (MT) performance. It is a new approach to MT, which tries to learn a set of parameters to maximize the conditional probability of target sentences given source sentences. In this paper, we present a novel approach to improve the translation performance in NMT by conveying topic knowledge during translation. The proposed topic-informed NMT can increase the likelihood of selecting words from the same topic and domain for translation. Experimentally, we demonstrate that topic-informed NMT can achieve a 1.15 (3.3% relative) and 1.67 (5.4% relative) absolute improvement in BLEU score on the Chinese-to-English language pair using NIST 2004 and 2005 test sets, respectively, compared to NMT without topic information.
Jian Zhang 0003, Liangyou Li, Andy Way, Qun Liu 0001
COLING2
2016 Combining Translation Memories and Syntax-Based SMT: Experiments with Real Industrial Data
Liangyou Li, Carla Parra Escartín, Qun Liu 0001
EAMT1
2016 Combining translation memories and statistical machine translation using sparse features
Liangyou Li, Carla Parra Escartín, Andy Way, Qun Liu 0001
Mach. Transl.1
2015 Dependency Graph-to-String Translation
abstract
Compared to tree grammars, graph grammars have stronger generative capacity over structures.Based on an edge replacement grammar, in this paper we propose to use a synchronous graph-to-string grammar for statistical machine translation.The graph we use is directly converted from a dependency tree by labelling edges.We build our translation model in the log-linear framework with standard features.Large-scale experiments on Chinese-English and German-English tasks show that our model is significantly better than the state-of-the-art hierarchical phrase-based (HPB) model and a recently improved dependency tree-to-string model on BLEU, METEOR and TER scores.Experiments also suggest that our model has better capability to perform long-distance reordering and is more suitable for translating long sentences.
Liangyou Li, Andy Way, Qun Liu 0001
EMNLP1
2011 Improve SMT with Source-Side "Topic-Document" Distributions
Zhengxian Gong, Guodong Zhou 0001, Liangyou Li
MTSummit3