Wei Liu 0302

dblp:49/3283-302 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
28since 2021 · last 2026
0009-0009-4327-1920ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 1 first-author · 25 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
abstract
Large language models (LLMs) are versatile, yet their deployment in complex real-world settings is limited by static knowledge cutoffs and the difficulty of producing controllable behavior within a single inference.Multi-agent search systems (MASS), which coordinate specialized LLM agents equipped with search tools, mitigate these issues via task decomposition and retrieval-augmented problem solving.However, optimizing LLMs for agent-specific roles remains labor-intensive with prompt engineering or supervised fine-tuning, motivating automated end-to-end training.Existing multiagent reinforcement learning (MARL) methods such as Multi-Agent Proximal Policy Optimization (MAPPO) typically depend on large critic networks to evaluate joint actions, leading to instability and high memory costs.We introduce Multi-Agent Heterogeneous Group Policy Optimization (MHGPO), which updates policies by estimating relative advantages across heterogeneous groups of multi-agent rollouts, shifting the optimization focus from local agent performance to global system success.We further study three group rollout sampling strategies to trade off sample efficiency and optimization quality.Experiments show that MHGPO captures implicit inter-agent dependencies and consistently outperforms strong baselines in both task performance and computational efficiency.
Shaoxiong Yang, Chao Li 0033, Wei Liu 0302, Jian Luan 0001, Zenglin Xu
ACL (1)4
2026 VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
abstract
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin, Wei Liu, Jian Luan, Weiping Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin 0001, Wei Liu 0302, Jian Luan 0001, Weiping Wang 0005
ACL (1)5
2026 Attention Basin: Why Contextual Position Matters in Large Language Models
abstract
Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu, Wei Liu, Jian Luan, Wanxia Cao, Ying Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu 0009, Wei Liu 0302, Jian Luan 0001, Wanxia Cao, Ying Shen 0001
ACL (1)6
2026 Mobile GUI Agents under Real-world Threats: Are We There Yet?
abstract
Recent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. These contents may inevitably appear in the agents' observation space and influence the task execution process. Systematic investigation of this problem is challenging since the real-world app contents are significantly skewed—testing on normal real-world apps usually cannot uncover any potential risk since most app contents are benign. To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at https://agenthazard.github.io.
Guohong Liu 0002, Jialei Ye, Wei Liu 0302, Pengzhi Gao, Jian Luan 0001, Yuanchun Li 0003, Yunxin Liu 0001
MobiSys4
2026 C2-Cite: Contextual-Aware Citation Generation for Attributed Large Language Models
abstract
The attribution technique enhances the credibility of LLMs by adding citations to the generated sentences, enabling users to trace back to the original sources and verify the reliability of the output. However, existing instruction-tuned attributed LLMs often fail to properly interpret the contextual semantics of citation symbols (e.g., [i]) during text generation. This shortcoming arises from their insufficient awareness of the context information surrounding citation markers, which in turn leads to disjointed references and poor integration of retrieved knowledge into the generated content. To address this issue, we propose a novel Contextual-aware Citation generation framework (C²-Cite) that explicitly integrates the semantic relationships between citation markers and their referenced content. Specifically, a contextual citation alignment mechanism is adopted: it first encodes the retrieved document contexts into the symbol representation of citations, then aligns the marker numbers by decoding information from a citation router function. This mechanism enables the transformation of citation markers from generic placeholders into active knowledge pointers that link to the referenced source information. Experimental results on the ALCE benchmark across three datasets validate our framework C²-Cite++: it outperforms the SOTA baseline by an average of 5.8% in citation quality and 17.4% in response correctness. The implementation is publicly available at https://github.com/BAI-LAB/c2cite
Yue Yu 0007, Ting Bai 0004, Hengzhi Lan, Jie Wu 0017, Wei Liu 0302, Jian Luan 0001, Chuan Shi 0001
WSDM7
2026 FwdLLM+: Accelerating Forward-Only FedLLM With Low-Rank Perturbations
abstract
Federated Learning (FL) facilitates privacy-preserving fine-tuning of Large Language Models (LLMs) for mobile applications, termed FedLLM. A vital challenge of FedLLM is the tension between LLM complexity and resource constraint of mobile devices. In response to this challenge, we first introduceFwdLLM(our conference version), an innovative FL framework designed to enhance the FedLLM efficiency. The key idea ofFwdLLMis to employ backpropagation (BP)-free training methods, requiring devices only to execute memory-efficient “perturbed inference”. Enabled by mobile NPU acceleration and an expanded array of participating devices,FwdLLMdelivers substantially better wall-clock efficiency than BP-based FedLLM. However,FwdLLMbased on vanilla BP-free optimization theoretically requires more optimization steps to converge. In this work, we further enhanceFwdLLMtoFwdLLM+, which incorporates advanced zeroth-order optimization techniques and low-rank perturbation decomposition to reduce convergence steps. Finally, we conduct extensive experiments on 4 models (ranging from 110 M to 7B) and 8 more datasets, demonstrating thatFwdLLM+achieves up to 151× faster training, a 93$\%$memory reduction compared to vanilla BP-based FedLLM, and superior performance compared toFwdLLM, enabling efficient federated fine-tuning of billion-parameter LLMs on commodity mobile devices.
Mengwei Xu 0001, Zhenyan Lu, Wei Liu 0302, Shangguang Wang, Nicholas D. Lane, Qibo Sun, Dongqi Cai 0001
IEEE Trans. Mob. Comput.4
2025 Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology
abstract
Zeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been many efforts to analyze the optimization convergence rate of zeroth-order decentralized stochastic gradient descent (ZO-DSGD) algorithms. However, the generalization of these methods has not been well studied. In this paper, we provide a generalization analysis of ZO-DSGD with changing topology, where the clients run zeroth-order SGD with local data and communicate with each other according to time-varying topology. We systematically analyze the generalization error in convex, strongly convex, and non-convex cases. The obtained results in the convex and strongly convex cases with zeroth-order oracles recover the results of SGD. Moreover, the generalization bounds derived in non-convex cases align with that of DSGD. To capture the influence of communication topology on the generalization performance, we analyze local generalization bounds concerning local models held at different clients. The obtained results reflect the influence of the number of clients, local sample size, and topology on the generalization error. To the best of our knowledge, this is the first work that provides a generalization analysis of zeroth-order decentralized stochastic gradient descent methods and recovers the results of SGD.
Xiaolin Hu 0001, Zixuan Gong, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
AAAI4
2025 HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
abstract
Many positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position carry less relevant information. We argue that long-term decay is outdated in the era of LLMs, as LLMs are now applied to tasks demanding precise retrieval of in-context information from arbitrary positions. Firstly, we present empirical analyses on various PEs, demonstrating that models inherently learn attention with only a local-decay pattern while forming a U-shape pattern globally, contradicting the principle of long-term decay. Furthermore, we conduct a detailed analysis of rotary position encoding (RoPE, a prevalent relative positional encoding in LLMs), and found that the U-shape attention is caused by some learned components, which are also the key factor limiting RoPE’s expressiveness and extrapolation. Inspired by these insights, we propose High-frequency rotary Position Encoding (HoPE). HoPE replaces the specific components in RoPE with position-independent ones, retaining only high-frequency signals, which also breaks the principle of long-term decay in theory. HoPE achieves two major advantages: (1) Without constraints imposed by long-term decay, contradictory factors that limit attention optimization are removed. Thus, the model’s context awareness is enhanced. (2) HoPE exhibits greater robustness to the out-of-distribution behavior in attention patterns during extrapolation. The effectiveness of HoPE is validated through extensive experiments and with a large language model of up to 3 billion parameters.
Yuhan Chen 0001, Ang Lv, Jian Luan 0001, Bin Wang 0004, Wei Liu 0302
ACL (1)5
2025 Demystifying Small Language Models for Edge Deployment
abstract
Small language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show that state-of-the-art SLMs outperform 7B models in general tasks, proving their practical viability. However, SLMs’ in-context learning capabilities remain limited, and their efficiency has significant optimization potential. We identify key SLM optimization opportunities, including dynamic task-specific routing, model-hardware co-design, and vocabulary/KV cache compression. Overall, we expect the work to reveal an all-sided landscape of SLMs, benefiting the research community across algorithm, model, system, and hardware levels.
Zhenyan Lu, Xiang Li 0067, Dongqi Cai 0001, Rongjie Yi, Fangming Liu, Wei Liu 0302, Jian Luan 0001, Nicholas D. Lane, Mengwei Xu 0001
ACL (1)6
2025 Global Eye: Breaking the "Fixed Thinking Pattern" during the Instruction Expansion Process
abstract
An extensive high-quality instruction dataset is crucial for the instruction tuning process of Large Language Models (LLMs).Recent instruction expansion methods have demonstrated their capability to improve the quality and quantity of existing datasets, by prompting high-performance LLM to generate multiple new instructions from the original ones.However, existing methods focus on constructing multi-perspective prompts (e.g., increasing complexity or difficulty) to expand instructions, overlooking the "Fixed Thinking Pattern" issue of LLMs.This issue arises when repeatedly using the same set of prompts, causing LLMs to rely on a limited set of certain expressions to expand all instructions, potentially compromising the diversity of the final expanded dataset.This paper theoretically analyzes the causes of the "Fixed Thinking Pattern", and corroborates this phenomenon through multi-faceted empirical research.Furthermore, we propose a novel method based on dynamic prompt updating: Global Eye.Specifically, after a fixed number of instruction expansions, we analyze the statistical characteristics of newly generated instructions and then update the prompts.Experimental results show that our method enables LLaMA3-8B and LLaMA2-13B to surpass the performance of open-source LLMs and GPT3.5 across various metrics.
Wenxuan Lu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Songhao Jiang, Tianning Zang
ACL (1)2
2025 Browsing Like Human: A Multimodal Web Agent with Experiential Fast-and-Slow Thinking
abstract
Automating web navigation which aims to build a web agent that follows user instructions to complete tasks like booking flights by interacting with websites, has received increasing attention due to its practical value.Although existing web agents are mostly equipped with visual perception, planning, and memory abilities, their reasoning process are still deviate from human cognition.In this work, we study the human thought pattern to empower agent with more human-like abilities in web navigation.To tackle this problem, we propose a novel multimodal web agent framework called WebExperT, which is designed to emulate the human planning process of "thinking fast and slow" to effectively decompose complex user instructions.Furthermore, WebExperT leverages experiential learning by reflecting from failure for continuously refining planning and decision-making outcomes.Experimental results on the MIND2WEB benchmark demonstrate the superiority of WebExperT in both supervised and unsupervised settings.
Haohao Luo, Jiayi Kuang, Wei Liu 0302, Ying Shen 0001, Jian Luan 0001, Yang Deng 0002
ACL (1)3
2025 Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
abstract
Vision-language models (VLMs) achieve remarkable success in single-image tasks.However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features.In this work, we propose Focus-Centric Visual Chain, a novel paradigm that enhances VLMs' perception, comprehension, and reasoning abilities in multi-image scenarios.To facilitate this paradigm, we propose Focus-Centric Data Synthesis, a scalable bottom-up approach for synthesizing high-quality data with elaborate reasoning paths.Through this approach, We construct VISC-150K, a large-scale dataset with reasoning data in the form of Focus-Centric Visual Chain, specifically designed for multi-image tasks.Experimental results on seven multi-image benchmarks demonstrate that our method achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities.Our study represents a significant step toward more robust and capable vision-language systems that can handle complex visual scenarios: VISC. * Corresponding authors.Which of the following images contains the same object as the first image and shares the same attribute weight?
Juntian Zhang, Chuanqi Cheng, Yuhan Liu 0023, Wei Liu 0302, Jian Luan 0001, Rui Yan 0001
ACL (1)4
2025 More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
abstract
Xiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung, Wei Liu, Jian Luan, Shuo Shang, Xiuying Chen, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaoqing Zhang 0017, Ang Lv, Yuhan Liu 0023, Flood Sung, Wei Liu 0302, Jian Luan 0001, Shuo Shang, Xiuying Chen, Rui Yan 0001
ACL (1)5
2025 PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
abstract
Low-rank adaptation (LoRA) and its variants have recently gained much interest due to their ability to avoid excessive inference costs. However, LoRA still encounters the following challenges: (1) Limitation of low-rank assumption; and (2) Its initialization method may be suboptimal. To this end, we propose PMSS(Pre-trained Matrices Skeleton Selection), which enables high-rank updates with low costs while leveraging semantic and linguistic information inherent in pre-trained weight. It achieves this by selecting skeletons from the pre-trained weight matrix and only learning a small matrix instead. Experiments demonstrate that PMSS outperforms LoRA and other fine-tuning methods across tasks with much less trainable parameters. We demonstrate its effectiveness, especially in handling complex tasks such as DROP benchmark(+3.4%/+5.9% on LLaMA2-7B/13B) and math reasoning (+12.89%/+5.61%/+3.11% on LLaMA2-7B, Mistral-7B and Gemma-7B of GSM8K).The code and model will be released soon.
Qibin Wang, Xiaolin Hu 0001, Weikai Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
COLING4
2025 MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognition
abstract
Xinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang, Hongzhang Mu, Yubin Wang, Gaopeng Gou, Li Qian, Li Peng, Wei Liu, Jian Luan, Hongbo Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xinkui Lin, Yongxiu Xu, Hongzhang Mu, Gaopeng Gou, Wei Liu 0302, Jian Luan 0001
EMNLP10
2025 BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
abstract
Graphical User Interface (GUI) agents have gained substantial attention due to their impressive capabilities to complete tasks through multiple interactions within GUI environments.However, existing agents primarily focus on enhancing the accuracy of individual actions and often lack effective mechanisms for detecting and recovering from errors.To address these shortcomings, we propose the BacktrackAgent, a robust framework that incorporates a backtracking mechanism to improve task completion efficiency.BacktrackAgent includes verifier, judger, and reflector components as modules for error detection and recovery, while also applying judgment rewards to further enhance the agent's performance.Additionally, we develop a training dataset specifically designed for the backtracking mechanism, which considers the outcome pages after action executions.Experimental results show that BacktrackAgent has achieved performance improvements in both task success rate and step accuracy on Mobile3M and Auto-UI benchmarks.Our data and code will be released upon acceptance.
Qinzhuo Wu, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001
EMNLP3
2025 Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization
abstract
Large Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomena related to the attention mechanism during the fine-tuning of LLMs (where Wq, Wk, and Wv denote the weights of the query, key, and value layers, respectively). The first phenomenon, termed “Unequal Importance of Attention Matrices”, highlights the impact of fine-tuning different weight matrices. It shows that optimizing the Wv matrix yields significantly better performance than optimizing the Wk matrix. Fine-tuning only the Wq and Wv matrices is computationally efficient while delivering results comparable to, or even better than fine-tuning all three matrices (Wq, Wk, and Wv). The second phenomenon, “Attention Matrices with Customized Learning Rate Lead to Better Convergence”, emphasizes the importance of assigning distinct learning rates to these matrices. Specifically, a higher learning rate for the Wv matrix compared to Wq and Wk accelerates convergence and improves performance. Building on these insights, we propose a new strategy that improves fine-tuning efficiency in terms of both storage and time. Experimental results on benchmark datasets validate the effectiveness of this approach, supporting our theoretical findings. Our analysis lays the theoretical groundwork for configuring and improving algorithms in LLMs fine-tuning.
Xinhao Yao, Hongjin Qian, Xiaolin Hu 0001, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
IJCAI5
2025 MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
abstract
Mobile phone agents can assist people in automating daily tasks on their phones, which have emerged as a pivotal research spotlight. However, existing procedure-oriented agents struggle with cross-app instructions, due to the following challenges: (1) complex task relationships, (2) diverse app environment, and (3) error propagation and information loss in multi-step execution. Drawing inspiration from object-oriented programming principles, we recognize that object-oriented solutions is more suitable for cross-app instruction. To address these challenges, we propose a self-evolving multi-agent framework named MobileSteward which integrates multiple app-oriented StaffAgents coordinated by a centralized StewardAgent. We design three specialized modules in MobileSteward: (1) Dynamic Recruitment generates a scheduling graph guided by information flow to explicitly associate tasks among apps. (2) Assigned Execution assigns the task to app-oriented StaffAgents, each equipped with app-specialized expertise to address the diversity between apps. (3) Adjusted Evaluation conducts evaluation to provide reflection tips or deliver key information, which alleviates error propagation and information loss during multi-step execution. To continuously improve the performance of MobileSteward, we develop a Memory-based Self-evolution mechanism, which summarizes the experience from successful execution, to improve the performance of MobileSteward. We establish the first English Cross-APP Benchmark (CAPBench) in the real-world environment to evaluate the agents' capabilities of solving complex cross-app instructions. Experimental results demonstrate that MobileSteward achieves the best performance compared to both single-agent and multi-agent frameworks, highlighting the superiority of MobileSteward in better handling user instructions with diverse complexity.
Yuxuan Liu 0009, Hongda Sun 0001, Wei Liu 0302, Jian Luan 0001, Bo Du 0001, Rui Yan 0001
KDD (1)3
2025 Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
abstract
Menglong Cui, Pengzhi Gao, Wei Liu, Jian Luan, Bin Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Menglong Cui, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
NAACL (Long Papers)3
2025 ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
abstract
Qinzhuo Wu, Wei Liu, Jian Luan, Bin Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qinzhuo Wu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
NAACL (Long Papers)2
2025 DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
Xiaolin Hu 0001, Xiang Cheng 0007, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
PAKDD (5)4
2025 Rethinking Correlation Filter Trackers for Small Unmanned Aircraft Systems
abstract
ABSTRACT To achieve spatiotemporal continuity or some sparsity for robust tracking, most current discriminative correlation filter (DCF) methods introduce new regularization terms or self‐adaption hyperparameters to restrict the trackers. However, regardless of the validity of the pseudo‐Gaussian label, previous DCF trackers generally suffer from aberrance, mismatching. In this work, we rethink the DCF tracker from the label matching and propose a label approximation DCF tracker (LACF) focusing on analyzing the commonly used Gaussian pseudo labels in the DCF. Specifically, based on the assumption that the same objects should contain a similar response between two frames, we construct a new pseudo label that combines the original pseudo‐Gaussian labels and the previous response map. On the other hand, we introduce a windowing strategy to focus the DCF model on matching crucial labels for the right position. The experimental results demonstrate that LACF significantly achieves competitive performance for real‐time CPU small unmanned aircraft tracking.
Wei Liu 0302, Xin Yun, Youfa Liu
Comput. Intell.1
2024 Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
abstract
Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Jianfeng Liu, Ang Li, Jian Luan, Bin Wang, Rui Yan, Shuo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shihan Deng, Weikai Xu, Hongda Sun 0001, Wei Liu 0302, Tao Tan 0005, Jianfeng Liu 0005, Jian Luan 0001, Bin Wang 0004, Rui Yan 0001, Shuo Shang
ACL (1)4
2024 DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to Determinacy
abstract
Hongda Sun, Weikai Xu, Wei Liu, Jian Luan, Bin Wang, Shuo Shang, Ji-Rong Wen, Rui Yan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongda Sun 0001, Weikai Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Shuo Shang, Ji-Rong Wen, Rui Yan 0001
ACL (1)3
2024 ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval
abstract
Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address this challenge, previous studies have proposed retrieving suitable tools for the LLM based on the user query. However, previously proposed methods do not consider the differences between seen and unseen tools, nor do they take the hierarchy of the tool library into account, which may lead to suboptimal performance for tool retrieval. Therefore, to address the aforementioned issues, we propose ToolRerank, an adaptive and hierarchy-aware reranking method for tool retrieval to further refine the retrieval results. Specifically, our proposed ToolRerank includes Adaptive Truncation, which truncates the retrieval results related to seen and unseen tools at different positions, and Hierarchy-Aware Reranking, which makes retrieval results more concentrated for single-tool queries and more diverse for multi-tool queries. Experimental results show that ToolRerank can improve the quality of the retrieval results, leading to better execution results generated by the LLM.
Yuanhang Zheng, Peng Li 0021, Wei Liu 0302, Yang Liu 0005, Jian Luan 0001, Bin Wang 0004
LREC/COLING3
2024 SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
abstract
While Large Language Models (LLMs) have achieved remarkable success in various fields, the efficiency of training and inference remains a major challenge. To address this issue, we propose SUBLLM, short for Subsampling-Upsampling-Bypass Large Language Model, an innovative architecture that extends the core decoder-only framework by incorporating subsampling, upsampling, and bypass modules. The subsampling modules are responsible for shortening the sequence, while the upsampling modules restore the sequence length, and the bypass modules enhance convergence. In comparison to LLaMA, the proposed SUBLLM exhibits significant enhancements in both training and inference speeds as well as memory usage, while maintaining competitive few-shot performance. During training, SUBLLM increases speeds by 26% and cuts memory by 10GB per GPU. In inference, it boosts speeds by up to 37% and reduces memory by 1GB per GPU. The training and inference speeds can be enhanced by 34% and 52% respectively when the context window is expanded to 8192. Our code is available at https://github.com/XiaoMi/subllm.
Quandong Wang, Xiaoyu Yang 0005, Ruike Zhang, Wei Liu 0302, Jian Luan 0001, Daniel Povey, Bin Wang 0004
ECAI6
2024 ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
abstract
Recently, tool-augmented LLMs have gained increasing attention.Given an instruction, toolaugmented LLMs can interact with various external tools in multiple rounds and provide a final answer.However, previous LLMs were trained on overly detailed instructions, which included API names or parameters, while real users would not explicitly mention these API details.This leads to a gap between trained LLMs and real-world scenarios.In addition, most works ignore whether the interaction process follows the instruction.To address these issues, we constructed a training dataset called MGToolBench, which contains statement and category-level instructions to better reflect realworld scenarios.In addition, we propose Tool-Planner, a two-stage reinforcement learning framework that utilizes path planning and two feedback mechanisms to enhance the LLM's task completion and instruction-following capabilities.Experimental results show that Tool-Planner significantly improves the Match Rate, Pass Rate and Win Rate by 26.8%, 20.2%, and 5.6% compared to the SOTA model.Human evaluation verifies that the multi-granularity instructions can better align with users' usage habits.Our data and code are available at https://github.com/XiaoMi/toolplanner.
Qinzhuo Wu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
EMNLP2
2023 Overview of the NLPCC 2023 Shared Task 9: User Feedback Prediction and Response Generation
Hanlin Teng, Hongda Sun 0001, Wei Liu 0302, Shuang Dong, Rui Yan 0001, Jian Luan 0001, Bin Wang 0004
NLPCC (3)3