Hanyu Lai

dblp:330/5182 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-3106-320XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 AndroidGen: Building an Android Language Agent under Data Scarcity
abstract
Large language models have opened up a world of possibilities for various NLP tasks, sparking optimism for the future.Despite their potential, LLMs have yet to be widely used as agents on real mobile devices.The main challenge is the need for high-quality data sources.Time constraints and labor intensity often hinder human annotation.On the other hand, existing LLMs exhibit inadequate completion rates and need a robust data filtration strategy.Given these challenges, we develop a framework called ANDROIDGEN to enhance the capabilities of LLM-based agents under data scarcity.In addition, we leverage AN-DROIDGEN to collect trajectories given human tasks and train open-source LLMs on these trajectories to develop an open-source mobile agent without manually labeled trajectories.We extensively evaluate ANDROIDGEN with AndroidWorld, AitW, and various popular applications, demonstrating its improvements and revealing potential areas for future improvement.Code, model, and data are available at https://github.com/THUDM/AndroidGen.
Hanyu Lai, Xiao Liu 0036, Yifan Xu 0014, Shudan Zhang, Yuxiao Dong, Jie Tang 0001
ACL (1)1
2025 A Survey of Post-Training Scaling in Large Language Models
abstract
Hanyu Lai, Xiao Liu, Junjie Gao, Jiale Cheng, Zehan Qi, Yifan Xu, Shuntian Yao, Dan Zhang, Jinhua Du, Zhenyu Hou, Xin Lv, Minlie Huang, Yuxiao Dong, Jie Tang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hanyu Lai, Xiao Liu 0036, Zehan Qi, Yifan Xu 0014, Shuntian Yao, Jinhua Du, Minlie Huang, Yuxiao Dong, Jie Tang 0001
ACL (1)1
2025 AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
abstract
Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, Yuxiao Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yifan Xu 0014, Xiao Liu 0036, Xueqiao Sun, Hao Yu 0030, Hanyu Lai, Shudan Zhang, Jie Tang 0001, Yuxiao Dong
ACL (1)6
2025 VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
abstract
Large Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}.
Xiao Liu 0036, Tianjie Zhang, Yu Gu 0016, Iat Long Iong, Xixuan Song, Yifan Xu 0014, Shudan Zhang, Hanyu Lai, Jiadai Sun, Zehan Qi, Shuntian Yao, Xueqiao Sun, Qinkai Zheng, Hao Yu 0030, Hanchen Zhang, Wenyi Hong, Ming Ding 0004, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su 0001, Yuxiao Dong, Jie Tang 0001
ICLR8
2025 WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
abstract
Large language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks. However, existing LLM web agents face significant limitations: high-performing agents rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabilities. This paper introduces WebRL, a novel self-evolving online curriculum reinforcement learning framework designed to train high-performance web agents using open LLMs. Our approach addresses key challenges in this domain, including the scarcity of training tasks, sparse feedback signals, and policy distribution drift in online learning. WebRL incorporates a self-evolving curriculum that generates new tasks from unsuccessful attempts, a robust outcome-supervised reward model (ORM), and adaptive reinforcement learning strategies to ensure consistent improvement. We apply WebRL to transform Llama-3.1 models into proficient web agents, achieving remarkable results on the WebArena-Lite benchmark. Our Llama-3.1-8B agent improves from an initial 4.8\% success rate to 42.4\%, while the Llama-3.1-70B agent achieves a 47.3\% success rate across five diverse websites. These results surpass the performance of GPT-4-Turbo (17.6\%) by over 160\% relatively and significantly outperform previous state-of-the-art web agents trained on open LLMs (AutoWebGLM, 18.2\%). Our findings demonstrate WebRL's effectiveness in bridging the gap between open and proprietary LLM-based web agents, paving the way for more accessible and powerful autonomous web interaction systems.
Zehan Qi, Xiao Liu 0036, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Jiadai Sun, Shuntian Yao, Wei Xu 0017, Jie Tang 0001, Yuxiao Dong
ICLR4
2025 WebGLM: Towards an Efficient and Reliable Web-Enhanced Question-Answering System
abstract
We present WebGLM, an enhanced Large Language Model (LLM)-based retrieval question-answering system based on the ChatGLM3-6B, offering significant improvements over previous systems. We aim to augment a pre-trained LLM with web search and reliable retrieval capabilities while being efficient for real-world deployments. Leveraging LLM’s in-context learning ability and a robust filter strategy, we create a high-quality training dataset and address the hallucination issue with a self-check mechanism. Our base model, ChatGLM3-6B, excels in extracting critical information and generating desired responses. We tackle the decline in retrieval effectiveness for complex queries with a keywording technique and incorporate more web content for references. We align with user preferences by training a human preference-aware scorer and employing DPO training for direct alignment. Extensive experiments, including human evaluations and the Turing test, demonstrate WebGLM’s superior performance against leading web-enhanced question-answering systems, significantly enhancing performance and efficiency. The code, demo, and data are at https://github.com/THUDM/WebGLM .
Hanyu Lai, Xiao Liu 0036, Hao Yu 0030, Yifan Xu 0014, Iat Long Iong, Shuntian Yao, Aohan Zeng, Zhengxiao Du, Yuxiao Dong, Jie Tang 0001
ACM Trans. Inf. Syst.1
2024 AgentBench: Evaluating LLMs as Agents
abstract
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001
ICLR6
2024 AutoWebGLM: A Large Language Model-based Web Navigating Agent
abstract
Large language models (LLMs) have fueled many intelligent web agents, but most existing ones perform far from satisfying in real-world web navigation tasks due to three factors: (1) the complexity of HTML text data (2) versatility of actions on webpages, and (3) task difficulty due to the open-domain nature of the web. In light of these challenges, we develop the open AutoWebGLM based on ChatGLM3-6B. AutoWebGLM can serve as a powerful automated web navigation agent that outperform GPT-4. Inspired by human browsing patterns, we first design an HTML simplification algorithm to represent webpages with vital information preserved succinctly. We then employ a hybrid human-AI method to build web browsing data for curriculum training. Finally, we bootstrap the model by reinforcement learning and rejection sampling to further facilitate webpage comprehension, browser operations, and efficient task decomposition by itself. For comprehensive evaluation, we establish a bilingual benchmark---AutoWebBench---for real-world web navigation tasks. We evaluate AutoWebGLM across diverse web navigation benchmarks, demonstrating its potential to tackle challenging tasks in real environments. Related code, model, and data are released at https://github.com/THUDM/AutoWebGLM.
Hanyu Lai, Xiao Liu 0036, Iat Long Iong, Shuntian Yao, Pengbo Shen, Hao Yu 0030, Hanchen Zhang, Yuxiao Dong, Jie Tang 0001
KDD1
2023 GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu 0036, Zhengxiao Du, Hanyu Lai, Ming Ding 0004, Zhuoyi Yang, Yifan Xu 0014, Wendi Zheng, Weng Lam Tam, Zixuan Ma, Jidong Zhai, Zhiyuan Liu 0001, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001
ICLR5
2023 WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences
abstract
We present WebGLM, a web-enhanced question-answering system based on the General Language Model (GLM). Its goal is to augment a pre-trained large language model (LLM) with web search and retrieval capabilities while being efficient for real-world deployments. To achieve this, we develop WebGLM with strategies for the LLM-augmented retriever, bootstrapped generator, and human preference-aware scorer. Specifically, we identify and address the limitations of WebGPT (OpenAI), through which WebGLM is enabled with accuracy, efficiency, and cost-effectiveness advantages. In addition, we propose systematic criteria for evaluating web-enhanced QA systems. We conduct multi-dimensional human evaluation and quantitative ablation studies, which suggest the outperformance of the proposed WebGLM designs over existing systems. WebGLM with the 10-billion-parameter GLM (10B) is shown to perform better than the similar-sized WebGPT (13B) and even comparably to WebGPT (175B) in human evaluation. The code, demo, and data are at https://github.com/THUDM/WebGLM.
Xiao Liu 0036, Hanyu Lai, Hao Yu 0030, Yifan Xu 0014, Aohan Zeng, Zhengxiao Du, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001
KDD2