Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hao Yu 0030

dblp:64/4832-30 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0004-1695-8499ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 64% Multi-agent systems · 25% Vision and language · 9%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
2.332024
AutoWebGLM: A Large Language Model-based Web Navigating Agent · KDD 2024
AgentBench: Evaluating LLMs as Agents · ICLR 2024
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments · EMNLP 2024
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents
0.912025
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents · ACL (1) 2025
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
0.912025
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM agent evaluation
0.812024
AgentBench: Evaluating LLMs as Agents · ICLR 2024
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
tool-augmented agents
0.812024
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments · EMNLP 2024
Natural language and speech › Language models and text generation › LLM agents
web navigation
0.812024
AutoWebGLM: A Large Language Model-based Web Navigating Agent · KDD 2024
Human-AI interaction › LLM-based agents
LLM-based web agents
0.812024
AutoWebGLM: A Large Language Model-based Web Navigating Agent · KDD 2024
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.712023
WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences · KDD 2023
Information retrieval
retrieval models
0.712023
WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences · KDD 2023
Natural language and speech › Language models and text generation
large language model
0.312025
WebGLM: Towards an Efficient and Reliable Web-Enhanced Question-Answering System · ACM Trans. Inf. Syst. 2025
Human-AI interaction
GUI agent
0.312025
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · ICLR 2025
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.212024
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments · EMNLP 2024
Natural language and speech › Language models and text generation › large language model
large language model augmentation
0.212023
WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences · KDD 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.4program-based solver · 1.7human demonstration · 1.7agent bootstrapping · 1.7self-check · 0.9large language model · 0.9in-context learning · 0.9DPO training · 0.9rejection sampling · 0.8middleware · 0.8curriculum training · 0.8HTML simplification · 0.8GPT-4 · 0.8human preference learning · 0.7bootstrapped generation · 0.7
YearPublicationVenuePosition
2025 AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
abstract
Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, Yuxiao Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yifan Xu 0014, Xiao Liu 0036, Xueqiao Sun, Hao Yu 0030, Hanyu Lai, Shudan Zhang, Jie Tang 0001, Yuxiao Dong
ACL (1)5
2025 VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
abstract
Large Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}.
Xiao Liu 0036, Tianjie Zhang, Yu Gu 0016, Iat Long Iong, Xixuan Song, Yifan Xu 0014, Shudan Zhang, Hanyu Lai, Jiadai Sun, Zehan Qi, Shuntian Yao, Xueqiao Sun, Qinkai Zheng, Hao Yu 0030, Hanchen Zhang, Wenyi Hong, Ming Ding 0004, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su 0001, Yuxiao Dong, Jie Tang 0001
ICLR17
2025 WebGLM: Towards an Efficient and Reliable Web-Enhanced Question-Answering System
abstract
We present WebGLM, an enhanced Large Language Model (LLM)-based retrieval question-answering system based on the ChatGLM3-6B, offering significant improvements over previous systems. We aim to augment a pre-trained LLM with web search and reliable retrieval capabilities while being efficient for real-world deployments. Leveraging LLM’s in-context learning ability and a robust filter strategy, we create a high-quality training dataset and address the hallucination issue with a self-check mechanism. Our base model, ChatGLM3-6B, excels in extracting critical information and generating desired responses. We tackle the decline in retrieval effectiveness for complex queries with a keywording technique and incorporate more web content for references. We align with user preferences by training a human preference-aware scorer and employing DPO training for direct alignment. Extensive experiments, including human evaluations and the Turing test, demonstrate WebGLM’s superior performance against leading web-enhanced question-answering systems, significantly enhancing performance and efficiency. The code, demo, and data are at https://github.com/THUDM/WebGLM .
Hanyu Lai, Xiao Liu 0036, Hao Yu 0030, Yifan Xu 0014, Iat Long Iong, Shuntian Yao, Aohan Zeng, Zhengxiao Du, Yuxiao Dong, Jie Tang 0001
ACM Trans. Inf. Syst.3
2024 Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
abstract
The applications of large language models (LLMs) have expanded well beyond the confines of text processing, signaling a new era where LLMs are envisioned as generalist agents capable of operating within complex environments.These environments are often highly expansive, making it impossible for the LLM to process them within its short-term memory.Motivated by recent research on extending the capabilities of LLMs with tools, we seek to investigate the intriguing potential of tools to augment LLMs in handling such complexity by introducing a novel class of tools, termed middleware, to aid in the proactive exploration within these massive environments.Such specialized tools can serve as a middleware layer shielding the LLM from environmental complexity.In two representative complex environmentsknowledge bases (KBs) and databases-we demonstrate the significant potential of augmenting language agents with tools in complex environments.Notably, equipped with the middleware, GPT-4 achieves 2.8× the performance of the best baseline in tasks requiring access to database content and 2.2× in KB tasks.Our findings illuminate the path for advancing language agents in real-world applications.1
Yu Gu 0016, Yiheng Shu, Hao Yu 0030, Xiao Liu 0036, Yuxiao Dong, Jie Tang 0001, Jayanth Srinivasa, Hugo Latapie, Yu Su 0001
EMNLP3
2024 AgentBench: Evaluating LLMs as Agents
abstract
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001
ICLR2
2024 AutoWebGLM: A Large Language Model-based Web Navigating Agent
abstract
Large language models (LLMs) have fueled many intelligent web agents, but most existing ones perform far from satisfying in real-world web navigation tasks due to three factors: (1) the complexity of HTML text data (2) versatility of actions on webpages, and (3) task difficulty due to the open-domain nature of the web. In light of these challenges, we develop the open AutoWebGLM based on ChatGLM3-6B. AutoWebGLM can serve as a powerful automated web navigation agent that outperform GPT-4. Inspired by human browsing patterns, we first design an HTML simplification algorithm to represent webpages with vital information preserved succinctly. We then employ a hybrid human-AI method to build web browsing data for curriculum training. Finally, we bootstrap the model by reinforcement learning and rejection sampling to further facilitate webpage comprehension, browser operations, and efficient task decomposition by itself. For comprehensive evaluation, we establish a bilingual benchmark---AutoWebBench---for real-world web navigation tasks. We evaluate AutoWebGLM across diverse web navigation benchmarks, demonstrating its potential to tackle challenging tasks in real environments. Related code, model, and data are released at https://github.com/THUDM/AutoWebGLM.
Hanyu Lai, Xiao Liu 0036, Iat Long Iong, Shuntian Yao, Pengbo Shen, Hao Yu 0030, Hanchen Zhang, Yuxiao Dong, Jie Tang 0001
KDD7
2023 WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences
abstract
We present WebGLM, a web-enhanced question-answering system based on the General Language Model (GLM). Its goal is to augment a pre-trained large language model (LLM) with web search and retrieval capabilities while being efficient for real-world deployments. To achieve this, we develop WebGLM with strategies for the LLM-augmented retriever, bootstrapped generator, and human preference-aware scorer. Specifically, we identify and address the limitations of WebGPT (OpenAI), through which WebGLM is enabled with accuracy, efficiency, and cost-effectiveness advantages. In addition, we propose systematic criteria for evaluating web-enhanced QA systems. We conduct multi-dimensional human evaluation and quantitative ablation studies, which suggest the outperformance of the proposed WebGLM designs over existing systems. WebGLM with the 10-billion-parameter GLM (10B) is shown to perform better than the similar-sized WebGPT (13B) and even comparably to WebGPT (175B) in human evaluation. The code, demo, and data are at https://github.com/THUDM/WebGLM.
Xiao Liu 0036, Hanyu Lai, Hao Yu 0030, Yifan Xu 0014, Aohan Zeng, Zhengxiao Du, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001
KDD3