Xiang Deng 0001

dblp:95/4545-1 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0002-9214-7151ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Dual-View Visual Contextualization for Web Navigation
abstract
Automatic web navigation aims to build a web agent that can follow language instructions to execute complex and diverse tasks on real-world websites. Existing work primarily takes HTML documents as input, which define the contents and action spaces (i.e., actionable elements and operations) of webpages. Nevertheless, HTML documents may not provide a clear task-related context for each element, making it hard to select the right (sequence of) actions. In this paper, we propose to contextualize HTML elements through their “dual views” in webpage screenshots: each HTML element has its corresponding bounding box and visual content in the screenshot. We build upon the insight-web developers tend to arrange task-related elements nearby on webpages to enhance user experiences-and propose to contextualize each element with its neighbor elements, using both tex-tual and visual features. The resulting representations of HTML elements are more informative for the agent to take action. We validate our method on the recently released Mind2Web dataset, which features diverse navigation domains and tasks on real-world websites. Our method consistently outperforms the baseline in all the scenarios, in-cluding cross-task, cross-website, and cross-domain ones.
Jihyung Kil, Chan Hee Song, Boyuan Zheng 0001, Xiang Deng 0001, Yu Su 0001, Wei-Lun Chao
CVPR4
2024 AgentBench: Evaluating LLMs as Agents
abstract
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001
ICLR12
2023 Don't Generate, Discriminate: A Proposal for Grounding Language Models to Real-World Environments
abstract
A key missing capacity of current language models (LMs) is grounding to real-world environments.Most existing work for grounded language understanding uses LMs to directly generate plans that can be executed in the environment to achieve the desired effects.It thereby casts the burden of ensuring grammaticality, faithfulness, and controllability all on the LMs.We propose Pangu, a generic framework for grounded language understanding that capitalizes on the discriminative ability of LMs instead of their generative ability.Pangu consists of a symbolic agent and a neural LM working in a concerted fashion: The agent explores the environment to incrementally construct valid plans, and the LM evaluates the plausibility of the candidate plans to guide the search process.A case study on the challenging problem of knowledge base question answering (KBQA), which features a massive environment, demonstrates the remarkable effectiveness and flexibility of Pangu: A BERT-base LM is sufficient for setting a new record on standard KBQA datasets, and larger LMs further bring substantial gains.Pangu also enables, for the first time, effective few-shot in-context learning for KBQA with large LMs such as Codex. 1
Yu Gu 0016, Xiang Deng 0001, Yu Su 0001
ACL (1)2
2023 Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters
abstract
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, Huan Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Boshi Wang, Sewon Min, Xiang Deng 0001, You Wu 0001, Luke Zettlemoyer, Huan Sun 0001
ACL (1)3
2023 Exploring Chain of Thought Style Prompting for Text-to-SQL
abstract
In-context learning with large language models (LLMs) has recently caught increasing attention due to its superior few-shot performance on various tasks.However, its performance on text-to-SQL parsing still has much room for improvement.In this paper, we hypothesize that a crucial aspect of LLMs to improve for text-to-SQL parsing is their multi-step reasoning ability.Thus, we systematically study how to enhance LLMs' reasoning ability through chain of thought (CoT) style prompting, including the original chain-of-thought prompting (Wei et al., 2022b) and least-to-most prompting (Zhou et al., 2023).Our experiments demonstrate that iterative prompting as in Zhou et al. (2023) may be unnecessary for text-to-SQL parsing, and using detailed reasoning steps tends to have more error propagation issues.Based on these findings, we propose a new CoT-style prompting method for text-to-SQL parsing.It brings 5.2 and 6.5 point absolute gains on the Spider development set and the Spider Realistic set, respectively, compared to the standard prompting method without reasoning steps; 2.4 and 1.5 point absolute gains, compared to the least-to-most prompting method 1 .
Chang-Yu Tai, Ziru Chen, Tianshu Zhang 0001, Xiang Deng 0001, Huan Sun 0001
EMNLP4
2023 Mind2Web: Towards a Generalist Agent for the Web
abstract
We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either use simulated websites or only cover a limited set of websites and tasks, thus not suitable for generalist web agents. With over 2,000 open-ended tasks collected from 137 websites spanning 31 domains and crowdsourced action sequences for the tasks, Mind2Web provides three necessary ingredients for building generalist web agents: 1) diverse domains, websites, and tasks, 2) use of real-world websites instead of simulated and simplified ones, and 3) a broad spectrum of user interaction patterns. Based on Mind2Web, we conduct an initial exploration of using large language models (LLMs) for building generalist web agents. While the raw HTML of real-world websites are often too large to be fed to LLMs, we show that first filtering it with a small LM significantly improves the effectiveness and efficiency of LLMs. Our solution demonstrates a decent level of performance, even on websites or entire domains the model has never seen before, but there is still a substantial room to improve towards truly generalizable agents. We open-source our dataset, model implementation, and trained models (https://osu-nlp-group.github.io/Mind2Web) to facilitate further research on building a generalist agent for the web.
Xiang Deng 0001, Yu Gu 0016, Boyuan Zheng 0001, Samual Stevens, Boshi Wang, Huan Sun 0001, Yu Su 0001
NeurIPS1
2023 Roll Up Your Sleeves: Working with a Collaborative and Engaging Task-Oriented Dialogue System
abstract
Lingbo Mo, Shijie Chen, Ziru Chen, Xiang Deng, Ashley Lewis, Sunit Singh, Samuel Stevens, Chang-You Tai, Zhen Wang, Xiang Yue, Tianshu Zhang, Yu Su, Huan Sun. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Lingbo Mo, Ziru Chen, Xiang Deng 0001, Ashley Lewis, Sunit Singh, Samuel Stevens 0001, Chang-You Tai, Zhen Wang 0041, Xiang Yue, Tianshu Zhang 0001, Yu Su 0001, Huan Sun 0001
SIGDIAL4
2022 Iteratively Prompt Pre-trained Language Models for Chain of Thought
abstract
While Pre-trained Language Models (PLMs) internalize a great amount of world knowledge, they have been shown incapable of recalling these knowledge to solve tasks requiring complex & multi-step reasoning.Similar to how humans develop a "chain of thought" for these tasks, how can we equip PLMs with such abilities?In this work, we explore an iterative prompting framework, a new prompting paradigm which progressively elicits relevant knowledge from PLMs for multi-step inference.We identify key limitations of existing prompting methods, namely they are either restricted to queries with a single identifiable relation/predicate, or being agnostic to input contexts, which makes it difficult to capture variabilities across different inference steps.We propose an iterative context-aware prompter, which addresses these limitations by learning to dynamically synthesize prompts conditioned on the current step's contexts.Experiments on three datasets involving multi-step reasoning show the effectiveness of the iterative scheme and the context-aware prompter design. 1
Boshi Wang, Xiang Deng 0001, Huan Sun 0001
EMNLP2
2021 ReasonBERT: Pre-trained to Reason with Distant Supervision
abstract
We present ReasonBERT, a pre-training method that augments language models with the ability to reason over long-range relations and multiple, possibly hybrid, contexts.Unlike existing pre-training methods that only harvest learning signals from local contexts of naturally occurring texts, we propose a generalized notion of distant supervision to automatically connect multiple pieces of text and tables to create pre-training examples that require long-range reasoning.Different types of reasoning are simulated, including intersecting multiple pieces of evidence, bridging from one piece of evidence to another, and detecting unanswerable cases.We conduct a comprehensive evaluation on a variety of extractive question answering datasets ranging from single-hop to multi-hop and from text-only to table-only to hybrid that require various reasoning capabilities and show that ReasonBERT achieves remarkable improvement over an array of strong baselines.Fewshot experiments further demonstrate that our pre-training method substantially improves sample efficiency. 1
Xiang Deng 0001, Yu Su 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001, Huan Sun 0001
EMNLP (1)1
2021 Structure-Grounded Pretraining for Text-to-SQL
abstract
Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, Matthew Richardson. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xiang Deng 0001, Ahmed Awadallah 0001, Christopher Meek, Oleksandr Polozov, Huan Sun 0001, Matthew Richardson
NAACL-HLT1
2020 TURL: Table Understanding through Representation Learning
abstract
Relational tables on the Web store a vast amount of knowledge. Owing to the wealth of such tables, there has been tremendous progress on a variety of tasks in the area of table understanding. However, existing work generally relies on heavily-engineered task-specific features and model architectures. In this paper, we present TURL, a novel framework that introduces the pre-training/fine-tuning paradigm to relational Web tables. During pre-training, our framework learns deep contextualized representations on relational tables in an unsupervised manner. Its universal model design with pre-trained representations can be applied to a wide range of tasks with minimal task-specific fine-tuning. Specifically, we propose a structure-aware Transformer encoder to model the row-column structure of relational tables, and present a new Masked Entity Recovery (MER) objective for pre-training to capture the semantics and knowledge in large-scale unlabeled data. We systematically evaluate TURL with a benchmark consisting of 6 different tasks for table understanding (e.g., relation extraction, cell filling). We show that TURL generalizes well to all tasks and substantially outperforms existing methods in almost all instances.
Xiang Deng 0001, Huan Sun 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001
Proc. VLDB Endow.1
2019 Leveraging 2-hop Distant Supervision from Table Entity Pairs for Relation Extraction
abstract
Xiang Deng, Huan Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiang Deng 0001, Huan Sun 0001
EMNLP/IJCNLP (1)1