EDBT 2026 Demo / reviewers in the wild / expert
Tao Yu 0009
dblp:67/1014-9
· DBLP profile ↗
38ranked-venue papers
8as first author
30since 2021 · last 2025
0000-0001-9939-2216ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 8 first-author · 29 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attacking Vision-Language Computer Agents via Pop-upsabstractAutonomous agents powered by large vision and language models (VLM) have demonstrated significant potential in completing daily computer tasks, such as browsing the web to book travel and operating desktop software, which requires agents to understand these interfaces. Despite such visual inputs becoming more integrated into agentic applications, what types of risks and attacks exist around them still remain unclear. In this work, we demonstrate that VLM agents can be easily attacked by a set of carefully designed adversarial pop-ups, which human users would typically recognize and ignore. This distraction leads agents to click these pop-ups instead of performing their tasks as usual. Integrating these pop-ups into existing agent testing environments like OSWorld and VisualWebArena leads to an attack success rate (the frequency of the agent clicking the pop-ups) of 86% on average and decreases the task success rate by 47%. Basic defense techniques, such as asking the agent to ignore pop-ups or including an advertisement notice, are ineffective against the attack. Code is available at this link. Tao Yu 0009, Diyi Yang |
ACL (1) | 2 |
| 2025 | Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsabstractReal-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics.
We introduce Spider 2.0, an evaluation framework comprising $632$ real-world text-to-SQL workflow problems derived from enterprise-level database use cases.
The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake.
We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases.
This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding $100$ lines, which goes far beyond traditional text-to-SQL challenges.
Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3\% of the tasks, compared with 91.2\% on Spider 1.0 and 73.0\% on BIRD.
Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation --- especially in prior text-to-SQL benchmarks --- they require significant improvement in order to achieve adequate performance for real-world enterprise usage.
Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings.
Our code, baseline models, and data are available at [spider2-sql.github.io](spider2-sql.github.io) . Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Victor Zhong, Caiming Xiong, Ruoxi Sun 0002, Qian Liu 0033, Sida I. Wang, Tao Yu 0009 |
ICLR | 16 |
| 2025 | Generative Representational Instruction TuningabstractAll text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM-7B is among the top models on the Massive Text Embedding Benchmark (MTEB) and outperforms various models up to its size on a range of generative tasks. By scaling up further, GritLM-8x7B achieves even stronger generative performance while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm. Niklas Muennighoff, Hongjin Su, Liang Wang 0046, Nan Yang 0002, Furu Wei, Tao Yu 0009, Amanpreet Singh, Douwe Kiela |
ICLR | 6 |
| 2025 | Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic EnvironmentsabstractAutonomous agents powered by large language models (LLMs) have the potential to enhance human capabilities, assisting with digital tasks from sending emails to performing data analysis. The abilities of existing LLMs at such tasks are often hindered by the lack of high-quality agent data from the corresponding environments they interact with. We propose LEARN-BY-INTERACT, a data-centric framework to adapt LLM agents to any given environments without human annotations. LEARN-BY-INTERACT synthesizes trajectories of agent-environment interactions based on documentations, and constructs instructions by summarizing or abstracting the interaction histories, a process called backward construction. We assess the quality of our synthetic data by using them in both training-based scenarios and training-free in-context learning (ICL), where we craft innovative retrieval approaches optimized for agents. Extensive experiments on SWE-bench, WebArena, OSWorld, and Spider2-V spanning across realistic coding, web, and desktop environments show the effectiveness of LEARN-BY-INTERACT in various downstream agentic tasks — baseline results are improved up to 11.1% for ICL with Claude-3.5 and 23.1% for training with Codestral-22B. We further demonstrate the critical role of backward construction, which provides up to 10.6% improvement for training. Our ablation studies demonstrate the efficiency provided by our synthesized data in ICL and the superiority of our retrieval pipeline over alternative approaches like conventional retrieval-augmented generation (RAG). We expect that LEARN-BY-INTERACT will serve as a foundation for agent data synthesis as LLMs are increasingly deployed at real-world environments. Hongjin Su, Ruoxi Sun 0002, Jinsung Yoon, Tao Yu 0009, Sercan Ö. Arik |
ICLR | 5 |
| 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalabstractExisting retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings. Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009 |
ICLR | 15 |
| 2025 | AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsabstractGraphical User Interface (GUI) agents hold great potential for automating complex tasks across diverse digital environments, from web applications to desktop software. However, the development of such agents is hindered by the lack of high-quality, multi-step trajectory data required for effective training. Existing approaches rely on expensive and labor-intensive human annotation, making them unsustainable at scale. To address this challenge, we propose AgentTrek, a scalable data synthesis pipeline that generates high-quality web agent trajectories by leveraging web tutorials. Our method automatically gathers tutorial-like texts from the internet, transforms them into task goals with step-by-step instructions, and employs a visual-language model (VLM) agent to simulate their execution in a real digital environment. A VLM-based evaluator ensures the correctness of the generated trajectories. We demonstrate that training GUI agents with these synthesized trajectories significantly improves their grounding and planning performance over the current models. Moreover, our approach is more cost-efficient compared to traditional human annotation methods. This work underscores the potential of guided replay with web tutorials as a viable strategy for large-scale GUI agent training, paving the way for more capable and autonomous digital agents. Yiheng Xu, Dunjie Lu, Zhennan Shen, Caiming Xiong, Tao Yu 0009 |
ICLR | 8 |
| 2025 | Aguvis: Unified Pure Vision Agents for Autonomous GUI InteractionabstractAutomating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via inner monologue. To enable this, we construct Aguvis data collection, a large-scale dataset with multimodal grounding and reasoning annotations, and develop a two-stage training pipeline that separates GUI grounding from planning and reasoning. Experiments show that Aguvis achieves state-of-the-art performance across offline and real-world online benchmarks, marking the first fully autonomous vision-based GUI agent that operates without closed-source models. We open-source all datasets, models, and training recipes at https://aguvis-project.github.io to advance future research. Yiheng Xu, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu 0009, Caiming Xiong |
ICML | 8 |
| 2025 | Agentic AI for Enterprise: Emerging Applications and Real-world ChallengesabstractLarge language models (LLMs) have revolutionized natural language processing, enabling unprecedented capabilities in reasoning, planning, and tool utilization. Enterprises are increasingly adopting LLM-powered agents to automate complex workflows, from meeting summarization (e.g., Microsoft Copilot) to supply chain optimization and customer service orchestration. However, deploying agentic AI systems in enterprise settings introduces unique challenges, including decision making under uncertainty, multi-agent collaboration, security vulnerabilities, and trust gaps in mission-critical applications. This workshop aims to bridge the gap between academia and industry to explore LLM-driven agentic systems tailored for enterprise needs. We focus on three pillars: 1) emerging architectures that enable dynamic task decomposition and tool invocation; 2) domain-specific applications such as case studies in supply chain and employee productivity domain; 3) evaluation and governance such as the AAEF (Agentic Application Evaluation Framework) and security strategies. Anbang Xu, Min Du 0003, Meghana Puvvadi, Tao Yu 0009, Justin Emile Gottschlich |
KDD (2) | 5 |
| 2025 | OpenCUA: Open Foundations for Computer-Use AgentsabstractVision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state–action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld‑Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research. Xinyuan Wang 0010, Dunjie Lu, Junlin Yang, Tianbao Xie, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Xiaochuan Li 0003, Junda Chen, Boyuan Zheng 0001, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu 0006, Jixuan Chen, Yuxiao Ye, Yipu Wang, Diyi Yang, Victor Zhong, Y. Charles, Tao Yu 0009 |
NeurIPS | 34 |
| 2025 | Scaling Computer-Use Grounding via User Interface Decomposition and SynthesisabstractGraphical user interface (GUI) grounding, the ability to map natural language instructions to specific actions on graphical user interfaces, remains a critical bottleneck in computer use agent development. Current benchmarks oversimplify grounding tasks as short referring expressions, failing to capture the complexity of real-world interactions that require software commonsense, layout understanding, and fine-grained manipulation capabilities. To address these limitations, we introduce OSWorld-G, a comprehensive benchmark comprising 564 finely annotated samples across diverse task types including text matching, element recognition, layout understanding, and precise manipulation. Additionally, we synthesize and release the largest computer use grounding dataset Jedi, which contains 4 million examples through multi-perspective decoupling of tasks. Our multi-scale models trained on Jedi demonstrate its effectiveness by outperforming existing approaches on ScreenSpot-v2, ScreenSpot-Pro, and our OSWorld-G. Furthermore, we demonstrate that improved grounding with Jedi directly enhances agentic capabilities of general foundation models on complex computer tasks with state-of-the-art performance, improving from 23% to 51% on OSWorld. Through detailed ablation studies, we identify key factors contributing to grounding performance and verify that combining specialized data for different interface elements enables compositional generalization to novel interfaces. All benchmark, data, checkpoints, and code are open-sourced and available at https://osworld-grounding.github.io. Tianbao Xie, Xiaochuan Li 0003, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang 0010, Yiheng Xu, Doyen Sahoo, Tao Yu 0009, Caiming Xiong |
NeurIPS | 14 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 23 |
| 2024 | Text2Reward: Reward Shaping with Language Models for Reinforcement LearningabstractDesigning reward functions is a longstanding challenge in reinforcement learning (RL); it requires specialized knowledge or domain data, leading to high costs for development. To address this, we introduce Text2Reward, a data-free framework that automates the generation and shaping of dense reward functions based on large language models (LLMs). Given a goal described in natural language, Text2Reward generates shaped dense reward functions as an executable program grounded in a compact representation of the environment. Unlike inverse RL and recent work that uses LLMs to write sparse reward codes or unshaped dense rewards with a constant function across timesteps, Text2Reward produces interpretable, free-form dense reward codes that cover a wide range of tasks, utilize existing packages, and allow iterative refinement with human feedback. We evaluate Text2Reward on two robotic manipulation benchmarks (ManiSkill2, MetaWorld) and two locomotion environments of MuJoCo. On 13 of the 17 manipulation tasks, policies trained with generated reward codes achieve similar or better task success rates and convergence speed than expert-written reward codes. For locomotion tasks, our method learns six novel locomotion behaviors with a success rate exceeding 94%. Furthermore, we show that the policies trained in the simulator with our method can be deployed in the real world. Finally, Text2Reward further improves the policies by refining their reward functions with human feedback. Video results are available at https://text-to-reward.github.io/ Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Victor Zhong, Yanchao Yang 0001, Tao Yu 0009 |
ICLR | 8 |
| 2024 | Lemur: Harmonizing Natural Language and Code for Language AgentsabstractWe introduce Lemur and Lemur-Chat, openly accessible language models optimized
for both natural language and coding capabilities to serve as the backbone
of versatile language agents. The evolution from language chat models to
functional language agents demands that models not only master human interaction,
reasoning, and planning but also ensure grounding in the relevant environments.
This calls for a harmonious blend of language and coding capabilities
in the models. Lemur and Lemur-Chat are proposed to address this necessity,
demonstrating balanced proficiencies in both domains, unlike existing
open-source models that tend to specialize in either. Through meticulous pretraining
using a code-intensive corpus and instruction fine-tuning on text and code
data, our models achieve state-of-the-art averaged performance across diverse
text and coding benchmarks. Comprehensive experiments demonstrate Lemur’s
superiority over existing open-source models and its proficiency across various
agent tasks involving human communication, tool usage, and interaction under
fully- and partially- observable environments. The harmonization between natural
and programming languages enables Lemur-Chat to significantly narrow the
gap with proprietary models on agent abilities, providing key insights into developing
advanced open-source agents adept at reasoning, planning, and operating
seamlessly across environments. Our model and code have been open-sourced at
https://github.com/OpenLemur/Lemur. Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu 0033, Binyuan Hui, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu 0009 |
ICLR | 16 |
| 2024 | Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?abstractData science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io. Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009 |
NeurIPS | 23 |
| 2024 | OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsabstractAutonomous agents that accomplish complex computer tasks with minimal human interventions have the potential to transform human-computer interaction, significantly enhancing accessibility and productivity. However, existing benchmarks either lack an interactive environment or are limited to environments specific to certain applications or domains, failing to reflect the diverse and complex nature of real-world computer use, thereby limiting the scope of tasks and agent scalability. To address this issue, we introduce OSWorld, the first-of-its-kind scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems such as Ubuntu, Windows, and macOS. OSWorld can serve as a unified, integrated computer environment for assessing open-ended computer tasks that involve arbitrary applications. Building upon OSWorld, we create a benchmark of 369 computer tasks involving real web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications. Each task example is derived from real-world computer use cases and includes a detailed initial state setup configuration and a custom execution-based evaluation script for reliable, reproducible evaluation. Extensive evaluation of state-of-the-art LLM/VLM-based agents on OSWorld reveals significant deficiencies in their ability to serve as computer assistants. While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success, primarily struggling with GUI grounding and operational knowledge. Comprehensive analysis using OSWorld provides valuable insights for developing multimodal generalist agents that were not possible with previous benchmarks. Our code, environment, baseline models, and data are publicly available at this https URL. Tianbao Xie, Jixuan Chen, Xiaochuan Li 0003, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, Tao Yu 0009 |
NeurIPS | 17 |
| 2023 | Generating Data for Symbolic Language with Large Language ModelsabstractWhile large language models (LLMs) bring not only performance but also complexity, recent work has started to turn LLMs into data generators rather than task inferencers, where another affordable task model is trained for efficient deployment and inference.However, such an approach has primarily been applied to natural language tasks, and has not yet been explored for symbolic language tasks with complex structured outputs (e.g., semantic parsing and code generation).In this paper, we propose SYMGEN which utilizes LLMs for generating various annotationexpensive symbolic language data.SYMGEN consists of an informative prompt to steer generation and an agreement-based verifier to improve data correctness.We conduct extensive experiments on six symbolic language tasks across various settings.Compared with the LLMs, we demonstrate the 1%-sized task model can achieve comparable or better performance, largely cutting inference and deployment costs.We also show that generated data with only a few human demonstrations can be as effective as over 10 times the amount of human-annotated data when training the task model, saving a considerable amount of annotation effort.SYMGEN takes a step toward data generation for annotation-expensive complex tasks, and we release the code at https://github.com/HKUNLP/SymGen. Jiacheng Ye, Chengzu Li, Lingpeng Kong, Tao Yu 0009 |
EMNLP | 4 |
| 2023 | Binding Language Models in Symbolic Languages
Zhoujun Cheng, Tianbao Xie, Peng Shi 0010, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir R. Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 12 |
| 2023 | Selective Annotation Makes Language Models Better Few-Shot Learners
Hongjin Su, Jungo Kasai, Chen Henry Wu, Jiayi Xin, Rui Zhang 0037, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 11 |
| 2023 | DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationabstractWe introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io. Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009 |
ICML | 10 |
| 2023 | Compositional Exemplars for In-context LearningabstractLarge pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we systematically formulate in-context example selection as a subset selection problem, and optimize it in an end-to-end fashion. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through carefully-designed contrastive learning to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, phraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation and semantic parsing. Extensive experiments demonstrate the effectiveness, transferability, compositionality of CEIL, shedding new lights on in-context leaning. Our code is released at https://github.com/HKUNLP/icl-ceil. Jiacheng Ye, Zhiyong Wu 0003, Jiangtao Feng, Tao Yu 0009, Lingpeng Kong |
ICML | 4 |
| 2023 | Coder Reviewer Reranking for Code GenerationabstractSampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions. Inspired by collaborative programming, we propose Coder-Reviewer reranking. We augment Coder language models from past work, which generate programs given language instructions, with Reviewer models, which evaluate the likelihood of the instruction given the generated programs. We perform an extensive study across six datasets with eight models from three model families. Experimental results show that Coder-Reviewer reranking leads to consistent and significant improvement (up to 17% absolute accuracy gain) over reranking with the Coder model only. When combined with executability filtering, Coder-Reviewer reranking can often outperform the minimum Bayes risk method. Coder-Reviewer reranking is easy to implement by prompting, can generalize to different programming languages, and works well with off-the-shelf hyperparameters. Tao Yu 0009, Tatsunori B. Hashimoto, Mike Lewis, Scott Yih, Daniel Fried, Sida I. Wang |
ICML | 2 |
| 2023 | Automated Self-Supervised Learning for RecommendationabstractGraph neural networks (GNNs) have emerged as the state-of-the-art paradigm for collaborative filtering (CF). To improve the representation quality over limited labeled data, contrastive learning has attracted attention in recommendation and benefited graph-based CF model recently. However, the success of most contrastive methods heavily relies on manually generating effective contrastive views for heuristic-based data augmentation. This does not generalize across different datasets and downstream recommendation tasks, which is difficult to be adaptive for data augmentation and robust to noise perturbation. To fill this crucial gap, this work proposes a unified Automated Collaborative Filtering (AutoCF) to automatically perform data augmentation for recommendation. Specifically, we focus on the generative self-supervised learning framework with a learnable augmentation paradigm that benefits the automated distillation of important self-supervised signals. To enhance the representation discrimination ability, our masked graph autoencoder is designed to aggregate global information during the augmentation via reconstructing the masked subgraph structures. Experiments and ablation studies are performed on several public datasets for recommending products, venues, and locations. Results demonstrate the superiority of AutoCF against various baseline methods. We release the model implementation at https://github.com/HKUDS/AutoCF. Lianghao Xia, Chao Huang 0001, Chunzhen Huang, Kangyi Lin, Tao Yu 0009, Ben Kao |
WWW | 5 |
| 2022 | DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationabstractZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Awadallah, Dragomir Radev. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 0001, Rui Zhang 0037, Tao Yu 0009, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev |
ACL (1) | 6 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 23 |
| 2022 | ZeroGen: Efficient Zero-shot Learning via Dataset GenerationabstractThere is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs).In this paper, we study a flexible and efficient zero-short learning method, ZEROGEN.Given a zero-shot task, we first generate a dataset from scratch using PLMs in an unsupervised manner.Then, we train a tiny task model (e.g., LSTM) under the supervision of the synthesized dataset.This approach allows highly efficient inference as the final task model only has orders of magnitude fewer parameters comparing to PLMs (e.g., GPT2-XL).Apart from being annotation-free and efficient, we argue that ZEROGEN can also provide useful insights from the perspective of datafree model-agnostic knowledge distillation, and unreferenced text generation evaluation.Experiments and analysis on different NLP tasks, namely, text classification, question answering, and natural language inference, show the effectiveness of ZEROGEN. Jiacheng Ye, Jiahui Gao 0002, Qintong Li, Hang Xu 0004, Jiangtao Feng, Zhiyong Wu 0003, Tao Yu 0009, Lingpeng Kong |
EMNLP | 7 |
| 2021 | Effective Fine-Tuning Methods for Cross-lingual AdaptationabstractLarge scale multilingual pre-trained language models have shown promising results in zeroand few-shot cross-lingual tasks.However, recent studies have shown their lack of generalizability when the languages are structurally dissimilar.In this work, we propose a novel fine-tuning method based on co-training that aims to learn more generalized semantic equivalences as complementary to multilingual language modeling using the unlabeled data in the target language.We also propose an adaption method based on contrastive learning to better capture the semantic relationship in the parallel data, when a few translation pairs are available.To show our method's effectiveness, we conduct extensive experiments on cross-lingual inference and review classification tasks across various languages.We report significant gains compared to directly finetuning multilingual pre-trained models and other semi-supervised alternatives.1 Tao Yu 0009, Shafiq R. Joty |
EMNLP (1) | 1 |
| 2021 | GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
Tao Yu 0009, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang 0002, Dragomir R. Radev, Richard Socher, Caiming Xiong |
ICLR | 1 |
| 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Tao Yu 0009, Rui Zhang 0037, Oleksandr Polozov, Christopher Meek, Ahmed Awadallah 0001 |
ICLR | 1 |
| 2021 | DART: Open-Domain Structured Data Record to Text GenerationabstractLinyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Linyong Nan, Dragomir R. Radev, Rui Zhang 0037, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma 0001, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta 0015, Tao Yu 0009, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani |
NAACL-HLT | 19 |
| 2021 | QMSum: A New Benchmark for Query-based Multi-domain Meeting SummarizationabstractMing Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ming Zhong 0005, Da Yin, Tao Yu 0009, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Awadallah 0001, Asli Celikyilmaz, Yang Liu 0124, Xipeng Qiu, Dragomir R. Radev |
NAACL-HLT | 3 |
| 2020 | Online Conversation Disentanglement with Pointer NetworksabstractHuge amounts of textual conversations occur online every day, where multiple conversations take place concurrently.Interleaved conversations lead to difficulties in not only following the ongoing discussions but also extracting relevant information from simultaneous messages.Conversation disentanglement aims to separate intermingled messages into detached conversations.However, existing disentanglement methods rely mostly on handcrafted features that are dataset specific, which hinders generalization and adaptability.In this work, we propose an end-to-end online framework for conversation disentanglement that avoids time-consuming domain-specific feature engineering.We design a novel way to embed the whole utterance that comprises timestamp, speaker, and message text, and propose a custom attention mechanism that models disentanglement as a pointing problem while effectively capturing inter-utterance interactions in an end-to-end fashion.We also introduce a joint-learning objective to better capture contextual information.Our experiments on the Ubuntu IRC dataset show that our method achieves state-of-the-art performance in both link and conversation prediction tasks. Tao Yu 0009, Shafiq R. Joty |
EMNLP (1) | 1 |
| 2020 | Semantic Evaluation for Text-to-SQL with Distilled Test SuitesabstractWe propose test suite accuracy to approximate semantic accuracy for Text-to-SQL models.Our method distills a small test suite of databases that achieves high code coverage for the gold query from a large number of randomly generated databases.At evaluation time, it computes the denotation accuracy of the predicted queries on the distilled test suite, hence calculating a tight upper-bound for semantic accuracy efficiently.We use our proposed method to evaluate 21 models submitted to the Spider leader board and manually verify that our method is always correct on 100 examples.In contrast, the current Spider metric leads to a 2.5% false negative rate on average and 8.1% in the worst case, indicating that test suite accuracy is needed.Our implementation, along with distilled test suites for eleven Textto-SQL datasets, is publicly available. Ruiqi Zhong, Tao Yu 0009, Daniel Klein 0001 |
EMNLP (1) | 2 |
| 2019 | SParC: Cross-Domain Semantic Parsing in ContextabstractTao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev |
ACL (1) | 1 |
| 2019 | CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to DatabasesabstractTao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Editing-Based SQL Query Generation for Cross-Domain Context-Dependent QuestionsabstractRui Zhang, Tao Yu, Heyang Er, Sungrok Shim, Eric Xue, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Rui Zhang 0037, Tao Yu 0009, Heyang Er, Sungrok Shim, Eric Xue 0001, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir R. Radev |
EMNLP/IJCNLP (1) | 2 |
| 2018 | SyntaxSQLNet: Syntax Tree Networks for Complex and Cross-Domain Text-to-SQL TaskabstractMost existing studies in text-to-SQL tasks do not require generating complex SQL queries with multiple clauses or sub-queries, and generalizing to new, unseen databases.In this paper we propose SyntaxSQLNet, a syntax tree network to address the complex and crossdomain text-to-SQL generation task.Syn-taxSQLNet employs a SQL specific syntax tree-based decoder with SQL generation path history and table-aware column attention encoders.We evaluate SyntaxSQLNet on a new large-scale text-to-SQL corpus containing databases with multiple tables and complex SQL queries containing multiple SQL clauses and nested queries.We use a database split setting where databases in the test set are unseen during training.Experimental results show that SyntaxSQLNet can handle a significantly greater number of complex SQL examples than prior work, outperforming the previous state-of-the-art model by 9.5% in exact matching accuracy.To our knowledge, we are the first to study this complex text-to-SQL task.Our task and models with the latest updates are available at https://yale-lily. github.io/seq2sql/spider. Tao Yu 0009, Michihiro Yasunaga, Rui Zhang 0037, Zifan Li, Dragomir R. Radev |
EMNLP | 1 |
| 2018 | Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL TaskabstractTao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, Dragomir Radev. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Dragomir R. Radev |
EMNLP | 1 |
| 2018 | Cross-lingual sentiment transfer with limited resources
Mohammad Sadegh Rasooli, Noura Farra, Axinia Radeva, Tao Yu 0009, Kathy McKeown |
Mach. Transl. | 4 |