EDBT 2026 Demo / reviewers in the wild / expert
Victor Zhong
dblp:182/8931
· DBLP profile ↗
25ranked-venue papers
8as first author
14since 2021 · last 2026
0009-0001-0500-6802ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 8 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSCode Comparator: An Interactive Interface for Comparing Models and Evaluating Code for Data Science TasksabstractCode-generating models are increasingly used to support data science tasks. Yet reviewing their outputs, both to understand how the code works and to assess its quality, remains largely manual and time-consuming. Instead of eliminating effort, these models shift the burden from writing code to verifying it. Complicating matters further, different models often produce divergent solutions of varying efficacy, creating additional challenges for code interrogation. To address this unmet need, we introduce DSCode Comparator, an interactive interface designed to support code understanding, evaluation, refinement, and comparison in data science workflows. DSCode Comparator allows code to be viewed from different levels of granularity, from individual lines of code to comparisons across prompts and tasks. The individual code views automatically annotate lines of code via an agentic pipeline we developed to facilitate quick functional overviews. The individual views also automatic diagnosis code quality according to efficiency, readability, and resource computation. The comparison and historical views leverage the annotations to create compact visual summaries of code, allowing for direct comparisons of its functionality, length, and efficiency across multiple models, data science tasks, and prompts. To evaluate DSCode Comparator, we conducted a user study with 22 participants, all with varying levels of proficiency in writing Data science code. Our findings show that, especially for non-expert users, DSCode Comparator sped up the pace of code comprehension and increased participant confidence. The majority report finding DSCode Comparator easier to use and more efficient than manual efforts in reviewing and refining code. Overall, our systems and their findings present an intelligent, human-centered approach to address the verification gap when using code generation models for Data Science. Victor Zhong, Anamaria Crisan |
IUI | 2 |
| 2025 | Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsabstractReal-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics.
We introduce Spider 2.0, an evaluation framework comprising $632$ real-world text-to-SQL workflow problems derived from enterprise-level database use cases.
The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake.
We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases.
This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding $100$ lines, which goes far beyond traditional text-to-SQL challenges.
Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3\% of the tasks, compared with 91.2\% on Spider 1.0 and 73.0\% on BIRD.
Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation --- especially in prior text-to-SQL benchmarks --- they require significant improvement in order to achieve adequate performance for real-world enterprise usage.
Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings.
Our code, baseline models, and data are available at [spider2-sql.github.io](spider2-sql.github.io) . Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Victor Zhong, Caiming Xiong, Ruoxi Sun 0002, Qian Liu 0033, Sida I. Wang, Tao Yu 0009 |
ICLR | 11 |
| 2025 | OpenCUA: Open Foundations for Computer-Use AgentsabstractVision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state–action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld‑Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research. Xinyuan Wang 0010, Dunjie Lu, Junlin Yang, Tianbao Xie, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Xiaochuan Li 0003, Junda Chen, Boyuan Zheng 0001, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu 0006, Jixuan Chen, Yuxiao Ye, Yipu Wang, Diyi Yang, Victor Zhong, Y. Charles, Tao Yu 0009 |
NeurIPS | 31 |
| 2024 | Text2Reward: Reward Shaping with Language Models for Reinforcement LearningabstractDesigning reward functions is a longstanding challenge in reinforcement learning (RL); it requires specialized knowledge or domain data, leading to high costs for development. To address this, we introduce Text2Reward, a data-free framework that automates the generation and shaping of dense reward functions based on large language models (LLMs). Given a goal described in natural language, Text2Reward generates shaped dense reward functions as an executable program grounded in a compact representation of the environment. Unlike inverse RL and recent work that uses LLMs to write sparse reward codes or unshaped dense rewards with a constant function across timesteps, Text2Reward produces interpretable, free-form dense reward codes that cover a wide range of tasks, utilize existing packages, and allow iterative refinement with human feedback. We evaluate Text2Reward on two robotic manipulation benchmarks (ManiSkill2, MetaWorld) and two locomotion environments of MuJoCo. On 13 of the 17 manipulation tasks, policies trained with generated reward codes achieve similar or better task success rates and convergence speed than expert-written reward codes. For locomotion tasks, our method learns six novel locomotion behaviors with a success rate exceeding 94%. Furthermore, we show that the policies trained in the simulator with our method can be deployed in the real world. Finally, Text2Reward further improves the policies by refining their reward functions with human feedback. Video results are available at https://text-to-reward.github.io/ Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Victor Zhong, Yanchao Yang 0001, Tao Yu 0009 |
ICLR | 6 |
| 2024 | Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?abstractData science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io. Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009 |
NeurIPS | 20 |
| 2024 | OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsabstractAutonomous agents that accomplish complex computer tasks with minimal human interventions have the potential to transform human-computer interaction, significantly enhancing accessibility and productivity. However, existing benchmarks either lack an interactive environment or are limited to environments specific to certain applications or domains, failing to reflect the diverse and complex nature of real-world computer use, thereby limiting the scope of tasks and agent scalability. To address this issue, we introduce OSWorld, the first-of-its-kind scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems such as Ubuntu, Windows, and macOS. OSWorld can serve as a unified, integrated computer environment for assessing open-ended computer tasks that involve arbitrary applications. Building upon OSWorld, we create a benchmark of 369 computer tasks involving real web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications. Each task example is derived from real-world computer use cases and includes a detailed initial state setup configuration and a custom execution-based evaluation script for reliable, reproducible evaluation. Extensive evaluation of state-of-the-art LLM/VLM-based agents on OSWorld reveals significant deficiencies in their ability to serve as computer assistants. While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success, primarily struggling with GUI grounding and operational knowledge. Comprehensive analysis using OSWorld provides valuable insights for developing multimodal generalist agents that were not possible with previous benchmarks. Our code, environment, baseline models, and data are publicly available at this https URL. Tianbao Xie, Jixuan Chen, Xiaochuan Li 0003, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, Tao Yu 0009 |
NeurIPS | 16 |
| 2024 | Policy Improvement using Language Feedback ModelsabstractWe introduce Language Feedback Models (LFMs) that identify desirable behaviour --- actions that help achieve tasks specified in the instruction - for imitation learning in instruction following. To train LFMs, we obtain feedback from Large Language Models (LLMs) on visual trajectories verbalized to language descriptions. First, by using LFMs to identify desirable behaviour to imitate, we improve in task-completion rate over strong behavioural cloning baselines on three distinct language grounding environments (Touchdown, ScienceWorld, and ALFWorld). Second, LFMs outperform using LLMs as experts to directly predict actions, when controlling for the number of LLM output tokens. Third, LFMs generalize to unseen environments, improving task-completion rate by 3.5-12.0% through one round of adaptation. Finally, LFMs can be modified to provide human-interpretable feedback without performance loss, allowing human verification of desirable behaviour for imitation learning. Victor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre Côté |
NeurIPS | 1 |
| 2023 | When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesabstractAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 3 |
| 2022 | M2D2: A Massively Multi-Domain Language Modeling DatasetabstractWe present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs).M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar.Using ontologies derived from Wikipedia and ArXiv categories, we organize the domains in each data source into 22 groups.This two-level hierarchy enables the study of relationships between domains and their effects on in-and out-of-domain performance after adaptation.We also present a number of insights into the nature of effective domain adaptation in LMs, as examples of the new types of studies M2D2 enables.To improve in-domain performance, we show the benefits of adapting the LM along a domain hierarchy; adapting to smaller amounts of fine-grained domainspecific data can lead to larger in-domain performance gains than larger amounts of weakly relevant data.We further demonstrate a tradeoff between in-domain specialization and outof-domain generalization within and across ontologies, as well as a strong correlation between out-of-domain performance and lexical overlap between domains. Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyer |
EMNLP | 2 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 11 |
| 2022 | Improving Intrinsic Exploration with Language AbstractionsabstractReinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 47-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites. Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah D. Goodman, Tim Rocktäschel, Edward Grefenstette |
NeurIPS | 2 |
| 2022 | Improving Policy Learning via Language Dynamics DistillationabstractRecent work has shown that augmenting environments with language descriptions improves policy learning. However, for environments with complex language abstractions, learning how to ground language to observations is difficult due to sparse, delayed rewards. We propose Language Dynamics Distillation (LDD), which pretrains a model to predict environment dynamics given demonstrations with language descriptions, and then fine-tunes these language-aware pretrained representations via reinforcement learning (RL). In this way, the model is trained to both maximize expected reward and retain knowledge about how language relates to environment dynamics. On SILG, a benchmark of five tasks with language descriptions that evaluate distinct generalization challenges on unseen environments (NetHack, ALFWorld, RTFM, Messenger, and Touchdown), LDD outperforms tabula-rasa RL, VAE pretraining, and methods that learn from unlabeled demonstrations in inverse RL and reward shaping with pretrained experts. In our analyses, we show that language descriptions in demonstrations improve sample-efficiency and generalization across environments, and that dynamics modeling with expert demonstrations is more effective than with non-experts. Victor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette, Tim Rocktäschel |
NeurIPS | 1 |
| 2021 | Grounding Language to Entities and Dynamics for Generalization in Reinforcement LearningabstractWe investigate the use of natural language to drive the generalization of control policies and introduce the new multi-task environment Messenger with free-form text manuals describing the environment dynamics. Unlike previous work, Messenger does not assume prior knowledge connecting text and state observations {—} the control policy must simultaneously ground the game manual to entity symbols and dynamics in the environment. We develop a new model, EMMA (Entity Mapper with Multi-modal Attention) which uses an entity-conditioned attention module that allows for selective focus over relevant descriptions in the manual for each entity in the environment. EMMA is end-to-end differentiable and learns a latent grounding of entities and dynamics from text to observations using only environment rewards. EMMA achieves successful zero-shot generalization to unseen games with new dynamics, obtaining a 40% higher win rate compared to multiple baselines. However, win rate on the hardest stage of Messenger remains low (10%), demonstrating the need for additional work in this direction. Austin W. Hanjie, Victor Zhong, Karthik Narasimhan |
ICML | 2 |
| 2021 | SILG: The Multi-domain Symbolic Interactive Language Grounding BenchmarkabstractExisting work in language grounding typically study single environments. How do we build unified models that apply across multiple environments? We propose the multi-environment Symbolic Interactive Language Grounding benchmark (SILG), which unifies a collection of diverse grounded language learning environments under a common interface. SILG consists of grid-world environments that require generalization to new dynamics, entities, and partially observed worlds (RTFM, Messenger, NetHack), as well as symbolic counterparts of visual worlds that re- quire interpreting rich natural language with respect to complex scenes (ALFWorld, Touchdown). Together, these environments provide diverse grounding challenges in richness of observation space, action space, language specification, and plan com- plexity. In addition, we propose the first shared model architecture for RL on these environments, and evaluate recent advances such as egocentric local convolution, recurrent state-tracking, entity-centric attention, and pretrained LM using SILG. Our shared architecture achieves comparable performance to environment-specific architectures. Moreover, we find that many recent modelling advances do not result in significant gains on environments other than the one they were designed for. This highlights the need for a multi-environment benchmark. Finally, the best models significantly underperform humans on SILG, which suggests ample room for future work. We hope SILG enables the community to quickly identify new methodolo- gies for language grounding that generalize to a diverse set of environments and their associated challenges. Victor Zhong, Austin W. Hanjie, Sida I. Wang, Karthik Narasimhan, Luke Zettlemoyer |
NeurIPS | 1 |
| 2020 | Grounded Adaptation for Zero-shot Executable Semantic ParsingabstractWe propose Grounded Adaptation for Zeroshot Executable Semantic Parsing (GAZP) to adapt an existing semantic parser to new environments (e.g.new database schemas).GAZP combines a forward semantic parser with a backward utterance generator to synthesize data (e.g.utterances and SQL queries) in the new environment, then selects cycleconsistent examples to adapt the parser.Unlike data-augmentation, which typically synthesizes unverified examples in the training environment, GAZP synthesizes examples in the new environment whose inputoutput consistency are verified.On the Spider, Sparc, and CoSQL zero-shot semantic parsing tasks, GAZP improves logical form and execution accuracy of the baseline parser.Our analyses show that GAZP outperforms dataaugmentation in the training environment, performance increases with the amount of GAZPsynthesized data, and cycle-consistency is central to successful adaptation. Victor Zhong, Mike Lewis, Sida I. Wang, Luke Zettlemoyer |
EMNLP (1) | 1 |
| 2020 | RTFM: Generalising to New Environment Dynamics via Reading
Victor Zhong, Tim Rocktäschel, Edward Grefenstette |
ICLR | 1 |
| 2019 | Multi-hop Reading Comprehension through Question Decomposition and RescoringabstractMulti-hop Reading Comprehension (RC) requires reasoning and aggregation across several paragraphs.We propose a system for multi-hop RC that decomposes a compositional question into simpler sub-questions that can be answered by off-the-shelf single-hop RC models.Since annotations for such decomposition are expensive, we recast subquestion generation as a span prediction problem and show that our method, trained using only 400 labeled examples, generates sub-questions that are as effective as humanauthored sub-questions.We also introduce a new global rescoring approach that considers each decomposition (i.e. the sub-questions and their answers) to select the best final answer, greatly improving overall performance.Our experiments on HOTPOTQA show that this approach achieves the state-of-the-art results, while providing explainable evidence for its decision making in the form of sub-questions. Sewon Min, Victor Zhong, Luke Zettlemoyer, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2019 | E3: Entailment-driven Extracting and Editing for Conversational Machine ReadingabstractConversational machine reading systems help users answer high-level questions (e.g.determine if they qualify for particular government benefits) when they do not know the exact rules by which the determination is made (e.g.whether they need certain income levels or veteran status).The key challenge is that these rules are only provided in the form of a procedural text (e.g.guidelines from government website) which the system must read to figure out what to ask the user.We present a new conversational machine reading model that jointly extracts a set of decision rules from the procedural text while reasoning about which are entailed by the conversational history and which still need to be edited to create questions for the user.On the recently introduced ShARC conversational machine reading dataset, our Entailment-driven Extract and Edit network (E 3 ) achieves a new state-of-theart, outperforming existing systems as well as a new BERT-based baseline.In addition, by explicitly highlighting which information still needs to be gathered, E 3 provides a more explainable alternative to prior work.We release source code for our models and experiments at https://github.com/vzhong/e3. Victor Zhong, Luke Zettlemoyer |
ACL (1) | 1 |
| 2019 | Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering
Victor Zhong, Caiming Xiong, Nitish Shirish Keskar, Richard Socher |
ICLR (Poster) | 1 |
| 2018 | Global-Locally Self-Attentive Encoder for Dialogue State TrackingabstractDialogue state tracking, which estimates user goals and requests given the dialogue context, is an essential part of taskoriented dialogue systems.In this paper, we propose the Global-Locally Self-Attentive Dialogue State Tracker (GLAD), which learns representations of the user utterance and previous system actions with global-local modules.Our model uses global modules to share parameters between estimators for different types (called slots) of dialogue states, and uses local modules to learn slot-specific features.We show that this significantly improves tracking of rare states and achieves stateof-the-art performance on the WoZ and DSTC2 state tracking tasks.GLAD obtains 88.1% joint goal accuracy and 97.1% request accuracy on WoZ, outperforming prior work by 3.7% and 5.5%.On DSTC2, our model obtains 74.5% joint goal accuracy and 97.5% request accuracy, outperforming prior work by 1.1% and 1.0%. Victor Zhong, Caiming Xiong, Richard Socher |
ACL (1) | 1 |
| 2018 | Efficient and Robust Question Answering from Minimal Context over DocumentsabstractNeural models for question answering (QA) over documents have achieved significant performance improvements.Although effective, these models do not scale to large corpora due to their complex modeling of interactions between the document and the question.Moreover, recent work has shown that such models are sensitive to adversarial inputs.In this paper, we study the minimal context required to answer the question, and find that most questions in existing datasets can be answered with a small set of sentences.Inspired by this observation, we propose a simple sentence selector to select the minimal set of sentences to feed into the QA model.Our overall system achieves significant reductions in training (up to 15 times) and inference times (up to 13 times), with accuracy comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQuAD-Open.Furthermore, our experimental results and analyses show that our approach is more robust to adversarial inputs. Sewon Min, Victor Zhong, Richard Socher, Caiming Xiong |
ACL (1) | 2 |
| 2018 | DCN+: Mixed Objective And Deep Residual Coattention for Question Answering
Caiming Xiong, Victor Zhong, Richard Socher |
ICLR (Poster) | 2 |
| 2017 | Position-aware Attention and Supervised Data Improve Slot FillingabstractOrganized relational knowledge in the form of "knowledge graphs" is important for many applications.However, the ability to populate knowledge bases with facts automatically extracted from documents has improved frustratingly slowly.This paper simultaneously addresses two issues that have held back prior work.We first propose an effective new model, which combines an LSTM sequence model with a form of entity position-aware attention that is better suited to relation extraction.Then we build TACRED, a large (119,474 examples) supervised relation extraction dataset, obtained via crowdsourcing and targeted towards TAC KBP relations.The combination of better supervised data and a more appropriate high-capacity model enables much better relation extraction performance.When the model trained on this new dataset replaces the previous relation extraction component of the best TAC KBP 2015 slot filling system, its F 1 score increases markedly from 22.2% to 26.7%. Yuhao Zhang 0004, Victor Zhong, Danqi Chen 0001, Gabor Angeli, Christopher D. Manning |
EMNLP | 2 |
| 2017 | Dynamic Coattention Networks For Question Answering
Caiming Xiong, Victor Zhong, Richard Socher |
ICLR (Poster) | 2 |
| 2016 | Ask Me Anything: Dynamic Memory Networks for Natural Language ProcessingabstractMost tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers. Questions trigger an iterative attention process which allows the model to condition its attention on the inputs and the result of previous iterations. These results are then reasoned over in a hierarchical recurrent sequence model to generate answers. The DMN can be trained end-to-end and obtains state-of-the-art results on several types of tasks and datasets: question answering (Facebook’s bAbI dataset), text classification for sentiment analysis (Stanford Sentiment Treebank) and sequence modeling for part-of-speech tagging (WSJ-PTB). The training for these different tasks relies exclusively on trained word vector representations and input-question-answer triplets. Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury 0002, Ishaan Gulrajani, Victor Zhong, Romain Paulus, Richard Socher |
ICML | 7 |