VLDB 2026 Research / reviewers in the wild / expert
Kyle Richardson 0001
dblp:38/9169 · also Kyle D. Richardson
· DBLP profile ↗
31ranked-venue papers
10as first author
20since 2021 · last 2025
0000-0003-4836-0753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding the Logic of Direct Preference Alignment through LogicabstractRecent direct preference alignment algorithms (DPA), such as DPO, have shown great promise in aligning large language models to human preferences. While this has motivated the development of many new variants of the original DPO loss, understanding the differences between these recent proposals, as well as developing new DPA loss functions, remains difficult given the lack of a technical and conceptual framework for reasoning about the underlying semantics of these algorithms. In this paper, we attempt to remedy this by formalizing DPA losses in terms of discrete reasoning problems. Specifically, we ask: Given an existing DPA loss, can we systematically derive a symbolic program that characterizes its semantics? We propose a novel formalism for characterizing preference losses for single model and reference model based approaches, and identify symbolic forms for a number of commonly used DPA variants. Further, we show how this formal view of preference learning sheds new light on both the size and structure of the DPA loss landscape, making it possible to not only rigorously characterize the relationships between recent loss proposals but also to systematically explore the landscape and derive new loss functions from first principles. We hope our framework and findings will help provide useful guidance to those working on human AI alignment. Kyle Richardson 0001, Vivek Srikumar, Ashish Sabharwal |
ICML | 1 |
| 2025 | ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningabstractWe investigate the logical reasoning capabilities of Large Language Models (LLMs) and their scalability across complex deductive tasks. Using ZebraLogic, a newly developed benchmark dataset of logic grid puzzles derived from constraint satisfaction problems (CSPs), we systematically evaluate LLM performance. ZebraLogic spans a broad range of search space complexities and incorporates diverse logical constraints, providing a controlled environment to assess reasoning abilities. Our results reveal a significant decline in accuracy as problem complexity increases—a phenomenon we term the “curse of complexity.” Notably, this limitation persists even with scaling model size and inference-time computation, suggesting fundamental constraints in current LLM reasoning capabilities. Additionally, we explore strategies such as Best-of-N sampling, backtracking mechanisms, and self-verification prompts to enhance logical reasoning performance. Our findings provide critical insights into the scaling behavior of LLMs, highlight their limitations, and outline potential directions for advancing their reasoning capabilities. Bill Y. Lin, Ronan Le Bras 0001, Kyle Richardson 0001, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi 0001 |
ICML | 3 |
| 2025 | SELFGOAL: Your Language Agents Already Know How to Achieve High-level GoalsabstractRuihan Yang, Jiangjie Chen, Yikai Zhang, Siyu Yuan, Aili Chen, Kyle Richardson, Yanghua Xiao, Deqing Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruihan Yang, Jiangjie Chen, Yikai Zhang 0004, Aili Chen, Kyle Richardson 0001, Yanghua Xiao, Deqing Yang |
NAACL (Long Papers) | 6 |
| 2025 | Language Modeling by Language Modelsabstract*Can we leverage LLMs to model the process of discovering novel language model (LM) architectures?* Inspired by real research, we propose a multi-agent LLM approach that simulates the conventional stages of research, from ideation and literature search (proposal stage) to design implementation (code generation), generative pre-training, and downstream evaluation (verification). Using ideas from scaling laws, our system *Genesys* employs a *Ladder of Scales* approach; new designs are proposed, adversarially reviewed, implemented, and selectively verified at increasingly larger model scales (14M$\sim$350M parameters) with a narrowing budget (the number of models we can train at each scale). To help make discovery efficient and factorizable, Genesys uses a novel genetic programming backbone, which we show has empirical advantages over commonly used direct prompt generation workflows (e.g., $\sim$86\% percentage point improvement in successful design generation, a key bottleneck). We report experiments involving 1,162 newly discovered designs (1,062 fully verified) and find the best designs to be competitive with known architectures (e.g., outperform GPT2, Mamba2, etc., on 6/9 common benchmarks). We couple these results with comprehensive system-level ablations and formal results, which give broader insights into the design of effective autonomous discovery systems. Junyan Cheng, Peter Clark, Kyle Richardson 0001 |
NeurIPS | 3 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 37 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 24 |
| 2024 | TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware SimulationabstractDespite remarkable advancements in emulating human-like behavior through Large Language Models (LLMs), current textual simulations do not adequately address the notion of time.To this end, we introduce TIMEARENA, a novel textual simulated environment that incorporates complex temporal dynamics and constraints that better reflect real-life planning scenarios.In TIMEARENA, agents are asked to complete multiple tasks as soon as possible, allowing for parallel processing to save time.We implement the dependency between actions, the time duration for each action, and the occupancy of the agent and the objects in the environment.TIMEARENA grounds to 30 real-world tasks in cooking, household activity, and laboratory work.We conduct extensive experiments with various LLMs using TIMEARENA.Our findings reveal that even the most powerful models, e.g., GPT-4, still lag behind humans in effective multitasking, underscoring the need for enhanced temporal awareness in the development of language agents. 1 Yikai Zhang 0004, Caiyu Hu, Kyle Richardson 0001, Yanghua Xiao, Jiangjie Chen |
ACL (1) | 4 |
| 2024 | Event Causality Identification with Synthetic ControlabstractEvent causality identification (ECI), a process that extracts causal relations between events from text, is crucial for distinguishing causation from correlation.Traditional approaches to ECI have primarily utilized linguistic patterns and multi-hop relational inference, risking false causality identification due to informal usage of causality and specious graphical inference.In this paper, we adopt the Rubin Causal Model to identify event causality: given two temporally ordered events, we see the first event as the treatment and the second one as the observed outcome.Determining their causality involves manipulating the treatment and estimating the resultant change in the likelihood of the outcome.Given that it is only possible to implement manipulation conceptually in the text domain, as a work-around, we try to find a 'twin' for the protagonist from existing corpora.This 'twin' should have identical life experiences with the protagonist before the treatment but undergoes an intervention of treatment.However, the practical difficulty of locating such a match limits its feasibility.Addressing this issue, we use the synthetic control method to generate such a 'twin' from relevant historical data, leveraging text embedding synthesis and inversion techniques.This approach allows us to identify causal relations more robustly than previous methods, including GPT-4, which is demonstrated on a causality benchmark, COPES-hard. Haoyu Wang 0005, Fengze Liu, Jiayao Zhang 0001, Dan Roth 0001, Kyle Richardson 0001 |
EMNLP | 5 |
| 2024 | SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research RepositoriesabstractBen Bogin, Kejuan Yang, Shashank Gupta, Kyle Richardson, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ben Bogin, Kejuan Yang, Kyle Richardson 0001, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot |
EMNLP | 4 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 15 |
| 2023 | DISCO: Distilling Counterfactuals with Large Language ModelsabstractModels trained with counterfactually augmented data learn representations of the causal structure of tasks, enabling robust generalization.However, high-quality counterfactual data is scarce for most tasks and not easily generated at scale.When crowdsourced, such data is typically limited in scale and diversity; when generated using supervised methods, it is computationally expensive to extend to new counterfactual dimensions.In this work, we introduce DISCO (DIStilled COunterfactual Data), a new method for automatically generating highquality counterfactual data at scale.DISCO engineers prompts to generate phrasal perturbations with a large general language model.Then, a task-specific teacher model filters these generations to distill high-quality counterfactual data.While task-agnostic, we apply our pipeline to the task of natural language inference (NLI) and find that on challenging evaluations such as the NLI stress test, comparatively smaller student models trained with DISCOgenerated counterfactuals are more robust (6% absolute) and generalize better across distributions (2%) compared to models trained without data augmentation.Furthermore, DISCOaugmented models are 10% more consistent between counterfactual pairs on three evaluation sets, demonstrating that DISCO augmentation enables models to more reliably learn causal representations.Our Zeming Chen 0001, Qiyue Gao, Antoine Bosselut, Ashish Sabharwal, Kyle Richardson 0001 |
ACL (1) | 5 |
| 2023 | Language Models with RationalityabstractWhile large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs".This lack of interpretability is a growing impediment to widespread use of LLMs.To address this, our goals are to make model beliefs and their inferential relationships explicit, and to resolve inconsistencies that may exist, so that answers are supported by interpretable chains of reasoning drawn from a consistent network of beliefs.Our approach, which we call REFLEX, is to add a rational, self-reflecting layer on top of the LLM.First, given a question, we construct a belief graph using a backward-chaining process to materialize relevant model beliefs (including beliefs about answer candidates) and their inferential relationships.Second, we identify and minimize contradictions in that graph using a formal constraint reasoner.We find that REFLEX significantly improves consistency (by 8%-11% absolute) without harming overall answer accuracy, resulting in answers supported by faithful chains of reasoning drawn from a more consistent belief system.This suggests a new style of system architecture in which an LLM extended with a rational layer can provide an interpretable window into system beliefs, add a systematic reasoning capability, and repair latent inconsistencies present in the LLM. Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson 0001, Hinrich Schütze, Peter Clark |
EMNLP | 4 |
| 2023 | Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
ICLR | 5 |
| 2022 | Pushing the Limits of Rule Reasoning in Transformers through Natural Language SatisfiabilityabstractInvestigating the reasoning abilities of transformer models, and discovering new challenging tasks for them, has been a topic of much interest. Recent studies have found these models to be surprisingly strong at performing deductive reasoning over formal logical theories expressed in natural language. A shortcoming of these studies, however, is that they do not take into account that logical theories, when sampled uniformly at random, do not necessarily lead to hard instances. We propose a new methodology for creating challenging algorithmic reasoning datasets that focus on natural language satisfiability (NLSat) problems. The key idea is to draw insights from empirical sampling of hard propositional SAT problems and from complexity-theoretic studies of language. This methodology allows us to distinguish easy from hard instances, and to systematically increase the complexity of existing reasoning benchmarks such as RuleTaker. We find that current transformers, given sufficient training data, are surprisingly robust at solving the resulting NLSat problems of substantially increased difficulty. They also exhibit some degree of scale-invariance—the ability to generalize to problems of larger size and scope. Our results, however, reveal important limitations too: careful sampling of training data is crucial for building models that generalize to larger problems, and transformer models’ limited scale-invariance suggests they are far from learning robust deductive reasoning algorithms. Kyle Richardson 0001, Ashish Sabharwal |
AAAI | 1 |
| 2022 | Breakpoint Transformers for Modeling and Tracking Intermediate BeliefsabstractCan we teach natural language understanding models to track their beliefs through intermediate points in text?We propose a representation learning framework called breakpoint modeling that allows for learning of this type.Given any text encoder and data marked with intermediate states (breakpoints) along with corresponding textual queries viewed as true/false propositions (i.e., the candidate beliefs of a model, consisting of information changing through time) our approach trains models in an efficient and end-to-end fashion to build intermediate representations that facilitate teaching and direct querying of beliefs at arbitrary points alongside solving other end tasks.To show the benefit of our approach, we experiment with a diverse set of NLU tasks including relational reasoning on CLUTRR and narrative understanding on bAbI.Using novel belief prediction tasks for both tasks, we show the benefit of our main breakpoint transformer, based on T5, over conventional representation learning approaches in terms of processing efficiency, prediction accuracy and prediction consistency, all with minimal to no effect on corresponding QA endtasks.To show the feasibility of incorporating our belief tracker into more complex reasoning pipelines, we also obtain SOTA performance on the three-tiered reasoning challenge for the TRIP benchmark (around 23-32% absolute improvement on Tasks 2-3). 1 Kyle Richardson 0001, Ronen Tamari, Oren Sultan, Dafna Shahaf, Reut Tsarfaty, Ashish Sabharwal |
EMNLP | 1 |
| 2022 | What Makes Instruction Learning Hard? An Investigation and a New Challenge in a Synthetic EnvironmentabstractThe instruction learning paradigm-where a model learns to perform new tasks from task descriptions alone-has become popular in research on general-purpose models.The capabilities of large transformer models as instruction learners, however, remain poorly understood.We use a controlled synthetic environment to characterize such capabilities.Specifically, we use the task of deciding whether a given string matches a regular expression (viewed as an instruction) to identify properties of tasks, instructions, and instances that make instruction learning challenging.For instance, we find that our model, a fine-tuned T5-based text2text transformer, struggles with large regular languages, suggesting that less precise instructions are challenging for models.Instruction executions that require tracking longer contexts of prior steps are also difficult.We use our findings to systematically construct a challenging instruction learning dataset, which we call Hard RegSet.Fine-tuning on Hard RegSet, our large transformer learns to correctly interpret (with at least 90% accuracy) only 65.6% of test instructions, and 11%-24% of the instructions in out-of-distribution generalization settings.We thus propose Hard RegSet as a challenging instruction learning dataset, and a controlled environment for studying instruction learning.1 Matthew Finlayson, Kyle Richardson 0001, Ashish Sabharwal, Peter Clark |
EMNLP | 2 |
| 2022 | Learning to Decompose: Hypothetical Question Decomposition Based on Comparable TextsabstractExplicit decomposition modeling, which involves breaking down complex tasks into more straightforward and often more interpretable sub-tasks, has long been a central theme in developing robust and interpretable NLU systems.However, despite the many datasets and resources built as part of this effort, the majority have small-scale annotations and limited scope, which is insufficient to solve general decomposition tasks.In this paper, we look at large-scale intermediate pre-training of decomposition-based transformers using distant supervision from comparable texts, particularly large-scale parallel news.We show that with such intermediate pre-training, developing robust decomposition-based models for a diverse range of tasks becomes more feasible.For example, on semantic parsing, our model, DECOMPT5, improves 20% to 30% on two datasets, Overnight and TORQUE, over the baseline language model.We further use DECOMPT5 to build a novel decompositionbased QA system named DECOMPENTAIL, improving over state-of-the-art models, including GPT-3, on both HotpotQA and StrategyQA by 8% and 4%, respectively. Ben Zhou, Kyle Richardson 0001, Xiaodong Yu 0003, Dan Roth 0001 |
EMNLP | 2 |
| 2022 | Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous PromptsabstractDaniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001 |
NAACL-HLT | 5 |
| 2021 | Text Modular Networks: Learning to Decompose Tasks in the Language of Existing ModelsabstractTushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, Ashish Sabharwal. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tushar Khot, Daniel Khashabi, Kyle Richardson 0001, Peter Clark, Ashish Sabharwal |
NAACL-HLT | 3 |
| 2021 | Temporal Reasoning on Implicit Events from Distant SupervisionabstractBen Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ben Zhou, Kyle Richardson 0001, Qiang Ning, Tushar Khot, Ashish Sabharwal, Dan Roth 0001 |
NAACL-HLT | 2 |
| 2020 | Probing Natural Language Inference Models through Semantic FragmentsabstractDo state-of-the-art models for language understanding already have, or can they easily learn, abilities such as boolean coordination, quantification, conditionals, comparatives, and monotonicity reasoning (i.e., reasoning about word substitutions in sentential contexts)? While such phenomena are involved in natural language inference (NLI) and go beyond basic linguistic understanding, it is unclear the extent to which they are captured in existing NLI benchmarks and effectively learned by models. To investigate this, we propose the use of semantic fragments—systematically generated datasets that each target a different semantic phenomenon—for probing, and efficiently improving, such capabilities of linguistic models. This approach to creating challenge datasets allows direct control over the semantic diversity and complexity of the targeted linguistic phenomena, and results in a more precise characterization of a model's linguistic behavior. Our experiments, using a library of 8 such semantic fragments, reveal two remarkable findings: (a) State-of-the-art models, including BERT, that are pre-trained on existing NLI benchmark datasets perform poorly on these new fragments, even though the phenomena probed here are central to the NLI task; (b) On the other hand, with only a few minutes of additional fine-tuning—with a carefully selected learning rate and a novel variation of “inoculation”—a BERT-based model can master all of these logic and monotonicity fragments while retaining its performance on established NLI benchmarks. Kyle Richardson 0001, Hai Hu 0001, Lawrence S. Moss, Ashish Sabharwal |
AAAI | 1 |
| 2020 | CLUE: A Chinese Language Understanding Evaluation BenchmarkabstractLiang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan |
COLING | 31 |
| 2020 | A Dataset for Tracking Entities in Open Domain Procedural TextabstractNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson 0001, Eduard H. Hovy |
EMNLP (1) | 7 |
| 2020 | Transformers as Soft Reasoners over LanguageabstractBeginning with McCarthy's Advice Taker (1959), AI has pursued the goal of providing a system with explicit, general knowledge and having the system reason over that knowledge. However, expressing the knowledge in a formal (logical or probabilistic) representation has been a major obstacle to this research. This paper investigates a modern approach to this problem where the facts and rules are provided as natural language sentences, thus bypassing a formal representation. We train transformers to reason (or emulate reasoning) over these sentences using synthetically generated data. Our models, that we call RuleTakers, provide the first empirical demonstration that this kind of soft reasoning over language is learnable, can achieve high (99%) accuracy, and generalizes to test data requiring substantially deeper chaining than seen during training (95%+ scores). We also demonstrate that the models transfer well to two hand-authored rulebases, and to rulebases paraphrased into more natural language. These findings are significant as it suggests a new role for transformers, namely as limited "soft theorem provers" operating over explicit theories in language. This in turn suggests new possibilities for explainability, correctability, and counterfactual reasoning in question-answering. Peter Clark, Oyvind Tafjord, Kyle Richardson 0001 |
IJCAI | 3 |
| 2020 | What Does My QA Model Know? Devising Controlled Probes using ExpertabstractOpen-domain question answering (QA) involves many knowledge and reasoning challenges, but are successful QA models actually learning such knowledge when trained on benchmark QA tasks? We investigate this via several new diagnostic tasks probing whether multiple-choice QA models know definitions and taxonomic reasoning—two skills widespread in existing benchmarks and fundamental to more complex reasoning. We introduce a methodology for automatically building probe datasets from expert knowledge sources, allowing for systematic control and a comprehensive evaluation. We include ways to carefully control for artifacts that may arise during this process. Our evaluation confirms that transformer-based multiple-choice QA models are already predisposed to recognize certain types of structural linguistic knowledge. However, it also reveals a more nuanced picture: their performance notably degrades even with a slight increase in the number of “hops” in the underlying taxonomic hierarchy, and with more challenging distractor candidates. Further, existing models are far from perfect when assessed at the level of clusters of semantically connected probes, such as all hypernym questions about a single concept. Kyle Richardson 0001, Ashish Sabharwal |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Polyglot Semantic Parsing in APIsabstractKyle Richardson, Jonathan Berant, Jonas Kuhn. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Kyle Richardson 0001, Jonathan Berant, Jonas Kuhn |
NAACL-HLT | 1 |
| 2017 | Learning Semantic Correspondences in Technical DocumentationabstractWe consider the problem of translating high-level textual descriptions to formal representations in technical documentation as part of an effort to model the meaning of such documentation.We focus specifically on the problem of learning translational correspondences between text descriptions and grounded representations in the target documentation, such as formal representation of functions or code templates.Our approach exploits the parallel nature of such documentation, or the tight coupling between high-level text and the low-level representations we aim to learn.Data is collected by mining technical documents for such parallel text-representation pairs, which we use to train a simple semantic parsing model.We report new baseline results on sixteen novel datasets, including the standard library documentation for nine popular programming languages across seven natural languages, and a small collection of Unix utility manuals.1. Java Documentation * Returns the greater of two long values * * @param a an argument * @param b another argument * @return the larger of a and b * @see java.lang.Long#MAX VALUE * / public static long max(long a, long b) 2. Clojure Documentation (defn random-sample "Returns items from coll with random probability of prob (0.0 -1.0)" ([prob coll] ...)) 3. PHP documentation (French) Ajoute une valeur comme dernier élément * * @param value La valeur á ajouter * @see ArrayIterations::next() * / public void append(mixed $value) Kyle Richardson 0001, Jonas Kuhn |
ACL (1) | 1 |
| 2017 | The Code2Text Challenge: Text Generation in Source LibrariesabstractWe propose a new shared task for tactical datato-text generation in the domain of source code libraries.Specifically, we focus on text generation of function descriptions from example software projects.Data is drawn from existing resources used for studying the related problem of semantic parser induction (Richardson and Kuhn, 2017b; Richardson and Kuhn, 2017a), and spans a wide variety of both natural languages and programming languages.In this paper, we describe these existing resources, which will serve as training and development data for the task, and discuss plans for building new independent test sets.1. Java Documentation * Returns the greater of two long values * / ... public static long max(long a, long b) 2. Python Documentation # from decimal.Context max(self, a, b): """Compares two values numerically and returns the maximum""" 3. aNALoGuE Challenge (Novikova and Rieser, 2016) MR input: name[Bibmbap House] food[French] priceRange[cheap], area[riverside] near[Clare Hall] NL output: Near Clare Hall, in the riverside area, Bibimbap serves French food in the price range cheap. Kyle Richardson 0001, Sina Zarrieß, Jonas Kuhn |
INLG | 1 |
| 2016 | Learning to Make Inferences in a Semantic Parsing TaskabstractWe introduce a new approach to training a semantic parser that uses textual entailment judgements as supervision. These judgements are based on high-level inferences about whether the meaning of one sentence follows from another. When applied to an existing semantic parsing task, they prove to be a useful tool for revealing semantic distinctions and background knowledge not captured in the target representations. This information is used to improve the quality of the semantic representations being learned and to acquire generic knowledge for reasoning. Experiments are done on the benchmark Sportscaster corpus (Chen and Mooney, 2008), and a novel RTE-inspired inference dataset is introduced. On this new dataset our method strongly outperforms several strong baselines. Separately, we obtain state-of-the-art results on the original Sportscaster semantic parsing task. Kyle Richardson 0001, Jonas Kuhn |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | UnixMan Corpus: A Resource for Language Learning in the Unix Domain
Kyle Richardson 0001, Jonas Kuhn |
LREC | 1 |
| 2011 | Deducing answers to english questions from structured dataabstractWe describe ongoing research using natural English text queries as an intelligent interface for inferring answers from structured data in a specific domain. Users can express queries whose answers need to be deduced from data in different databases, without knowing the structures of those databases nor even the existence of the sources used. Users can pose queries incrementally, elaborating on an initial query, and ask follow-up questions based on answers to earlier queries. Daniel G. Bobrow, Cleo Condoravdi, Kyle Richardson 0001, Richard J. Waldinger, Amar Das |
IUI | 3 |