Sebastian Riedel 0001

dblp:18/3348-1 · DBLP profile ↗
← Back
90ranked-venue papers
11as first author
27since 2021 · last 2024
0000-0002-3655-2486ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 11 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 Do Large Language Models Latently Perform Multi-Hop Reasoning?
abstract
We study whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of 'Superstition' is".We look for evidence of a latent reasoning pathway where an LLM (1) latently identifies "the singer of 'Superstition"' as Stevie Wonder, the bridge entity, and (2) uses its knowledge of Stevie Wonder's mother to complete the prompt.We analyze these two hops individually and consider their co-occurrence as indicative of latent multi-hop reasoning.For the first hop, we test if changing the prompt to indirectly mention the bridge entity instead of any other entity increases the LLM's internal recall of the bridge entity.For the second hop, we test if increasing this recall causes the LLM to better utilize what it knows about the bridge entity.We find strong evidence of latent multi-hop reasoning for the prompts of certain relation types, with the reasoning pathway used in more than 80% of the prompts.However, the utilization is highly contextual, varying across different types of prompts.Also, on average, the evidence for the second hop and the full multi-hop traversal is rather moderate and only substantial for the first hop.Moreover, we find a clear scaling trend with increasing model size for the first hop of reasoning but not for the second hop.Our experimental findings suggest potential challenges and opportunities for future development and applications of LLMs.
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel 0001
ACL (1)5
2024 Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt Optimisation
abstract
Yao Lu, Jiayi Wang, Raphael Tang, Sebastian Riedel, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiayi Wang 0010, Raphael Tang, Sebastian Riedel 0001, Pontus Stenetorp
NAACL-HLT4
2023 Can discrete information extraction prompts generalize across language models?
Nathanaël Carraz Rakotonirina, Roberto Dessì, Fabio Petroni, Sebastian Riedel 0001, Marco Baroni
ICLR4
2023 PEER: A Collaborative Language Model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick S. H. Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, Sebastian Riedel 0001
ICLR10
2023 Improving Language Plasticity via Pretraining with Active Forgetting
abstract
Pretrained language models (PLMs) are today the primary model for natural language processing. Despite their impressive downstream performance, it can be difficult to apply PLMs to new languages, a barrier to making their capabilities universally accessible. While prior work has shown it possible to address this issue by learning a new embedding layer for the new language, doing so is both data and compute inefficient. We propose to use an active forgetting mechanism during pretraining, as a simple way of creating PLMs that can quickly adapt to new languages. Concretely, by resetting the embedding layer every K updates during pretraining, we encourage the PLM to improve its ability of learning new embeddings within limited number of updates, similar to a meta-learning effect. Experiments with RoBERTa show that models pretrained with our forgetting mechanism not only demonstrate faster convergence during language adaptation, but also outperform standard ones in a low-data regime, particularly for languages that are distant from English. Code will be available at https://github.com/facebookresearch/language-model-plasticity.
Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel 0001, Mikel Artetxe
NeurIPS6
2023 Atlas: Few-shot Learning with Retrieval Augmented Language Models
abstract
Large language models have shown impressive few-shot results on a wide range of tasks. However, when knowledge is key for such results, as is the case for tasks such as question answering and fact checking, massive parameter counts to store knowledge seem to be needed. Retrieval-augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings. In this work we present Atlas, a carefully designed and pre-trained retrieval-augmented language model able to learn knowledge intensive tasks with very few training examples. We perform evaluations on a wide range of tasks, including MMLU, KILT and Natural Questions, and study the impact of the content of the document index, showing that it can easily be updated. Notably, Atlas reaches over 42% accuracy on Natural Questions using only 64 examples, outperforming a 540B parameter model by 3% despite having 50x fewer parameters.
Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel 0001, Edouard Grave
J. Mach. Learn. Res.9
2022 Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
abstract
When primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, finetuned, large, pretrained language models.We demonstrate that the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance: essentially some permutations are "fantastic" and some not.We analyse this phenomenon in detail, establishing that: it is present across model sizes (even for the largest current models), it is not related to a specific subset of samples, and that a given good permutation for one model is not transferable to another.While one could use a development set to determine which permutations are performant, this would deviate from the true fewshot setting as it requires additional annotated data.Instead, we use the generative nature of language models to construct an artificial development set and based on entropy statistics of the candidate permutations on this set, we identify performant prompts.Our method yields a 13% relative improvement for GPTfamily models across eleven different established text classification tasks.
Max Bartolo, Alastair Moore, Sebastian Riedel 0001, Pontus Stenetorp
ACL (1)4
2022 EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and Indexing
abstract
Existing work on Entity Linking mostly assumes that the reference knowledge base is complete, and therefore all mentions can be linked.In practice this is hardly ever the case, as knowledge bases are incomplete and because novel concepts arise constantly.We introduce the temporally segmented Unknown Entity Discovery and Indexing (EDIN) -benchmark where unknown entities, that is entities not part of the knowledge base and without descriptions and labeled mentions, have to be integrated into an existing entity linking system.By contrasting EDIN with zero-shot entity linking, we provide insight on the additional challenges it poses.Building on denseretrieval based entity linking, we introduce the end-to-end EDIN-pipeline that detects, clusters, and indexes mentions of unknown entities in context.Experiments show that indexing a single embedding per entity unifying the information of multiple mentions works better than indexing mentions independently.
Nora Kassner, Fabio Petroni, Mikhail Plekhanov, Sebastian Riedel 0001, Nicola Cancedda
EMNLP4
2022 An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks
abstract
Access to external knowledge is essential for many natural language processing tasks, such as question answering and dialogue.Existing methods often rely on a parametric model that stores knowledge in its parameters, or use a retrieval-augmented model that has access to an external knowledge source.Parametric and retrieval-augmented models have complementary strengths in terms of computational efficiency and predictive accuracy.To combine the strength of both approaches, we propose the Efficient Memory-Augmented Transformer (EMAT) -it encodes external knowledge into a key-value memory and exploits the fast maximum inner product search for memory querying.We also introduce pre-training tasks that allow EMAT to encode informative key-value representations, and to learn an implicit strategy to integrate multiple memory slots into the transformer.Experiments on various knowledge-intensive tasks such as question answering and dialogue datasets show that, simply augmenting parametric models (T5-base) using our method produces more accurate results (e.g., 25.8 → 44.3 EM on NQ) while retaining a high throughput (e.g., 1000 queries/s on NQ).Compared to retrievalaugmented models, EMAT runs substantially faster across the board and produces more accurate results on WoW and ELI5. 1
Yuxiang Wu, Yu Zhao 0043, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
EMNLP6
2022 Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants
abstract
Max Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, Douwe Kiela. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Max Bartolo, Tristan Thrush, Sebastian Riedel 0001, Pontus Stenetorp, Robin Jia, Douwe Kiela
NAACL-HLT3
2022 Boosted Dense Retriever
abstract
Patrick Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Patrick S. H. Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel 0001
NAACL-HLT6
2022 Lifting the Curse of Multilinguality by Pre-training Modular Transformers
abstract
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe
NAACL-HLT6
2022 Autoregressive Search Engines: Generating Substrings as Document Identifiers
abstract
Knowledge-intensive language tasks require NLP systems to both provide the correct answer and retrieve supporting evidence for it in a given corpus. Autoregressive language models are emerging as the de-facto standard for generating answers, with newer and more powerful systems emerging at an astonishing pace. In this paper we argue that all this (and future) progress can be directly applied to the retrieval problem with minimal intervention to the models' architecture. Previous work has explored ways to partition the search space into hierarchical structures and retrieve documents by autoregressively generating their unique identifier. In this work we propose an alternative that doesn't force any structure in the search space: using all ngrams in a passage as its possible identifiers. This setup allows us to use an autoregressive model to generate and score distinctive ngrams, that are then mapped to full passages through an efficient data structure. Empirically, we show this not only outperforms prior autoregressive approaches but also leads to an average improvement of at least 10 points over more established retrieval solutions for passage-level retrieval on the KILT benchmark, establishing new state-of-the-art downstream performance on some datasets, while using a considerably lighter memory footprint than competing systems. Code available in the supplementary materials. Pre-trained models will be made available.
Michele Bevilacqua, Giuseppe Ottaviano, Patrick S. H. Lewis, Scott Yih, Sebastian Riedel 0001, Fabio Petroni
NeurIPS5
2022 ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
abstract
Factorisation-based Models (FMs), such as DistMult, have enjoyed enduring success for Knowledge Graph Completion (KGC) tasks, often outperforming Graph Neural Networks (GNNs). However, unlike GNNs, FMs struggle to incorporate node features and generalise to unseen nodes in inductive settings. Our work bridges the gap between FMs and GNNs by proposing ReFactor GNNs. This new architecture draws upon $\textit{both}$ modelling paradigms, which previously were largely thought of as disjoint. Concretely, using a message-passing formalism, we show how FMs can be cast as GNNs by reformulating the gradient descent procedure as message-passing operations, which forms the basis of our ReFactor GNNs. Across a multitude of well-established KGC benchmarks, our ReFactor GNNs achieve comparable transductive performance to FMs, and state-of-the-art inductive performance while using an order of magnitude fewer parameters.
Pushkar Mishra, Luca Franceschi 0001, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
NeurIPS6
2022 Multilingual Autoregressive Entity Linking
abstract
Abstract We present mGENRE, a sequence-to- sequence system for the Multilingual Entity Linking (MEL) problem—the task of resolving language-specific mentions to a multilingual Knowledge Base (KB). For a mention in a given language, mGENRE predicts the name of the target entity left-to-right, token-by-token in an autoregressive fashion. The autoregressive formulation allows us to effectively cross-encode mention string and entity names to capture more interactions than the standard dot product between mention and entity vectors. It also enables fast search within a large KB even for mentions that do not appear in mention tables and with no need for large-scale vector indices. While prior MEL works use a single representation for each entity, we match against entity names of as many languages as possible, which allows exploiting language connections between source input and target name. Moreover, in a zero-shot setting on languages with no training data at all, mGENRE treats the target language as a latent variable that is marginalized at prediction time. This leads to over 50% improvements in average accuracy. We show the efficacy of our approach through extensive evaluation including experiments on three popular MEL benchmarks where we establish new state-of-the-art results. Source code available at https://github.com/facebookresearch/GENRE.
Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal 0001, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel 0001, Fabio Petroni
Trans. Assoc. Comput. Linguistics9
2022 ProoFVer: Natural Logic Theorem Proving for Fact Verification
abstract
Abstract Fact verification systems typically rely on neural network classifiers for veracity prediction, which lack explainability. This paper proposes ProoFVer, which uses a seq2seq model to generate natural logic-based inferences as proofs. These proofs consist of lexical mutations between spans in the claim and the evidence retrieved, each marked with a natural logic operator. Claim veracity is determined solely based on the sequence of these operators. Hence, these proofs are faithful explanations, and this makes ProoFVer faithful by construction. Currently, ProoFVer has the highest label accuracy and the second best score in the FEVER leaderboard. Furthermore, it improves by 13.21% points over the next best model on a dataset with counterfactual instances, demonstrating its robustness. As explanations, the proofs show better overlap with human rationales than attention-based highlights and the proofs help humans predict model decisions correctly more often than using the evidence directly.1
Amrith Krishna, Sebastian Riedel 0001, Andreas Vlachos 0001
Trans. Assoc. Comput. Linguistics2
2021 Joint Verification and Reranking for Open Fact Checking Over Tables
abstract
Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, Sebastian Riedel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Scott Yih, Sebastian Riedel 0001
ACL/IJCNLP (1)6
2021 Database reasoning over text
abstract
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, Alon Halevy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel 0001, Alon Y. Halevy
ACL/IJCNLP (1)5
2021 Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets
abstract
Ideally Open-Domain Question Answering models should exhibit a number of competencies, ranging from simply memorizing questions seen at training time, to answering novel question formulations with answers seen during training, to generalizing to completely novel questions with novel answers.However, single aggregated test set scores do not show the full picture of what capabilities models truly have.In this work, we perform a detailed study of the test sets of three popular open-domain benchmark datasets with respect to these competencies.We find that 30% of test-set questions have a near-duplicate paraphrase in their corresponding train sets.In addition, we find that 60-70% of answers in the test sets are also present in the train sets.Using these findings, we evaluate a variety of popular open-domain models to obtain greater insight into what extent they can generalize, and what drives their overall performance.We find that all models perform substantially worse on questions that cannot be memorized from train sets, with a mean absolute performance difference of 61% between repeated and nonrepeated data.Finally we show that simple nearest-neighbor models outperform a BART closed-book QA model, further highlighting the role that train set memorization plays in these benchmarks.
Patrick S. H. Lewis, Pontus Stenetorp, Sebastian Riedel 0001
EACL3
2021 Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation
abstract
Despite recent progress, state-of-the-art question answering models remain vulnerable to a variety of adversarial attacks. While dynamic adversarial data collection, in which a human annotator tries to write examples that fool a model-in-the-loop, can improve model robustness, this process is expensive which limits the scale of the collected data. In this work, we are the first to use synthetic adversarial data generation to make question answering models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data.
Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 0001, Pontus Stenetorp, Douwe Kiela
EMNLP (1)4
2021 Autoregressive Entity Retrieval
Nicola De Cao, Gautier Izacard, Sebastian Riedel 0001, Fabio Petroni
ICLR3
2021 Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz
ICLR9
2021 Concept Matching for Low-Resource Classification
abstract
In many applications that rely on machine learning, the availability of labelled data is a matter of primary importance. However, when tackling new tasks, labels are usually missing and must be collected from scratch by the users. In this work, we address the problem of learning classifiers when the amount of labels is very scarce. We do so by learning multiple vectors, called prototypes, that represent relevant semantic concepts for the task at hand. We propose a theoretically inspired mechanism that computes probabilities of matching between the prototypes and the input elements, and we combine these probabilities to increase the expressiveness of the classifier. Moreover, by leveraging low-cost extra annotations in the training data, a simple error-boosting technique guides the learning process and provides substantial performance improvements. Empirical results confirm the benefits of the proposed approach in both balanced and unbalanced datasets. Our methodology is thus of practical use when gathering and labelling new examples is more expensive than annotating what we already have.
Federico Errica, Fabrizio Silvestri, Bora Edizel, Ludovic Denoyer, Fabio Petroni, Vassilis Plachouras, Sebastian Riedel 0001
IJCNN7
2021 Dynabench: Rethinking Benchmarking in NLP
abstract
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
NAACL-HLT13
2021 KILT: a Benchmark for Knowledge Intensive Language Tasks
abstract
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick S. H. Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel 0001
NAACL-HLT13
2021 From Natural Language Processing to Neural Databases
abstract
In recent years, neural networks have shown impressive performance gains on long-standing AI problems, such as answering queries from text and machine translation. These advances raise the question of whether neural nets can be used at the core of query processing to derive answers from facts, even when the facts are expressed in natural language. If so, it is conceivable that we could relax the fundamental assumption of database management, namely, that our data is represented as fields of a pre-defined schema. Furthermore, such technology would enable combining information from text, images, and structured data seamlessly. This paper introduces neural databases , a class of systems that use NLP transformers as localized answer derivation engines. We ground the vision in NeuralDB, a system for querying facts represented as short natural language sentences. We demonstrate that recent natural language processing models, specifically transformers, can answer select-project-join queries if they are given a set of relevant facts. However, they cannot scale to non-trivial databases nor answer set-based and aggregation queries. Based on these insights, we identify specific research challenges that are needed to build neural databases. Some of the challenges require drawing upon the rich literature in data management, and others pose new research opportunities to the NLP community. Finally, we show that with preliminary solutions, NeuralDB can already answer queries over thousands of sentences with very high accuracy.
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel 0001, Alon Y. Halevy
Proc. VLDB Endow.5
2021 PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them
abstract
Abstract Open-domain Question Answering models that directly leverage question-answer (QA) pairs, such as closed-book QA (CBQA) models and QA-pair retrievers, show promise in terms of speed and memory compared with conventional models which retrieve and read from text corpora. QA-pair retrievers also offer interpretable answers, a high degree of control, and are trivial to update at test time with new knowledge. However, these models fall short of the accuracy of retrieve-and-read systems, as substantially less knowledge is covered by the available QA-pairs relative to text corpora like Wikipedia. To facilitate improved QA-pair models, we introduce Probably Asked Questions (PAQ), a very large resource of 65M automatically generated QA-pairs. We introduce a new QA-pair retriever, RePAQ, to complement PAQ. We find that PAQ preempts and caches test questions, enabling RePAQ to match the accuracy of recent retrieve-and-read models, whilst being significantly faster. Using PAQ, we train CBQA models which outperform comparable baselines by 5%, but trail RePAQ by over 15%, indicating the effectiveness of explicit retrieval. RePAQ can be configured for size (under 500MB) or speed (over 1K questions per second) while retaining high accuracy. Lastly, we demonstrate RePAQ’s strength at selective QA, abstaining from answering when it is likely to be incorrect. This enables RePAQ to “back-off” to a more expensive state-of-the-art model, leading to a combined system which is both more accurate and 2x faster than the state-of-the-art model alone.
Patrick S. H. Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, Sebastian Riedel 0001
Trans. Assoc. Comput. Linguistics8
2020 Differentiable Reasoning on Large Knowledge Bases and Natural Language
abstract
Reasoning with knowledge expressed in natural language and Knowledge Bases (KBs) is a major challenge for Artificial Intelligence, with applications in machine reading, dialogue, and question answering. General neural architectures that jointly learn representations and transformations of text are very data-inefficient, and it is hard to analyse their reasoning process. These issues are addressed by end-to-end differentiable reasoning systems such as Neural Theorem Provers (NTPs), although they can only be used with small-scale symbolic KBs. In this paper we first propose Greedy NTPs (GNTPs), an extension to NTPs addressing their complexity and scalability limitations, thus making them applicable to real-world datasets. This result is achieved by dynamically constructing the computation graph of NTPs and including only the most promising proof paths during inference, thus obtaining orders of magnitude more efficient models 1. Then, we propose a novel approach for jointly reasoning over KBs and textual mentions, by embedding logic facts and natural language sentences in a shared embedding space. We show that GNTPs perform on par with NTPs at a fraction of their cost while achieving competitive link prediction results on large datasets, providing explanations for predictions, and inducing interpretable models.
Pasquale Minervini, Matko Bosnjak, Tim Rocktäschel, Sebastian Riedel 0001, Edward Grefenstette
AAAI4
2020 MLQA: Evaluating Cross-lingual Extractive Question Answering
abstract
Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets.Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making building QA systems that work well in other languages challenging.In order to develop such systems, it is crucial to invest in high quality multilingual evaluation benchmarks to measure progress.We present MLQA, a multi-way aligned extractive QA evaluation benchmark intended to spur research in this area.1 MLQA contains QA instances in 7 languages, English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese.MLQA has over 12K instances in English and 5K in each other language, with each instance parallel between 4 languages on average.We evaluate stateof-the-art cross-lingual models and machinetranslation-based baselines on MLQA.In all cases, transfer results are significantly behind training-language performance.
Patrick S. H. Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel 0001, Holger Schwenk
ACL4
2020 TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
abstract
Recent years have witnessed the burgeoning of pretrained language models (LMs) for textbased natural language (NL) understanding tasks.Such models are typically trained on free-form NL text, hence may not be suitable for tasks like semantic parsing over structured data, which require reasoning over both free-form NL questions and structured tabular data (e.g., database tables).In this paper we present TABERT, a pretrained LM that jointly learns representations for NL sentences and (semi-)structured tables.TABERT is trained on a large corpus of 26 million tables and their English contexts.In experiments, neural semantic parsers using TABERT as feature representation layers achieve new best results on the challenging weakly-supervised semantic parsing benchmark WIKITABLEQUESTIONS, while performing competitively on the text-to-SQL dataset SPIDER. 1
Graham Neubig, Scott Yih, Sebastian Riedel 0001
ACL4
2020 Generating Fact Checking Briefs
abstract
Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, Sebastian Riedel. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos 0001, Antoine Bordes, Sebastian Riedel 0001
EMNLP (1)8
2020 AxCell: Automatic Extraction of Results from Machine Learning Papers
abstract
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, Robert Stojnic. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel 0001, Ross Taylor, Robert Stojnic
EMNLP (1)5
2020 Avoiding the Hypothesis-Only Bias in Natural Language Inference via Ensemble Adversarial Training
abstract
Natural Language Inference (NLI) datasets contain annotation artefacts resulting in spurious correlations between the natural language utterances and their respective entailment classes.These artefacts are exploited by neural networks even when only considering the hypothesis and ignoring the premise, leading to unwanted biases.Belinkov et al. (2019b) proposed tackling this problem via adversarial training, but this can lead to learned sentence representations that still suffer from the same biases.We show that the bias can be reduced in the sentence representations by using an ensemble of adversaries, encouraging the model to jointly decrease the accuracy of these different adversaries while fitting the data.This approach produces more robust NLI models, outperforming previous de-biasing efforts when generalised to 12 other NLI datasets (Belinkov et al., 2019a;Mahabadi et al., 2020).In addition, we find that the optimal number of adversarial classifiers depends on the dimensionality of the sentence representations, with larger sentence representations being more difficult to de-bias while benefiting from using a greater number of adversaries.
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel 0001, Tim Rocktäschel
EMNLP (1)4
2020 Scalable Zero-shot Entity Linking with Dense Entity Retrieval
abstract
This paper introduces a conceptually simple, scalable, and highly effective BERT-based entity linking model, along with an extensive evaluation of its accuracy-speed trade-off.We present a two-stage zero-shot linking algorithm, where each entity is defined only by a short textual description.The first stage does retrieval in a dense space defined by a bi-encoder that independently embeds the mention context and the entity descriptions.Each candidate is then re-ranked with a crossencoder, that concatenates the mention and entity text.Experiments demonstrate that this approach is state of the art on recent zeroshot benchmarks (6 point absolute gains) and also on more established non-zero-shot evaluations (e.g.TACKBP-2010), despite its relative simplicity (e.g.no explicit entity embeddings or manually engineered mention tables).We also show that bi-encoder linking is very fast with nearest neighbour search (e.g.linking with 5.9 million candidates in 2 milliseconds), and that much of the accuracy gain from the more expensive crossencoder can be transferred to the bi-encoder via knowledge distillation.Our code and models are available at https://github. com/facebookresearch/BLINK. wikipedia dense spaceMy kids really enjoyed a ride in the Jaguar!Jaguar is the luxury vehicle brand. Jaguar_carsJaguar! is a junior roller coaster.
Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel 0001, Luke Zettlemoyer
EMNLP (1)4
2020 Don't Read Too Much Into It: Adaptive Computation for Open-Domain Question Answering
abstract
Most approaches to Open-Domain Question Answering consist of a light-weight retriever that selects a set of candidate passages, and a computationally expensive reader that examines the passages to identify the correct answer.Previous works have shown that as the number of retrieved passages increases, so does the performance of the reader.However, they assume all retrieved passages are of equal importance and allocate the same amount of computation to them, leading to a substantial increase in computational cost.To reduce this cost, we propose the use of adaptive computation to control the computational budget allocated for the passages to be read.We first introduce a technique operating on individual passages in isolation which relies on anytime prediction and a per-layer estimation of an early exit probability.We then introduce SKY-LINEBUILDER, an approach for dynamically deciding on which passage to allocate computation at each step, based on a resource allocation policy trained via reinforcement learning.Our results on SQuAD-Open show that adaptive computation with global prioritisation improves over several strong static and adaptive methods, leading to a 4.3x reduction in computation while retaining 95% performance of the full model.
Yuxiang Wu, Sebastian Riedel 0001, Pasquale Minervini, Pontus Stenetorp
EMNLP (1)2
2020 Learning Reasoning Strategies in End-to-End Differentiable Proving
abstract
Attempts to render deep learning models interpretable, data-efficient, and robust have seen some success through hybridisation with rule-based systems, for example, in Neural Theorem Provers (NTPs). These neuro-symbolic models can induce interpretable rules and learn representations from data via back-propagation, while providing logical explanations for their predictions. However, they are restricted by their computational complexity, as they need to consider all possible proof paths for explaining a goal, thus rendering them unfit for large-scale applications. We present Conditional Theorem Provers (CTPs), an extension to NTPs that learns an optimal rule selection strategy via gradient-based optimisation. We show that CTPs are scalable and yield state-of-the-art results on the CLUTRR dataset, which tests systematic generalisation of neural models by learning to reason over smaller graphs and evaluating on larger ones. Finally, CTPs show better link prediction results on standard benchmarks in comparison with other neural-symbolic models, while being explainable.
Pasquale Minervini, Sebastian Riedel 0001, Pontus Stenetorp, Edward Grefenstette, Tim Rocktäschel
ICML2
2020 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
abstract
Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal 0001, Heinrich Küttler, Mike Lewis, Scott Yih, Tim Rocktäschel, Sebastian Riedel 0001, Douwe Kiela
NeurIPS11
2020 Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension
abstract
Innovations in annotation methodology have been a catalyst for Reading Comprehension (RC) datasets and models. One recent trend to challenge current RC models is to involve a model in the annotation process: Humans create questions adversarially, such that the model fails to answer them correctly. In this work we investigate this annotation methodology and apply it in three different settings, collecting a total of 36,000 samples with progressively stronger models in the annotation loop. This allows us to explore questions such as the reproducibility of the adversarial effect, transfer from data collected with varying model-in-the-loop strengths, and generalization to data collected without a model. We find that training on adversarially collected samples leads to strong generalization to non-adversarially collected datasets, yet with progressive performance deterioration with increasingly stronger models-in-the-loop. Furthermore, we find that stronger models can still learn from datasets collected with substantially weaker models-in-the-loop. When trained on data collected with a BiDAF model in the loop, RoBERTa achieves 39.9F1 on questions that it cannot answer when trained on SQuAD—only marginally lower than when trained on data collected using RoBERTa itself (41.0F1).
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel 0001, Pontus Stenetorp
Trans. Assoc. Comput. Linguistics4
2019 Unsupervised Question Answering by Cloze Translation
abstract
Obtaining training data for Question Answering (QA) is time-consuming and resourceintensive, and existing QA datasets are only available for limited domains and languages.In this work, we explore to what extent high quality training data is actually required for Extractive QA, and investigate the possibility of unsupervised Extractive QA.We approach this problem by first learning to generate context, question and answer triples in an unsupervised manner, which we then use to synthesize Extractive QA training data automatically.To generate such triples, we first sample random context paragraphs from a large corpus of documents and then random noun phrases or named entity mentions from these paragraphs as answers.Next we convert answers in context to "fill-in-the-blank" cloze questions and finally translate them into natural questions.We propose and compare various unsupervised ways to perform cloze-tonatural question translation, including training an unsupervised NMT model using nonaligned corpora of natural questions and cloze questions as well as a rule-based approach.We find that modern QA models can learn to answer human questions surprisingly well using only synthetic training data.We demonstrate that, without using the SQuAD training data at all, our approach achieves 56.4 F1 on SQuAD v1 (64.5 F1 when the answer is a Named entity mention), outperforming early supervised models.
Patrick S. H. Lewis, Ludovic Denoyer, Sebastian Riedel 0001
ACL (1)3
2019 Language Models as Knowledge Bases?
abstract
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander Miller. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel 0001, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller
EMNLP/IJCNLP (1)3
2018 Convolutional 2D Knowledge Graph Embeddings
abstract
Link prediction for knowledge graphs is the task of predicting missing relationships between entities. Previous work on link prediction has focused on shallow, fast models which can scale to large knowledge graphs. However, these models learn less expressive features than deep, multi-layer models — which potentially limits performance. In this work we introduce ConvE, a multi-layer convolutional network model for link prediction, and report state-of-the-art results for several established datasets. We also show that the model is highly parameter efficient, yielding the same performance as DistMult and R-GCN with 8x and 17x fewer parameters. Analysis of our model suggests that it is particularly effective at modelling nodes with high indegree — which are common in highly-connected, complex knowledge graphs such as Freebase and YAGO3. In addition, it has been noted that the WN18 and FB15k datasets suffer from test set leakage, due to inverse relations from the training set being present in the test set — however, the extent of this issue has so far not been quantified. We find this problem to be severe: a simple rule-based model can achieve state-of-the-art results on both WN18 and FB15k. To ensure that models are evaluated on datasets where simply exploiting inverse relations cannot yield competitive results, we investigate and validate several commonly used datasets — deriving robust variants where necessary. We then perform experiments on these robust datasets for our own and several previously proposed models, and find that ConvE achieves state-of-the-art Mean Reciprocal Rank across all datasets.
Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001
AAAI4
2018 Zero-Shot Transfer Learning for Event Extraction
abstract
Most previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort.We take a fresh look at event extraction and model it as a generic grounding problem: mapping each event mention to a specific type in a target event ontology.We design a transferable architecture of structural and compositional neural networks to jointly represent and map event mentions and types into a shared semantic space.Based on this new framework, we can select, for each event mention, the event type which is semantically closest in this space as its type.By leveraging manual annotations available for a small set of existing event types, our framework can be applied to new unseen event types without additional manual annotations.When tested on 23 unseen event types, this zeroshot framework, without manual annotations, achieves performance comparable to a supervised model trained from 3,000 sentences annotated with 500 event mentions.1
Lifu Huang, Heng Ji 0001, Kyunghyun Cho, Ido Dagan, Sebastian Riedel 0001, Clare R. Voss
ACL (1)5
2018 Numeracy for Language Models: Evaluating and Improving their Ability to Predict Numbers
abstract
Numeracy is the ability to understand and work with numbers.It is a necessary skill for composing and understanding documents in clinical, scientific, and other technical domains.In this paper, we explore different strategies for modelling numerals with language models, such as memorisation and digit-by-digit composition, and propose a novel neural architecture that uses a continuous probability density function to model numerals from an open vocabulary.Our evaluation on clinical and scientific datasets shows that using hierarchical models to distinguish numerals from words improves a perplexity metric on the subset of numerals by 2 and 4 orders of magnitude, respectively, over nonhierarchical models.A combination of strategies can further improve perplexity.Our continuous probability density function model reduces mean absolute percentage errors by 18% and 54% in comparison to the second best strategy for each dataset, respectively.
Georgios Spithourakis, Sebastian Riedel 0001
ACL (1)2
2018 Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge
abstract
Adversarial examples are inputs to machine learning models designed to cause the model to make a mistake.They are useful for understanding the shortcomings of machine learning models, interpreting their results, and for regularisation.In NLP, however, most example generation strategies produce input text by using known, pre-specified semantic transformations, requiring significant manual effort and in-depth understanding of the problem and domain.In this paper, we investigate the problem of automatically generating adversarial examples that violate a set of given First-Order Logic constraints in Natural Language Inference (NLI).We reduce the problem of identifying such adversarial examples to a combinatorial optimisation problem, by maximising a quantity measuring the degree of violation of such constraints and by using a language model for generating linguisticallyplausible examples.Furthermore, we propose a method for adversarially regularising neural NLI models for incorporating background knowledge.Our results show that, while the proposed method does not always improve results on the SNLI and MultiNLI datasets, it significantly and consistently increases the predictive accuracy on adversarially-crafted datasets -up to a 79.6% relative improvement -while drastically reducing the number of background knowledge violations.Furthermore, we show that adversarial examples transfer among model architectures, and that the proposed adversarial training procedure improves the robustness of NLI models to adversarial examples.
Pasquale Minervini, Sebastian Riedel 0001
CoNLL2
2018 Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection
abstract
Grammatical error correction, like other machine learning tasks, greatly benefits from large quantities of high quality training data, which is typically expensive to produce.While writing a program to automatically generate realistic grammatical errors would be difficult, one could learn the distribution of naturallyoccurring errors and attempt to introduce them into other datasets.Initial work on inducing errors in this way using statistical machine translation has shown promise; we investigate cheaply constructing synthetic samples, given a small corpus of human-annotated data, using an off-the-rack attentive sequence-to-sequence model and a straight-forward post-processing procedure.Our approach yields error-filled artificial data that helps a vanilla bi-directional LSTM to outperform the previous state of the art at grammatical error detection, and a previously introduced model to gain further improvements of over 5% F 0.5 score.When attempting to determine if a given sentence is synthetic, a human annotator at best achieves 39.39 F 1 score, indicating that our model generates mostly human-like instances.
Sudhanshu Kasewa, Pontus Stenetorp, Sebastian Riedel 0001
EMNLP3
2018 Interpretation of Natural Language Rules in Conversational Machine Reading
abstract
Marzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Marzieh Saeidi, Max Bartolo, Patrick S. H. Lewis, Sameer Singh 0001, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel 0001
EMNLP8
2018 Behavior Analysis of NLI Models: Uncovering the Influence of Three Factors on Robustness
abstract
Ivan Sanchez, Jeff Mitchell, Sebastian Riedel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Vicente Iván Sánchez Carmona, Jeff Mitchell 0001, Sebastian Riedel 0001
NAACL-HLT3
2018 Constructing Datasets for Multi-hop Reading Comprehension Across Documents
abstract
Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence — effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement.
Johannes Welbl, Pontus Stenetorp, Sebastian Riedel 0001
Trans. Assoc. Comput. Linguistics3
2017 A Supervised Approach to Extractive Summarisation of Scientific Papers
abstract
Automatic summarisation is a popular approach to reduce a document to its main arguments.Recent research in the area has focused on neural approaches to summarisation, which can be very data-hungry.However, few large datasets exist and none for the traditionally popular domain of scientific publications, which opens up challenging research avenues centered on encoding large, complex documents.In this paper, we introduce a new dataset for summarisation of computer science publications by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.We develop models on the dataset making use of both neural sentence encoding and traditionally used summarisation features and show that models which encode sentences as well as their local and global context perform best, significantly outperforming well-established baseline methods.
Ed Collins, Isabelle Augenstein, Sebastian Riedel 0001
CoNLL3
2017 Neural Architectures for Fine-grained Entity Type Classification
abstract
Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel 0001
EACL (1)4
2017 Frustratingly Short Attention Spans in Neural Language Modeling
Michal Daniluk, Tim Rocktäschel, Johannes Welbl, Sebastian Riedel 0001
ICLR (Poster)4
2017 Programming with a Differentiable Forth Interpreter
abstract
Given that in practice training data is scarce for all but a small set of problems, a core question is how to incorporate prior knowledge into a model. In this paper, we consider the case of prior procedural knowledge for neural networks, such as knowing how a program should traverse a sequence, but not what local actions should be performed at each step. To this end, we present an end-to-end differentiable interpreter for the programming language Forth which enables programmers to write program sketches with slots that can be filled with behaviour trained from program input-output data. We can optimise this behaviour directly through gradient descent techniques on user-specified objectives, and also integrate the program into any larger neural computation graph. We show empirically that our interpreter is able to effectively leverage different levels of prior program structure and learn complex behaviours such as sequence sorting and addition. When connected to outputs of an LSTM and trained jointly, our interpreter achieves state-of-the-art accuracy for end-to-end reasoning about quantities expressed in natural language stories.
Matko Bosnjak, Tim Rocktäschel, Jason Naradowsky, Sebastian Riedel 0001
ICML4
2017 End-to-end Differentiable Proving
abstract
We introduce deep neural networks for end-to-end differentiable theorem proving that operate on dense vector representations of symbols. These neural networks are recursively constructed by following the backward chaining algorithm as used in Prolog. Specifically, we replace symbolic unification with a differentiable computation on vector representations of symbols using a radial basis function kernel, thereby combining symbolic reasoning with learning subsymbolic vector representations. The resulting neural network can be trained to infer facts from a given incomplete knowledge base using gradient descent. By doing so, it learns to (i) place representations of similar symbols in close proximity in a vector space, (ii) make use of such similarities to prove facts, (iii) induce logical rules, and (iv) it can use provided and induced logical rules for complex multi-hop reasoning. On four benchmark knowledge bases we demonstrate that this architecture outperforms ComplEx, a state-of-the-art neural link prediction model, while at the same time inducing interpretable function-free first-order logic rules.
Tim Rocktäschel, Sebastian Riedel 0001
NIPS2
2017 Adversarial Sets for Regularising Neural Link Predictors
Pasquale Minervini, Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001
UAI4
2017 Knowledge Graph Completion via Complex Tensor Factorization
abstract
In statistical relational learning, knowledge graph completion deals with automatically understanding the structure of large knowledge graphs---labeled directed graphs---and predicting missing relationships---labeled edges. State-of-the-art embedding models propose different trade-offs between modeling expressiveness, and time and space complexity. We reconcile both expressiveness and complexity through the use of complex-valued embeddings and explore the link between such complex-valued embeddings and unitary diagonalization. We corroborate our approach theoretically and show that all real square matrices---thus all possible relation/adjacency matrices---are the real part of some unitarily diagonalizable matrix. This results opens the door to a lot of other applications of square matrices factorization. Our approach based on complex embeddings is arguably simple, as it only involves a Hermitian dot product, the complex counterpart of the standard dot product between real vectors, whereas other methods resort to more and more complicated composition functions to increase their expressiveness. The proposed complex embeddings are scalable to large data sets as it remains linear in both space and time, while consistently outperforming alternative approaches on standard link prediction benchmarks.
Théo Trouillon, Christopher R. Dance, Éric Gaussier, Johannes Welbl, Sebastian Riedel 0001, Guillaume Bouchard
J. Mach. Learn. Res.5
2016 Creating Interactive and Visual Educational Resources for AI
abstract
Teaching artificial intelligence is effective if the experience is a visual and interactive one, with educational materials that utilize combinations of various content types such as text, math, and code into an integrated experience. Unfortunately, easy-to-use tools for creating such pedagogical resources are not available to the educators, resulting in most courses being taught using a disconnected set of static materials, which is not only ineffective for learning AI, but further, requires repeated and redundant effort for the instructor. In this paper, we introduce Moro, a software tool for easily creating and presenting AI-friendly teaching materials. Moro notebooks integrate content of different types (text, math, code, images), allow real-time interactions via modifiable and executable code blocks, and are viewable in browsers both as long-form pages and as presentations. Creating notebooks is easy and intuitive; the creation tool is also in-browser, is WYSIWYG for quick iterations of editing, and supports a variety of shortcuts and customizations for efficiency. We present three deployed case studies of Moro that widely differ from each other, demonstrating its utility in a variety of scenarios such as in-class teaching and conference tutorials.
Sameer Singh 0001, Sebastian Riedel 0001
AAAI2
2016 SentiHood: Targeted Aspect Based Sentiment Analysis Dataset for Urban Neighbourhoods
abstract
In this paper, we introduce the task of targeted aspect-based sentiment analysis. The goal is to extract fine-grained information with respect to entities mentioned in user comments. This work extends both aspect-based sentiment analysis – that assumes a single entity per document — and targeted sentiment analysis — that assumes a single sentiment towards a target entity. In particular, we identify the sentiment towards each aspect of one or more entities. As a testbed for this task, we introduce the SentiHood dataset, extracted from a question answering (QA) platform where urban neighbourhoods are discussed by users. In this context units of text often mention several aspects of one or more neighbourhoods. This is the first time that a generic social media platform,i.e. QA, is used for fine-grained opinion mining. Text coming from QA platforms are far less constrained compared to text from review specific platforms which current datasets are based on. We develop several strong baselines, relying on logistic regression and state-of-the-art recurrent neural networks
Marzieh Saeidi, Guillaume Bouchard, Maria Liakata, Sebastian Riedel 0001
COLING4
2016 Learning to Generate Textual Data
abstract
To learn text understanding models with millions of parameters one needs massive amounts of data.In this work, we argue that generating data can compensate for this need.While defining generic data generators is difficult, we propose to allow generators to be "weakly" specified in the sense that a set of parameters controls how the data is generated.Consider for example generators where the example templates, grammar, and/or vocabulary is determined by this set of parameters.Instead of manually tuning these parameters, we learn them from the limited training data at our disposal.To achieve this, we derive an efficient algorithm called GENERE that jointly estimates the parameters of the model and the undetermined generation parameters.We illustrate its benefits by learning to solve math exam questions using a highly parametrized sequence-to-sequence neural network.
Guillaume Bouchard, Pontus Stenetorp, Sebastian Riedel 0001
EMNLP3
2016 Lifted Rule Injection for Relation Embeddings
abstract
Methods based on representation learning currently hold the state-of-the-art in many natural language processing and knowledge base inference tasks.Yet, a major challenge is how to efficiently incorporate commonsense knowledge into such models.A recent approach regularizes relation and entity representations by propositionalization of first-order logic rules.However, propositionalization does not scale beyond domains with only few entities and rules.In this paper we present a highly efficient method for incorporating implication rules into distributed representations for automated knowledge base construction.We map entity-tuple embeddings into an approximately Boolean space and encourage a partial ordering over relation embeddings based on implication rules mined from WordNet.Surprisingly, we find that the strong restriction of the entity-tuple embedding space does not hurt the expressiveness of the model and even acts as a regularizer that improves generalization.By incorporating few commonsense rules, we achieve an increase of 2 percentage points mean average precision over a matrix factorization baseline, while observing a negligible increase in runtime.
Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001
EMNLP3
2016 Numerically Grounded Language Models for Semantic Error Correction
abstract
Semantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction.Current approaches generally focus on relatively shallow semantics and do not account for numeric quantities.Our approach uses language models grounded in numbers within the text.Such groundings are easily achieved for recurrent neural language model architectures, which can be further conditioned on incomplete background knowledge bases.Our evaluation on clinical reports shows that numerical grounding improves perplexity by 33% and F1 for semantic error correction by 5 points when compared to ungrounded approaches.Conditioning on a knowledge base yields further improvements.
Georgios Spithourakis, Isabelle Augenstein, Sebastian Riedel 0001
EMNLP3
2016 Complex Embeddings for Simple Link Prediction
abstract
In statistical relational learning, the link prediction problem is key to automatically understand the structure of large knowledge bases. As in previous studies, we propose to solve this problem through latent factorization. However, here we make use of complex valued embeddings. The composition of complex embeddings can handle a large variety of binary relations, among them symmetric and antisymmetric relations. Compared to state-of-the-art models such as Neural Tensor Network and Holographic Embeddings, our approach based on complex embeddings is arguably simpler, as it only uses the Hermitian dot product, the complex counterpart of the standard dot product between real vectors. Our approach is scalable to large datasets as it remains linear in both space and time, while consistently outperforming alternative approaches on standard link prediction benchmarks.
Théo Trouillon, Johannes Welbl, Sebastian Riedel 0001, Éric Gaussier, Guillaume Bouchard
ICML3
2015 Identification and Verification of Simple Claims about Statistical Properties
abstract
In this paper we study the identification and verification of simple claims about statistical properties, e.g.claims about the population or the inflation rate of a country.We show that this problem is similar to extracting numerical information from text and following recent work, instead of annotating data for each property of interest in order to learn supervised models, we develop a distantly supervised baseline approach using a knowledge base and raw text.In experiments on 16 statistical properties about countries from Freebase we show that our approach identifies simple statistical claims about properties with 60% precision, while it is able to verify these claims without requiring any explicit supervision for either tasks.Furthermore, we evaluate our approach as a statistical property extractor and we show it achieves 0.11 mean absolute percentage error.
Andreas Vlachos 0001, Sebastian Riedel 0001
EMNLP2
2015 Injecting Logical Background Knowledge into Embeddings for Relation Extraction
abstract
Matrix factorization approaches to relation extraction provide several attractive features: they support distant supervision, handle open schemas, and leverage unlabeled data.Unfortunately, these methods share a shortcoming with all other distantly supervised approaches: they cannot learn to extract target relations without existing data in the knowledge base, and likewise, these models are inaccurate for relations with sparse data.Rule-based extractors, on the other hand, can be easily extended to novel relations and improved for existing but inaccurate relations, through first-order formulae that capture auxiliary domain knowledge.However, usually a large set of such formulae is necessary to achieve generalization.In this paper, we introduce a paradigm for learning low-dimensional embeddings of entity-pairs and relations that combine the advantages of matrix factorization with first-order logic domain knowledge.We introduce simple approaches for estimating such embeddings, as well as a novel training algorithm to jointly optimize over factual and first-order logic information.Our results show that this method is able to learn accurate extractors with little or no distant supervision alignments, while at the same time generalizing to textual patterns that do not appear in the formulae.
Tim Rocktäschel, Sameer Singh 0001, Sebastian Riedel 0001
HLT-NAACL3
2015 WOLFE: An NLP-friendly Declarative Machine Learning Stack
abstract
Sameer Singh, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015.
Sameer Singh 0001, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel 0001
HLT-NAACL5
2014 Message Passing for Soft Constraint Dual Decomposition
David Belanger 0002, Alexandre Tachard Passos, Sebastian Riedel 0001, Andrew McCallum
UAI3
2013 AKBC 2013: third workshop on automated knowledge base construction
abstract
The AKBC 2013 workshop aims to be a venue of excellence and vision in the area of knowledge base construction. This year's workshop will feature keynotes by ten leading researchers in the field, including from Google, Microsoft, Stanford, and CMU. The submissions focus on visionary ideas instead of on experimental evaluation. Nineteen accepted papers will be presented as posters, with nine exceptional papers also highlighted as spotlight talks. Thereby, the workshop aims provides a vivid forum of discussion about the field of automated knowledge base construction.
Fabian M. Suchanek, Sebastian Riedel 0001, Sameer Singh 0001, Partha P. Talukdar
CIKM2
2013 Relation Extraction with Matrix Factorization and Universal Schemas
Sebastian Riedel 0001, Limin Yao, Andrew McCallum, Benjamin M. Marlin
HLT-NAACL1
2013 Automorphism Groups of Graphical Models and Lifted Variational Inference
Hung Hai Bui, Tuyen N. Huynh, Sebastian Riedel 0001
UAI3
2012 Unsupervised Relation Discovery with Sense Disambiguation
Limin Yao, Sebastian Riedel 0001, Andrew McCallum
ACL (1)2
2012 Improving NLP through Marginalization of Hidden Syntactic Structure
Jason Naradowsky, Sebastian Riedel 0001, David A. Smith
EMNLP-CoNLL2
2012 Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers
Sebastian Riedel 0001, David A. Smith, Andrew McCallum
EMNLP-CoNLL1
2012 MAP Inference in Chains using Column Generation
abstract
Linear chains and trees are basic building blocks in many applications of graphical models. Although exact inference in these models can be performed by dynamic programming, this computation can still be prohibitively expensive with non-trivial target variable domain sizes due to the quadratic dependence on this size. Standard message-passing algorithms for these problems are inefficient because they compute scores on hypotheses for which there is strong negative local evidence. For this reason there has been significant previous interest in beam search and its variants; however, these methods provide only approximate inference. This paper presents new efficient exact inference algorithms based on the combination of it column generation and pre-computed bounds on the model's cost structure. Improving worst-case performance is impossible. However, our method substantially speeds real-world, typical-case inference in chains and trees. Experiments show our method to be twice as fast as exact Viterbi for Wall Street Journal part-of-speech tagging and over thirteen times faster for a joint part-of-speed and named-entity-recognition task. Our algorithm is also extendable to new techniques for approximate inference, to faster two-best inference, and new opportunities for connections between inference and learning.
David Belanger 0002, Alexandre Tachard Passos, Sebastian Riedel 0001, Andrew McCallum
NIPS3
2012 Combining joint models for biomedical event extraction
abstract
BACKGROUND: We explore techniques for performing model combination between the UMass and Stanford biomedical event extraction systems. Both sub-components address event extraction as a structured prediction problem, and use dual decomposition (UMass) and parsing algorithms (Stanford) to find the best scoring event structure. Our primary focus is on stacking where the predictions from the Stanford system are used as features in the UMass system. For comparison, we look at simpler model combination techniques such as intersection and union which require only the outputs from each system and combine them directly. RESULTS: First, we find that stacking substantially improves performance while intersection and union provide no significant benefits. Second, we investigate the graph properties of event structures and their impact on the combination of our systems. Finally, we trace the origins of events proposed by the stacked model to determine the role each system plays in different components of the output. We learn that, while stacking can propose novel event structures not seen in either base model, these events have extremely low precision. Removing these novel events improves our already state-of-the-art F1 to 56.6% on the test set of Genia (Task 1). Overall, the combined system formed via stacking ("FAUST") performed well in the BioNLP 2011 shared task. The FAUST system obtained 1st place in three out of four tasks: 1st place in Genia Task 1 (56.0% F1) and Task 2 (53.9%), 2nd place in the Epigenetics and Post-translational Modifications track (35.0%), and 1st place in the Infectious Diseases track (55.6%). CONCLUSION: We present a state-of-the-art event extraction system that relies on the strengths of structured prediction and model combination through stacking. Akin to results on other tasks, stacking outperforms intersection and union and leads to very strong results. The utility of model combination hinges on complementary views of the data, and we show that our sub-systems capture different graph properties of event structures. Finally, by removing low precision novel events, we show that performance from stacking can be further improved.
David McClosky, Sebastian Riedel 0001, Mihai Surdeanu, Andrew McCallum, Christopher D. Manning
BMC Bioinform.2
2011 Fast and Robust Joint Models for Biomedical Event Extraction
Sebastian Riedel 0001, Andrew McCallum
EMNLP1
2011 Structured Relation Discovery using Generative Models
Limin Yao, Aria Haghighi, Sebastian Riedel 0001, Andrew McCallum
EMNLP3
2011 U-Compare bio-event meta-service: compatible BioNLP event extraction services
abstract
BACKGROUND: Bio-molecular event extraction from literature is recognized as an important task of bio text mining and, as such, many relevant systems have been developed and made available during the last decade. While such systems provide useful services individually, there is a need for a meta-service to enable comparison and ensemble of such services, offering optimal solutions for various purposes. RESULTS: We have integrated nine event extraction systems in the U-Compare framework, making them intercompatible and interoperable with other U-Compare components. The U-Compare event meta-service provides various meta-level features for comparison and ensemble of multiple event extraction systems. Experimental results show that the performance improvements achieved by the ensemble are significant. CONCLUSIONS: While individual event extraction systems themselves provide useful features for bio text mining, the U-Compare meta-service is expected to improve the accessibility to the individual systems, and to enable meta-level uses over multiple event extraction systems such as comparison and ensemble.
Yoshinobu Kano, Jari Björne, Filip Ginter, Tapio Salakoski, Ekaterina Buyko, Udo Hahn, Kevin Cohen 0001, Karin Verspoor, Christophe Roeder, Lawrence Hunter, Halil Kilicoglu, Sabine Bergler, Sofie Van Landeghem, Thomas Van Parys, Yves Van de Peer, Makoto Miwa, Sophia Ananiadou, Mariana L. Neves, Alberto D. Pascual-Montano, Arzucan Özgür, Dragomir R. Radev, Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Jin-Dong Kim, Sampo Pyysalo, Tomoko Ohta, Jun'ichi Tsujii
BMC Bioinform.22
2011 Bio-molecular Event Extraction with Markov Logic
abstract
This article presents a novel approach to event extraction from biological text using Markov Logic. It can be described by three design decisions: (1) instead of building a pipeline using local classifiers, we design and learn a joint probabilistic model over events in a sentence; (2) instead of developing specific inference and learning algorithms for our joint model, we apply Markov Logic, a general purpose Statistical Relation Learning language, for this task; (3) we represent events as relations over the token indices of a sentence, as opposed to structures that relate event entities to gene or protein mentions. In this article, we extend our original work by providing an error analysis for binding events. Moreover, we investigate the impact of different loss functions to precision, recall and F-measure. Finally, we show how to extract events of different types that share the same event clue. This extension allowed us to improve our performance our performance even further, leading to the third best scores for task 1 (in close range to the second place) and the best results for task 2 with a 14% point margin.
Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Toshihisa Takagi, Jun'ichi Tsujii
Comput. Intell.1
2010 Collective Cross-Document Relation Extraction Without Labelled Data
Limin Yao, Sebastian Riedel 0001, Andrew McCallum
EMNLP2
2010 Relaxed Marginal Inference and its Application to Dependency Parsing
Sebastian Riedel 0001, David A. Smith
HLT-NAACL1
2010 Constraint-Driven Rank-Based Learning for Information Extraction
Sameer Singh 0001, Limin Yao, Sebastian Riedel 0001, Andrew McCallum
HLT-NAACL3
2010 Modeling Relations and Their Mentions without Labeled Text
Sebastian Riedel 0001, Limin Yao, Andrew McCallum
ECML/PKDD (3)1
2010 Inference by Minimizing Size, Divergence, or their Sum
Sebastian Riedel 0001, David A. Smith, Andrew McCallum
UAI1
2009 Jointly Identifying Temporal Relations with Markov Logic
Katsumasa Yoshikawa, Sebastian Riedel 0001, Masayuki Asahara, Yuji Matsumoto 0001
ACL/IJCNLP2
2009 Jointly Identifying Predicates, Arguments and Senses using Markov Logic
Iván V. Meza, Sebastian Riedel 0001
HLT-NAACL2
2008 Collective Semantic Role Labelling with Markov Logic
Sebastian Riedel 0001, Iván V. Meza
CoNLL1
2008 Accurate statistical spoken language understanding from limited development resources
abstract
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
Iván V. Meza, Sebastian Riedel 0001, Oliver Lemon
ICASSP2
2008 Improving the Accuracy and Efficiency of MAP Inference for Markov Logic
Sebastian Riedel 0001
UAI1
2007 The CoNLL 2007 Shared Task on Dependency Parsing
Joakim Nivre, Johan Hall, Sandra Kübler, Ryan T. McDonald, Jens Nilsson 0001, Sebastian Riedel 0001, Deniz Yuret
EMNLP-CoNLL6
2006 Multi-lingual Dependency Parsing with Incremental Integer Linear Programming
Sebastian Riedel 0001, Ruken Cakici, Iván V. Meza
CoNLL1
2006 Incremental Integer Linear Programming for Non-projective Dependency Parsing
Sebastian Riedel 0001, James Clarke
EMNLP1