EDBT 2026 Demo / reviewers in the wild / expert
Sebastian Riedel 0001
dblp:18/3348-1
· DBLP profile ↗
90ranked-venue papers
11as first author
27since 2021 · last 2024
0000-0002-3655-2486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 86 · 11 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Do Large Language Models Latently Perform Multi-Hop Reasoning?abstractWe study whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of 'Superstition' is".We look for evidence of a latent reasoning pathway where an LLM (1) latently identifies "the singer of 'Superstition"' as Stevie Wonder, the bridge entity, and (2) uses its knowledge of Stevie Wonder's mother to complete the prompt.We analyze these two hops individually and consider their co-occurrence as indicative of latent multi-hop reasoning.For the first hop, we test if changing the prompt to indirectly mention the bridge entity instead of any other entity increases the LLM's internal recall of the bridge entity.For the second hop, we test if increasing this recall causes the LLM to better utilize what it knows about the bridge entity.We find strong evidence of latent multi-hop reasoning for the prompts of certain relation types, with the reasoning pathway used in more than 80% of the prompts.However, the utilization is highly contextual, varying across different types of prompts.Also, on average, the evidence for the second hop and the full multi-hop traversal is rather moderate and only substantial for the first hop.Moreover, we find a clear scaling trend with increasing model size for the first hop of reasoning but not for the second hop.Our experimental findings suggest potential challenges and opportunities for future development and applications of LLMs. Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel 0001 |
ACL (1) | 5 |
| 2024 | Strings from the Library of Babel: Random Sampling as a Strong Baseline for Prompt OptimisationabstractYao Lu, Jiayi Wang, Raphael Tang, Sebastian Riedel, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiayi Wang 0010, Raphael Tang, Sebastian Riedel 0001, Pontus Stenetorp |
NAACL-HLT | 4 |
| 2023 | Can discrete information extraction prompts generalize across language models?
Nathanaël Carraz Rakotonirina, Roberto Dessì, Fabio Petroni, Sebastian Riedel 0001, Marco Baroni |
ICLR | 4 |
| 2023 | PEER: A Collaborative Language Model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick S. H. Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, Sebastian Riedel 0001 |
ICLR | 10 |
| 2023 | Improving Language Plasticity via Pretraining with Active ForgettingabstractPretrained language models (PLMs) are today the primary model for natural language processing. Despite their impressive downstream performance, it can be difficult to apply PLMs to new languages, a barrier to making their capabilities universally accessible. While prior work has shown it possible to address this issue by learning a new embedding layer for the new language, doing so is both data and compute inefficient. We propose to use an active forgetting mechanism during pretraining, as a simple way of creating PLMs that can quickly adapt to new languages. Concretely, by resetting the embedding layer every K updates during pretraining, we encourage the PLM to improve its ability of learning new embeddings within limited number of updates, similar to a meta-learning effect. Experiments with RoBERTa show that models pretrained with our forgetting mechanism not only demonstrate faster convergence during language adaptation, but also outperform standard ones in a low-data regime, particularly for languages that are distant from English. Code will be available at https://github.com/facebookresearch/language-model-plasticity. Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel 0001, Mikel Artetxe |
NeurIPS | 6 |
| 2023 | Atlas: Few-shot Learning with Retrieval Augmented Language ModelsabstractLarge language models have shown impressive few-shot results on a wide range of tasks. However, when knowledge is key for such results, as is the case for tasks such as question answering and fact checking, massive parameter counts to store knowledge seem to be needed. Retrieval-augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings. In this work we present Atlas, a carefully designed and pre-trained retrieval-augmented language model able to learn knowledge intensive tasks with very few training examples. We perform evaluations on a wide range of tasks, including MMLU, KILT and Natural Questions, and study the impact of the content of the document index, showing that it can easily be updated. Notably, Atlas reaches over 42% accuracy on Natural Questions using only 64 examples, outperforming a 540B parameter model by 3% despite having 50x fewer parameters. Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel 0001, Edouard Grave |
J. Mach. Learn. Res. | 9 |
| 2022 | Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityabstractWhen primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, finetuned, large, pretrained language models.We demonstrate that the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance: essentially some permutations are "fantastic" and some not.We analyse this phenomenon in detail, establishing that: it is present across model sizes (even for the largest current models), it is not related to a specific subset of samples, and that a given good permutation for one model is not transferable to another.While one could use a development set to determine which permutations are performant, this would deviate from the true fewshot setting as it requires additional annotated data.Instead, we use the generative nature of language models to construct an artificial development set and based on entropy statistics of the candidate permutations on this set, we identify performant prompts.Our method yields a 13% relative improvement for GPTfamily models across eleven different established text classification tasks. Max Bartolo, Alastair Moore, Sebastian Riedel 0001, Pontus Stenetorp |
ACL (1) | 4 |
| 2022 | EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and IndexingabstractExisting work on Entity Linking mostly assumes that the reference knowledge base is complete, and therefore all mentions can be linked.In practice this is hardly ever the case, as knowledge bases are incomplete and because novel concepts arise constantly.We introduce the temporally segmented Unknown Entity Discovery and Indexing (EDIN) -benchmark where unknown entities, that is entities not part of the knowledge base and without descriptions and labeled mentions, have to be integrated into an existing entity linking system.By contrasting EDIN with zero-shot entity linking, we provide insight on the additional challenges it poses.Building on denseretrieval based entity linking, we introduce the end-to-end EDIN-pipeline that detects, clusters, and indexes mentions of unknown entities in context.Experiments show that indexing a single embedding per entity unifying the information of multiple mentions works better than indexing mentions independently. Nora Kassner, Fabio Petroni, Mikhail Plekhanov, Sebastian Riedel 0001, Nicola Cancedda |
EMNLP | 4 |
| 2022 | An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP TasksabstractAccess to external knowledge is essential for many natural language processing tasks, such as question answering and dialogue.Existing methods often rely on a parametric model that stores knowledge in its parameters, or use a retrieval-augmented model that has access to an external knowledge source.Parametric and retrieval-augmented models have complementary strengths in terms of computational efficiency and predictive accuracy.To combine the strength of both approaches, we propose the Efficient Memory-Augmented Transformer (EMAT) -it encodes external knowledge into a key-value memory and exploits the fast maximum inner product search for memory querying.We also introduce pre-training tasks that allow EMAT to encode informative key-value representations, and to learn an implicit strategy to integrate multiple memory slots into the transformer.Experiments on various knowledge-intensive tasks such as question answering and dialogue datasets show that, simply augmenting parametric models (T5-base) using our method produces more accurate results (e.g., 25.8 → 44.3 EM on NQ) while retaining a high throughput (e.g., 1000 queries/s on NQ).Compared to retrievalaugmented models, EMAT runs substantially faster across the board and produces more accurate results on WoW and ELI5. 1 Yuxiang Wu, Yu Zhao 0043, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001 |
EMNLP | 6 |
| 2022 | Models in the Loop: Aiding Crowdworkers with Generative Annotation AssistantsabstractMax Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, Douwe Kiela. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Max Bartolo, Tristan Thrush, Sebastian Riedel 0001, Pontus Stenetorp, Robin Jia, Douwe Kiela |
NAACL-HLT | 3 |
| 2022 | Boosted Dense RetrieverabstractPatrick Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Patrick S. H. Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel 0001 |
NAACL-HLT | 6 |
| 2022 | Lifting the Curse of Multilinguality by Pre-training Modular TransformersabstractJonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe |
NAACL-HLT | 6 |
| 2022 | Autoregressive Search Engines: Generating Substrings as Document IdentifiersabstractKnowledge-intensive language tasks require NLP systems to both provide the correct answer and retrieve supporting evidence for it in a given corpus. Autoregressive language models are emerging as the de-facto standard for generating answers, with newer and more powerful systems emerging at an astonishing pace. In this paper we argue that all this (and future) progress can be directly applied to the retrieval problem with minimal intervention to the models' architecture. Previous work has explored ways to partition the search space into hierarchical structures and retrieve documents by autoregressively generating their unique identifier. In this work we propose an alternative that doesn't force any structure in the search space: using all ngrams in a passage as its possible identifiers. This setup allows us to use an autoregressive model to generate and score distinctive ngrams, that are then mapped to full passages through an efficient data structure. Empirically, we show this not only outperforms prior autoregressive approaches but also leads to an average improvement of at least 10 points over more established retrieval solutions for passage-level retrieval on the KILT benchmark, establishing new state-of-the-art downstream performance on some datasets, while using a considerably lighter memory footprint than competing systems. Code available in the supplementary materials. Pre-trained models will be made available. Michele Bevilacqua, Giuseppe Ottaviano, Patrick S. H. Lewis, Scott Yih, Sebastian Riedel 0001, Fabio Petroni |
NeurIPS | 5 |
| 2022 | ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing PerspectiveabstractFactorisation-based Models (FMs), such as DistMult, have enjoyed enduring success for Knowledge Graph Completion (KGC) tasks, often outperforming Graph Neural Networks (GNNs). However, unlike GNNs, FMs struggle to incorporate node features and generalise to unseen nodes in inductive settings. Our work bridges the gap between FMs and GNNs by proposing ReFactor GNNs. This new architecture draws upon $\textit{both}$ modelling paradigms, which previously were largely thought of as disjoint. Concretely, using a message-passing formalism, we show how FMs can be cast as GNNs by reformulating the gradient descent procedure as message-passing operations, which forms the basis of our ReFactor GNNs. Across a multitude of well-established KGC benchmarks, our ReFactor GNNs achieve comparable transductive performance to FMs, and state-of-the-art inductive performance while using an order of magnitude fewer parameters. Pushkar Mishra, Luca Franceschi 0001, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001 |
NeurIPS | 6 |
| 2022 | Multilingual Autoregressive Entity LinkingabstractAbstract We present mGENRE, a sequence-to- sequence system for the Multilingual Entity Linking (MEL) problem—the task of resolving language-specific mentions to a multilingual Knowledge Base (KB). For a mention in a given language, mGENRE predicts the name of the target entity left-to-right, token-by-token in an autoregressive fashion. The autoregressive formulation allows us to effectively cross-encode mention string and entity names to capture more interactions than the standard dot product between mention and entity vectors. It also enables fast search within a large KB even for mentions that do not appear in mention tables and with no need for large-scale vector indices. While prior MEL works use a single representation for each entity, we match against entity names of as many languages as possible, which allows exploiting language connections between source input and target name. Moreover, in a zero-shot setting on languages with no training data at all, mGENRE treats the target language as a latent variable that is marginalized at prediction time. This leads to over 50% improvements in average accuracy. We show the efficacy of our approach through extensive evaluation including experiments on three popular MEL benchmarks where we establish new state-of-the-art results. Source code available at https://github.com/facebookresearch/GENRE. Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal 0001, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel 0001, Fabio Petroni |
Trans. Assoc. Comput. Linguistics | 9 |
| 2022 | ProoFVer: Natural Logic Theorem Proving for Fact VerificationabstractAbstract Fact verification systems typically rely on neural network classifiers for veracity prediction, which lack explainability. This paper proposes ProoFVer, which uses a seq2seq model to generate natural logic-based inferences as proofs. These proofs consist of lexical mutations between spans in the claim and the evidence retrieved, each marked with a natural logic operator. Claim veracity is determined solely based on the sequence of these operators. Hence, these proofs are faithful explanations, and this makes ProoFVer faithful by construction. Currently, ProoFVer has the highest label accuracy and the second best score in the FEVER leaderboard. Furthermore, it improves by 13.21% points over the next best model on a dataset with counterfactual instances, demonstrating its robustness. As explanations, the proofs show better overlap with human rationales than attention-based highlights and the proofs help humans predict model decisions correctly more often than using the evidence directly.1 Amrith Krishna, Sebastian Riedel 0001, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Joint Verification and Reranking for Open Fact Checking Over TablesabstractMichael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, Sebastian Riedel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Scott Yih, Sebastian Riedel 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | Database reasoning over textabstractJames Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, Alon Halevy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel 0001, Alon Y. Halevy |
ACL/IJCNLP (1) | 5 |
| 2021 | Question and Answer Test-Train Overlap in Open-Domain Question Answering DatasetsabstractIdeally Open-Domain Question Answering models should exhibit a number of competencies, ranging from simply memorizing questions seen at training time, to answering novel question formulations with answers seen during training, to generalizing to completely novel questions with novel answers.However, single aggregated test set scores do not show the full picture of what capabilities models truly have.In this work, we perform a detailed study of the test sets of three popular open-domain benchmark datasets with respect to these competencies.We find that 30% of test-set questions have a near-duplicate paraphrase in their corresponding train sets.In addition, we find that 60-70% of answers in the test sets are also present in the train sets.Using these findings, we evaluate a variety of popular open-domain models to obtain greater insight into what extent they can generalize, and what drives their overall performance.We find that all models perform substantially worse on questions that cannot be memorized from train sets, with a mean absolute performance difference of 61% between repeated and nonrepeated data.Finally we show that simple nearest-neighbor models outperform a BART closed-book QA model, further highlighting the role that train set memorization plays in these benchmarks. Patrick S. H. Lewis, Pontus Stenetorp, Sebastian Riedel 0001 |
EACL | 3 |
| 2021 | Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationabstractDespite recent progress, state-of-the-art question answering models remain vulnerable to a variety of adversarial attacks. While dynamic adversarial data collection, in which a human annotator tries to write examples that fool a model-in-the-loop, can improve model robustness, this process is expensive which limits the scale of the collected data. In this work, we are the first to use synthetic adversarial data generation to make question answering models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data. Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 0001, Pontus Stenetorp, Douwe Kiela |
EMNLP (1) | 4 |
| 2021 | Autoregressive Entity Retrieval
Nicola De Cao, Gautier Izacard, Sebastian Riedel 0001, Fabio Petroni |
ICLR | 3 |
| 2021 | Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz |
ICLR | 9 |
| 2021 | Concept Matching for Low-Resource ClassificationabstractIn many applications that rely on machine learning, the availability of labelled data is a matter of primary importance. However, when tackling new tasks, labels are usually missing and must be collected from scratch by the users. In this work, we address the problem of learning classifiers when the amount of labels is very scarce. We do so by learning multiple vectors, called prototypes, that represent relevant semantic concepts for the task at hand. We propose a theoretically inspired mechanism that computes probabilities of matching between the prototypes and the input elements, and we combine these probabilities to increase the expressiveness of the classifier. Moreover, by leveraging low-cost extra annotations in the training data, a simple error-boosting technique guides the learning process and provides substantial performance improvements. Empirical results confirm the benefits of the proposed approach in both balanced and unbalanced datasets. Our methodology is thus of practical use when gathering and labelling new examples is more expensive than annotating what we already have. Federico Errica, Fabrizio Silvestri, Bora Edizel, Ludovic Denoyer, Fabio Petroni, Vassilis Plachouras, Sebastian Riedel 0001 |
IJCNN | 7 |
| 2021 | Dynabench: Rethinking Benchmarking in NLPabstractDouwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams |
NAACL-HLT | 13 |
| 2021 | KILT: a Benchmark for Knowledge Intensive Language TasksabstractFabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick S. H. Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel 0001 |
NAACL-HLT | 13 |
| 2021 | From Natural Language Processing to Neural DatabasesabstractIn recent years, neural networks have shown impressive performance gains on long-standing AI problems, such as answering queries from text and machine translation. These advances raise the question of whether neural nets can be used at the core of query processing to derive answers from facts, even when the facts are expressed in natural language. If so, it is conceivable that we could relax the fundamental assumption of database management, namely, that our data is represented as fields of a pre-defined schema. Furthermore, such technology would enable combining information from text, images, and structured data seamlessly. This paper introduces neural databases , a class of systems that use NLP transformers as localized answer derivation engines. We ground the vision in NeuralDB, a system for querying facts represented as short natural language sentences. We demonstrate that recent natural language processing models, specifically transformers, can answer select-project-join queries if they are given a set of relevant facts. However, they cannot scale to non-trivial databases nor answer set-based and aggregation queries. Based on these insights, we identify specific research challenges that are needed to build neural databases. Some of the challenges require drawing upon the rich literature in data management, and others pose new research opportunities to the NLP community. Finally, we show that with preliminary solutions, NeuralDB can already answer queries over thousands of sentences with very high accuracy. James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel 0001, Alon Y. Halevy |
Proc. VLDB Endow. | 5 |
| 2021 | PAQ: 65 Million Probably-Asked Questions and What You Can Do With ThemabstractAbstract Open-domain Question Answering models that directly leverage question-answer (QA) pairs, such as closed-book QA (CBQA) models and QA-pair retrievers, show promise in terms of speed and memory compared with conventional models which retrieve and read from text corpora. QA-pair retrievers also offer interpretable answers, a high degree of control, and are trivial to update at test time with new knowledge. However, these models fall short of the accuracy of retrieve-and-read systems, as substantially less knowledge is covered by the available QA-pairs relative to text corpora like Wikipedia. To facilitate improved QA-pair models, we introduce Probably Asked Questions (PAQ), a very large resource of 65M automatically generated QA-pairs. We introduce a new QA-pair retriever, RePAQ, to complement PAQ. We find that PAQ preempts and caches test questions, enabling RePAQ to match the accuracy of recent retrieve-and-read models, whilst being significantly faster. Using PAQ, we train CBQA models which outperform comparable baselines by 5%, but trail RePAQ by over 15%, indicating the effectiveness of explicit retrieval. RePAQ can be configured for size (under 500MB) or speed (over 1K questions per second) while retaining high accuracy. Lastly, we demonstrate RePAQ’s strength at selective QA, abstaining from answering when it is likely to be incorrect. This enables RePAQ to “back-off” to a more expensive state-of-the-art model, leading to a combined system which is both more accurate and 2x faster than the state-of-the-art model alone. Patrick S. H. Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, Sebastian Riedel 0001 |
Trans. Assoc. Comput. Linguistics | 8 |
| 2020 | Differentiable Reasoning on Large Knowledge Bases and Natural LanguageabstractReasoning with knowledge expressed in natural language and Knowledge Bases (KBs) is a major challenge for Artificial Intelligence, with applications in machine reading, dialogue, and question answering. General neural architectures that jointly learn representations and transformations of text are very data-inefficient, and it is hard to analyse their reasoning process. These issues are addressed by end-to-end differentiable reasoning systems such as Neural Theorem Provers (NTPs), although they can only be used with small-scale symbolic KBs. In this paper we first propose Greedy NTPs (GNTPs), an extension to NTPs addressing their complexity and scalability limitations, thus making them applicable to real-world datasets. This result is achieved by dynamically constructing the computation graph of NTPs and including only the most promising proof paths during inference, thus obtaining orders of magnitude more efficient models 1. Then, we propose a novel approach for jointly reasoning over KBs and textual mentions, by embedding logic facts and natural language sentences in a shared embedding space. We show that GNTPs perform on par with NTPs at a fraction of their cost while achieving competitive link prediction results on large datasets, providing explanations for predictions, and inducing interpretable models. Pasquale Minervini, Matko Bosnjak, Tim Rocktäschel, Sebastian Riedel 0001, Edward Grefenstette |
AAAI | 4 |
| 2020 | MLQA: Evaluating Cross-lingual Extractive Question AnsweringabstractQuestion answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets.Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making building QA systems that work well in other languages challenging.In order to develop such systems, it is crucial to invest in high quality multilingual evaluation benchmarks to measure progress.We present MLQA, a multi-way aligned extractive QA evaluation benchmark intended to spur research in this area.1 MLQA contains QA instances in 7 languages, English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese.MLQA has over 12K instances in English and 5K in each other language, with each instance parallel between 4 languages on average.We evaluate stateof-the-art cross-lingual models and machinetranslation-based baselines on MLQA.In all cases, transfer results are significantly behind training-language performance. Patrick S. H. Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel 0001, Holger Schwenk |
ACL | 4 |
| 2020 | TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataabstractRecent years have witnessed the burgeoning of pretrained language models (LMs) for textbased natural language (NL) understanding tasks.Such models are typically trained on free-form NL text, hence may not be suitable for tasks like semantic parsing over structured data, which require reasoning over both free-form NL questions and structured tabular data (e.g., database tables).In this paper we present TABERT, a pretrained LM that jointly learns representations for NL sentences and (semi-)structured tables.TABERT is trained on a large corpus of 26 million tables and their English contexts.In experiments, neural semantic parsers using TABERT as feature representation layers achieve new best results on the challenging weakly-supervised semantic parsing benchmark WIKITABLEQUESTIONS, while performing competitively on the text-to-SQL dataset SPIDER. 1 Graham Neubig, Scott Yih, Sebastian Riedel 0001 |
ACL | 4 |
| 2020 | Generating Fact Checking BriefsabstractAngela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, Sebastian Riedel. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos 0001, Antoine Bordes, Sebastian Riedel 0001 |
EMNLP (1) | 8 |
| 2020 | AxCell: Automatic Extraction of Results from Machine Learning PapersabstractMarcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, Robert Stojnic. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel 0001, Ross Taylor, Robert Stojnic |
EMNLP (1) | 5 |
| 2020 | Avoiding the Hypothesis-Only Bias in Natural Language Inference via Ensemble Adversarial TrainingabstractNatural Language Inference (NLI) datasets contain annotation artefacts resulting in spurious correlations between the natural language utterances and their respective entailment classes.These artefacts are exploited by neural networks even when only considering the hypothesis and ignoring the premise, leading to unwanted biases.Belinkov et al. (2019b) proposed tackling this problem via adversarial training, but this can lead to learned sentence representations that still suffer from the same biases.We show that the bias can be reduced in the sentence representations by using an ensemble of adversaries, encouraging the model to jointly decrease the accuracy of these different adversaries while fitting the data.This approach produces more robust NLI models, outperforming previous de-biasing efforts when generalised to 12 other NLI datasets (Belinkov et al., 2019a;Mahabadi et al., 2020).In addition, we find that the optimal number of adversarial classifiers depends on the dimensionality of the sentence representations, with larger sentence representations being more difficult to de-bias while benefiting from using a greater number of adversaries. Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel 0001, Tim Rocktäschel |
EMNLP (1) | 4 |
| 2020 | Scalable Zero-shot Entity Linking with Dense Entity RetrievalabstractThis paper introduces a conceptually simple, scalable, and highly effective BERT-based entity linking model, along with an extensive evaluation of its accuracy-speed trade-off.We present a two-stage zero-shot linking algorithm, where each entity is defined only by a short textual description.The first stage does retrieval in a dense space defined by a bi-encoder that independently embeds the mention context and the entity descriptions.Each candidate is then re-ranked with a crossencoder, that concatenates the mention and entity text.Experiments demonstrate that this approach is state of the art on recent zeroshot benchmarks (6 point absolute gains) and also on more established non-zero-shot evaluations (e.g.TACKBP-2010), despite its relative simplicity (e.g.no explicit entity embeddings or manually engineered mention tables).We also show that bi-encoder linking is very fast with nearest neighbour search (e.g.linking with 5.9 million candidates in 2 milliseconds), and that much of the accuracy gain from the more expensive crossencoder can be transferred to the bi-encoder via knowledge distillation.Our code and models are available at https://github. com/facebookresearch/BLINK. wikipedia dense spaceMy kids really enjoyed a ride in the Jaguar!Jaguar is the luxury vehicle brand. Jaguar_carsJaguar! is a junior roller coaster. Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel 0001, Luke Zettlemoyer |
EMNLP (1) | 4 |
| 2020 | Don't Read Too Much Into It: Adaptive Computation for Open-Domain Question AnsweringabstractMost approaches to Open-Domain Question Answering consist of a light-weight retriever that selects a set of candidate passages, and a computationally expensive reader that examines the passages to identify the correct answer.Previous works have shown that as the number of retrieved passages increases, so does the performance of the reader.However, they assume all retrieved passages are of equal importance and allocate the same amount of computation to them, leading to a substantial increase in computational cost.To reduce this cost, we propose the use of adaptive computation to control the computational budget allocated for the passages to be read.We first introduce a technique operating on individual passages in isolation which relies on anytime prediction and a per-layer estimation of an early exit probability.We then introduce SKY-LINEBUILDER, an approach for dynamically deciding on which passage to allocate computation at each step, based on a resource allocation policy trained via reinforcement learning.Our results on SQuAD-Open show that adaptive computation with global prioritisation improves over several strong static and adaptive methods, leading to a 4.3x reduction in computation while retaining 95% performance of the full model. Yuxiang Wu, Sebastian Riedel 0001, Pasquale Minervini, Pontus Stenetorp |
EMNLP (1) | 2 |
| 2020 | Learning Reasoning Strategies in End-to-End Differentiable ProvingabstractAttempts to render deep learning models interpretable, data-efficient, and robust have seen some success through hybridisation with rule-based systems, for example, in Neural Theorem Provers (NTPs). These neuro-symbolic models can induce interpretable rules and learn representations from data via back-propagation, while providing logical explanations for their predictions. However, they are restricted by their computational complexity, as they need to consider all possible proof paths for explaining a goal, thus rendering them unfit for large-scale applications. We present Conditional Theorem Provers (CTPs), an extension to NTPs that learns an optimal rule selection strategy via gradient-based optimisation. We show that CTPs are scalable and yield state-of-the-art results on the CLUTRR dataset, which tests systematic generalisation of neural models by learning to reason over smaller graphs and evaluating on larger ones. Finally, CTPs show better link prediction results on standard benchmarks in comparison with other neural-symbolic models, while being explainable. Pasquale Minervini, Sebastian Riedel 0001, Pontus Stenetorp, Edward Grefenstette, Tim Rocktäschel |
ICML | 2 |
| 2020 | Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksabstractLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline. Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal 0001, Heinrich Küttler, Mike Lewis, Scott Yih, Tim Rocktäschel, Sebastian Riedel 0001, Douwe Kiela |
NeurIPS | 11 |
| 2020 | Beat the AI: Investigating Adversarial Human Annotation for Reading ComprehensionabstractInnovations in annotation methodology have been a catalyst for Reading Comprehension (RC) datasets and models. One recent trend to challenge current RC models is to involve a model in the annotation process: Humans create questions adversarially, such that the model fails to answer them correctly. In this work we investigate this annotation methodology and apply it in three different settings, collecting a total of 36,000 samples with progressively stronger models in the annotation loop. This allows us to explore questions such as the reproducibility of the adversarial effect, transfer from data collected with varying model-in-the-loop strengths, and generalization to data collected without a model. We find that training on adversarially collected samples leads to strong generalization to non-adversarially collected datasets, yet with progressive performance deterioration with increasingly stronger models-in-the-loop. Furthermore, we find that stronger models can still learn from datasets collected with substantially weaker models-in-the-loop. When trained on data collected with a BiDAF model in the loop, RoBERTa achieves 39.9F1 on questions that it cannot answer when trained on SQuAD—only marginally lower than when trained on data collected using RoBERTa itself (41.0F1). Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel 0001, Pontus Stenetorp |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | Unsupervised Question Answering by Cloze TranslationabstractObtaining training data for Question Answering (QA) is time-consuming and resourceintensive, and existing QA datasets are only available for limited domains and languages.In this work, we explore to what extent high quality training data is actually required for Extractive QA, and investigate the possibility of unsupervised Extractive QA.We approach this problem by first learning to generate context, question and answer triples in an unsupervised manner, which we then use to synthesize Extractive QA training data automatically.To generate such triples, we first sample random context paragraphs from a large corpus of documents and then random noun phrases or named entity mentions from these paragraphs as answers.Next we convert answers in context to "fill-in-the-blank" cloze questions and finally translate them into natural questions.We propose and compare various unsupervised ways to perform cloze-tonatural question translation, including training an unsupervised NMT model using nonaligned corpora of natural questions and cloze questions as well as a rule-based approach.We find that modern QA models can learn to answer human questions surprisingly well using only synthetic training data.We demonstrate that, without using the SQuAD training data at all, our approach achieves 56.4 F1 on SQuAD v1 (64.5 F1 when the answer is a Named entity mention), outperforming early supervised models. Patrick S. H. Lewis, Ludovic Denoyer, Sebastian Riedel 0001 |
ACL (1) | 3 |
| 2019 | Language Models as Knowledge Bases?abstractFabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander Miller. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Fabio Petroni, Tim Rocktäschel, Sebastian Riedel 0001, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Convolutional 2D Knowledge Graph EmbeddingsabstractLink prediction for knowledge graphs is the task of predicting missing relationships between entities. Previous work on link prediction has focused on shallow, fast models which can scale to large knowledge graphs. However, these models learn less expressive features than deep, multi-layer models — which potentially limits performance. In this work we introduce ConvE, a multi-layer convolutional network model for link prediction, and report state-of-the-art results for several established datasets. We also show that the model is highly parameter efficient, yielding the same performance as DistMult and R-GCN with 8x and 17x fewer parameters. Analysis of our model suggests that it is particularly effective at modelling nodes with high indegree — which are common in highly-connected, complex knowledge graphs such as Freebase and YAGO3. In addition, it has been noted that the WN18 and FB15k datasets suffer from test set leakage, due to inverse relations from the training set being present in the test set — however, the extent of this issue has so far not been quantified. We find this problem to be severe: a simple rule-based model can achieve state-of-the-art results on both WN18 and FB15k. To ensure that models are evaluated on datasets where simply exploiting inverse relations cannot yield competitive results, we investigate and validate several commonly used datasets — deriving robust variants where necessary. We then perform experiments on these robust datasets for our own and several previously proposed models, and find that ConvE achieves state-of-the-art Mean Reciprocal Rank across all datasets. Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001 |
AAAI | 4 |
| 2018 | Zero-Shot Transfer Learning for Event ExtractionabstractMost previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort.We take a fresh look at event extraction and model it as a generic grounding problem: mapping each event mention to a specific type in a target event ontology.We design a transferable architecture of structural and compositional neural networks to jointly represent and map event mentions and types into a shared semantic space.Based on this new framework, we can select, for each event mention, the event type which is semantically closest in this space as its type.By leveraging manual annotations available for a small set of existing event types, our framework can be applied to new unseen event types without additional manual annotations.When tested on 23 unseen event types, this zeroshot framework, without manual annotations, achieves performance comparable to a supervised model trained from 3,000 sentences annotated with 500 event mentions.1 Lifu Huang, Heng Ji 0001, Kyunghyun Cho, Ido Dagan, Sebastian Riedel 0001, Clare R. Voss |
ACL (1) | 5 |
| 2018 | Numeracy for Language Models: Evaluating and Improving their Ability to Predict NumbersabstractNumeracy is the ability to understand and work with numbers.It is a necessary skill for composing and understanding documents in clinical, scientific, and other technical domains.In this paper, we explore different strategies for modelling numerals with language models, such as memorisation and digit-by-digit composition, and propose a novel neural architecture that uses a continuous probability density function to model numerals from an open vocabulary.Our evaluation on clinical and scientific datasets shows that using hierarchical models to distinguish numerals from words improves a perplexity metric on the subset of numerals by 2 and 4 orders of magnitude, respectively, over nonhierarchical models.A combination of strategies can further improve perplexity.Our continuous probability density function model reduces mean absolute percentage errors by 18% and 54% in comparison to the second best strategy for each dataset, respectively. Georgios Spithourakis, Sebastian Riedel 0001 |
ACL (1) | 2 |
| 2018 | Adversarially Regularising Neural NLI Models to Integrate Logical Background KnowledgeabstractAdversarial examples are inputs to machine learning models designed to cause the model to make a mistake.They are useful for understanding the shortcomings of machine learning models, interpreting their results, and for regularisation.In NLP, however, most example generation strategies produce input text by using known, pre-specified semantic transformations, requiring significant manual effort and in-depth understanding of the problem and domain.In this paper, we investigate the problem of automatically generating adversarial examples that violate a set of given First-Order Logic constraints in Natural Language Inference (NLI).We reduce the problem of identifying such adversarial examples to a combinatorial optimisation problem, by maximising a quantity measuring the degree of violation of such constraints and by using a language model for generating linguisticallyplausible examples.Furthermore, we propose a method for adversarially regularising neural NLI models for incorporating background knowledge.Our results show that, while the proposed method does not always improve results on the SNLI and MultiNLI datasets, it significantly and consistently increases the predictive accuracy on adversarially-crafted datasets -up to a 79.6% relative improvement -while drastically reducing the number of background knowledge violations.Furthermore, we show that adversarial examples transfer among model architectures, and that the proposed adversarial training procedure improves the robustness of NLI models to adversarial examples. Pasquale Minervini, Sebastian Riedel 0001 |
CoNLL | 2 |
| 2018 | Wronging a Right: Generating Better Errors to Improve Grammatical Error DetectionabstractGrammatical error correction, like other machine learning tasks, greatly benefits from large quantities of high quality training data, which is typically expensive to produce.While writing a program to automatically generate realistic grammatical errors would be difficult, one could learn the distribution of naturallyoccurring errors and attempt to introduce them into other datasets.Initial work on inducing errors in this way using statistical machine translation has shown promise; we investigate cheaply constructing synthetic samples, given a small corpus of human-annotated data, using an off-the-rack attentive sequence-to-sequence model and a straight-forward post-processing procedure.Our approach yields error-filled artificial data that helps a vanilla bi-directional LSTM to outperform the previous state of the art at grammatical error detection, and a previously introduced model to gain further improvements of over 5% F 0.5 score.When attempting to determine if a given sentence is synthetic, a human annotator at best achieves 39.39 F 1 score, indicating that our model generates mostly human-like instances. Sudhanshu Kasewa, Pontus Stenetorp, Sebastian Riedel 0001 |
EMNLP | 3 |
| 2018 | Interpretation of Natural Language Rules in Conversational Machine ReadingabstractMarzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Marzieh Saeidi, Max Bartolo, Patrick S. H. Lewis, Sameer Singh 0001, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel 0001 |
EMNLP | 8 |
| 2018 | Behavior Analysis of NLI Models: Uncovering the Influence of Three Factors on RobustnessabstractIvan Sanchez, Jeff Mitchell, Sebastian Riedel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Vicente Iván Sánchez Carmona, Jeff Mitchell 0001, Sebastian Riedel 0001 |
NAACL-HLT | 3 |
| 2018 | Constructing Datasets for Multi-hop Reading Comprehension Across DocumentsabstractMost Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence — effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement. Johannes Welbl, Pontus Stenetorp, Sebastian Riedel 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | A Supervised Approach to Extractive Summarisation of Scientific PapersabstractAutomatic summarisation is a popular approach to reduce a document to its main arguments.Recent research in the area has focused on neural approaches to summarisation, which can be very data-hungry.However, few large datasets exist and none for the traditionally popular domain of scientific publications, which opens up challenging research avenues centered on encoding large, complex documents.In this paper, we introduce a new dataset for summarisation of computer science publications by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.We develop models on the dataset making use of both neural sentence encoding and traditionally used summarisation features and show that models which encode sentences as well as their local and global context perform best, significantly outperforming well-established baseline methods. Ed Collins, Isabelle Augenstein, Sebastian Riedel 0001 |
CoNLL | 3 |
| 2017 | Neural Architectures for Fine-grained Entity Type ClassificationabstractSonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel 0001 |
EACL (1) | 4 |
| 2017 | Frustratingly Short Attention Spans in Neural Language Modeling
Michal Daniluk, Tim Rocktäschel, Johannes Welbl, Sebastian Riedel 0001 |
ICLR (Poster) | 4 |
| 2017 | Programming with a Differentiable Forth InterpreterabstractGiven that in practice training data is scarce for all but a small set of problems, a core question is how to incorporate prior knowledge into a model. In this paper, we consider the case of prior procedural knowledge for neural networks, such as knowing how a program should traverse a sequence, but not what local actions should be performed at each step. To this end, we present an end-to-end differentiable interpreter for the programming language Forth which enables programmers to write program sketches with slots that can be filled with behaviour trained from program input-output data. We can optimise this behaviour directly through gradient descent techniques on user-specified objectives, and also integrate the program into any larger neural computation graph. We show empirically that our interpreter is able to effectively leverage different levels of prior program structure and learn complex behaviours such as sequence sorting and addition. When connected to outputs of an LSTM and trained jointly, our interpreter achieves state-of-the-art accuracy for end-to-end reasoning about quantities expressed in natural language stories. Matko Bosnjak, Tim Rocktäschel, Jason Naradowsky, Sebastian Riedel 0001 |
ICML | 4 |
| 2017 | End-to-end Differentiable ProvingabstractWe introduce deep neural networks for end-to-end differentiable theorem proving that operate on dense vector representations of symbols. These neural networks are recursively constructed by following the backward chaining algorithm as used in Prolog. Specifically, we replace symbolic unification with a differentiable computation on vector representations of symbols using a radial basis function kernel, thereby combining symbolic reasoning with learning subsymbolic vector representations. The resulting neural network can be trained to infer facts from a given incomplete knowledge base using gradient descent. By doing so, it learns to (i) place representations of similar symbols in close proximity in a vector space, (ii) make use of such similarities to prove facts, (iii) induce logical rules, and (iv) it can use provided and induced logical rules for complex multi-hop reasoning. On four benchmark knowledge bases we demonstrate that this architecture outperforms ComplEx, a state-of-the-art neural link prediction model, while at the same time inducing interpretable function-free first-order logic rules. Tim Rocktäschel, Sebastian Riedel 0001 |
NIPS | 2 |
| 2017 | Adversarial Sets for Regularising Neural Link Predictors
Pasquale Minervini, Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001 |
UAI | 4 |
| 2017 | Knowledge Graph Completion via Complex Tensor FactorizationabstractIn statistical relational learning, knowledge graph completion deals with automatically understanding the structure of large knowledge graphs---labeled directed graphs---and predicting missing relationships---labeled edges. State-of-the-art embedding models propose different trade-offs between modeling expressiveness, and time and space complexity. We reconcile both expressiveness and complexity through the use of complex-valued embeddings and explore the link between such complex-valued embeddings and unitary diagonalization. We corroborate our approach theoretically and show that all real square matrices---thus all possible relation/adjacency matrices---are the real part of some unitarily diagonalizable matrix. This results opens the door to a lot of other applications of square matrices factorization. Our approach based on complex embeddings is arguably simple, as it only involves a Hermitian dot product, the complex counterpart of the standard dot product between real vectors, whereas other methods resort to more and more complicated composition functions to increase their expressiveness. The proposed complex embeddings are scalable to large data sets as it remains linear in both space and time, while consistently outperforming alternative approaches on standard link prediction benchmarks. Théo Trouillon, Christopher R. Dance, Éric Gaussier, Johannes Welbl, Sebastian Riedel 0001, Guillaume Bouchard |
J. Mach. Learn. Res. | 5 |
| 2016 | Creating Interactive and Visual Educational Resources for AIabstractTeaching artificial intelligence is effective if the experience is a visual and interactive one, with educational materials that utilize combinations of various content types such as text, math, and code into an integrated experience. Unfortunately, easy-to-use tools for creating such pedagogical resources are not available to the educators, resulting in most courses being taught using a disconnected set of static materials, which is not only ineffective for learning AI, but further, requires repeated and redundant effort for the instructor. In this paper, we introduce Moro, a software tool for easily creating and presenting AI-friendly teaching materials. Moro notebooks integrate content of different types (text, math, code, images), allow real-time interactions via modifiable and executable code blocks, and are viewable in browsers both as long-form pages and as presentations. Creating notebooks is easy and intuitive; the creation tool is also in-browser, is WYSIWYG for quick iterations of editing, and supports a variety of shortcuts and customizations for efficiency. We present three deployed case studies of Moro that widely differ from each other, demonstrating its utility in a variety of scenarios such as in-class teaching and conference tutorials. Sameer Singh 0001, Sebastian Riedel 0001 |
AAAI | 2 |
| 2016 | SentiHood: Targeted Aspect Based Sentiment Analysis Dataset for Urban NeighbourhoodsabstractIn this paper, we introduce the task of targeted aspect-based sentiment analysis. The goal is to extract fine-grained information with respect to entities mentioned in user comments. This work extends both aspect-based sentiment analysis – that assumes a single entity per document — and targeted sentiment analysis — that assumes a single sentiment towards a target entity. In particular, we identify the sentiment towards each aspect of one or more entities. As a testbed for this task, we introduce the SentiHood dataset, extracted from a question answering (QA) platform where urban neighbourhoods are discussed by users. In this context units of text often mention several aspects of one or more neighbourhoods. This is the first time that a generic social media platform,i.e. QA, is used for fine-grained opinion mining. Text coming from QA platforms are far less constrained compared to text from review specific platforms which current datasets are based on. We develop several strong baselines, relying on logistic regression and state-of-the-art recurrent neural networks Marzieh Saeidi, Guillaume Bouchard, Maria Liakata, Sebastian Riedel 0001 |
COLING | 4 |
| 2016 | Learning to Generate Textual DataabstractTo learn text understanding models with millions of parameters one needs massive amounts of data.In this work, we argue that generating data can compensate for this need.While defining generic data generators is difficult, we propose to allow generators to be "weakly" specified in the sense that a set of parameters controls how the data is generated.Consider for example generators where the example templates, grammar, and/or vocabulary is determined by this set of parameters.Instead of manually tuning these parameters, we learn them from the limited training data at our disposal.To achieve this, we derive an efficient algorithm called GENERE that jointly estimates the parameters of the model and the undetermined generation parameters.We illustrate its benefits by learning to solve math exam questions using a highly parametrized sequence-to-sequence neural network. Guillaume Bouchard, Pontus Stenetorp, Sebastian Riedel 0001 |
EMNLP | 3 |
| 2016 | Lifted Rule Injection for Relation EmbeddingsabstractMethods based on representation learning currently hold the state-of-the-art in many natural language processing and knowledge base inference tasks.Yet, a major challenge is how to efficiently incorporate commonsense knowledge into such models.A recent approach regularizes relation and entity representations by propositionalization of first-order logic rules.However, propositionalization does not scale beyond domains with only few entities and rules.In this paper we present a highly efficient method for incorporating implication rules into distributed representations for automated knowledge base construction.We map entity-tuple embeddings into an approximately Boolean space and encourage a partial ordering over relation embeddings based on implication rules mined from WordNet.Surprisingly, we find that the strong restriction of the entity-tuple embedding space does not hurt the expressiveness of the model and even acts as a regularizer that improves generalization.By incorporating few commonsense rules, we achieve an increase of 2 percentage points mean average precision over a matrix factorization baseline, while observing a negligible increase in runtime. Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001 |
EMNLP | 3 |
| 2016 | Numerically Grounded Language Models for Semantic Error CorrectionabstractSemantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction.Current approaches generally focus on relatively shallow semantics and do not account for numeric quantities.Our approach uses language models grounded in numbers within the text.Such groundings are easily achieved for recurrent neural language model architectures, which can be further conditioned on incomplete background knowledge bases.Our evaluation on clinical reports shows that numerical grounding improves perplexity by 33% and F1 for semantic error correction by 5 points when compared to ungrounded approaches.Conditioning on a knowledge base yields further improvements. Georgios Spithourakis, Isabelle Augenstein, Sebastian Riedel 0001 |
EMNLP | 3 |
| 2016 | Complex Embeddings for Simple Link PredictionabstractIn statistical relational learning, the link prediction problem is key to automatically understand the structure of large knowledge bases. As in previous studies, we propose to solve this problem through latent factorization. However, here we make use of complex valued embeddings. The composition of complex embeddings can handle a large variety of binary relations, among them symmetric and antisymmetric relations. Compared to state-of-the-art models such as Neural Tensor Network and Holographic Embeddings, our approach based on complex embeddings is arguably simpler, as it only uses the Hermitian dot product, the complex counterpart of the standard dot product between real vectors. Our approach is scalable to large datasets as it remains linear in both space and time, while consistently outperforming alternative approaches on standard link prediction benchmarks. Théo Trouillon, Johannes Welbl, Sebastian Riedel 0001, Éric Gaussier, Guillaume Bouchard |
ICML | 3 |
| 2015 | Identification and Verification of Simple Claims about Statistical PropertiesabstractIn this paper we study the identification and verification of simple claims about statistical properties, e.g.claims about the population or the inflation rate of a country.We show that this problem is similar to extracting numerical information from text and following recent work, instead of annotating data for each property of interest in order to learn supervised models, we develop a distantly supervised baseline approach using a knowledge base and raw text.In experiments on 16 statistical properties about countries from Freebase we show that our approach identifies simple statistical claims about properties with 60% precision, while it is able to verify these claims without requiring any explicit supervision for either tasks.Furthermore, we evaluate our approach as a statistical property extractor and we show it achieves 0.11 mean absolute percentage error. Andreas Vlachos 0001, Sebastian Riedel 0001 |
EMNLP | 2 |
| 2015 | Injecting Logical Background Knowledge into Embeddings for Relation ExtractionabstractMatrix factorization approaches to relation extraction provide several attractive features: they support distant supervision, handle open schemas, and leverage unlabeled data.Unfortunately, these methods share a shortcoming with all other distantly supervised approaches: they cannot learn to extract target relations without existing data in the knowledge base, and likewise, these models are inaccurate for relations with sparse data.Rule-based extractors, on the other hand, can be easily extended to novel relations and improved for existing but inaccurate relations, through first-order formulae that capture auxiliary domain knowledge.However, usually a large set of such formulae is necessary to achieve generalization.In this paper, we introduce a paradigm for learning low-dimensional embeddings of entity-pairs and relations that combine the advantages of matrix factorization with first-order logic domain knowledge.We introduce simple approaches for estimating such embeddings, as well as a novel training algorithm to jointly optimize over factual and first-order logic information.Our results show that this method is able to learn accurate extractors with little or no distant supervision alignments, while at the same time generalizing to textual patterns that do not appear in the formulae. Tim Rocktäschel, Sameer Singh 0001, Sebastian Riedel 0001 |
HLT-NAACL | 3 |
| 2015 | WOLFE: An NLP-friendly Declarative Machine Learning StackabstractSameer Singh, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Sameer Singh 0001, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel 0001 |
HLT-NAACL | 5 |
| 2014 | Message Passing for Soft Constraint Dual Decomposition
David Belanger 0002, Alexandre Tachard Passos, Sebastian Riedel 0001, Andrew McCallum |
UAI | 3 |
| 2013 | AKBC 2013: third workshop on automated knowledge base constructionabstractThe AKBC 2013 workshop aims to be a venue of excellence and vision in the area of knowledge base construction. This year's workshop will feature keynotes by ten leading researchers in the field, including from Google, Microsoft, Stanford, and CMU. The submissions focus on visionary ideas instead of on experimental evaluation. Nineteen accepted papers will be presented as posters, with nine exceptional papers also highlighted as spotlight talks. Thereby, the workshop aims provides a vivid forum of discussion about the field of automated knowledge base construction. Fabian M. Suchanek, Sebastian Riedel 0001, Sameer Singh 0001, Partha P. Talukdar |
CIKM | 2 |
| 2013 | Relation Extraction with Matrix Factorization and Universal Schemas
Sebastian Riedel 0001, Limin Yao, Andrew McCallum, Benjamin M. Marlin |
HLT-NAACL | 1 |
| 2013 | Automorphism Groups of Graphical Models and Lifted Variational Inference
Hung Hai Bui, Tuyen N. Huynh, Sebastian Riedel 0001 |
UAI | 3 |
| 2012 | Unsupervised Relation Discovery with Sense Disambiguation
Limin Yao, Sebastian Riedel 0001, Andrew McCallum |
ACL (1) | 2 |
| 2012 | Improving NLP through Marginalization of Hidden Syntactic Structure
Jason Naradowsky, Sebastian Riedel 0001, David A. Smith |
EMNLP-CoNLL | 2 |
| 2012 | Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers
Sebastian Riedel 0001, David A. Smith, Andrew McCallum |
EMNLP-CoNLL | 1 |
| 2012 | MAP Inference in Chains using Column GenerationabstractLinear chains and trees are basic building blocks in many applications of graphical models. Although exact inference in these models can be performed by dynamic programming, this computation can still be prohibitively expensive with non-trivial target variable domain sizes due to the quadratic dependence on this size. Standard message-passing algorithms for these problems are inefficient because they compute scores on hypotheses for which there is strong negative local evidence. For this reason there has been significant previous interest in beam search and its variants; however, these methods provide only approximate inference. This paper presents new efficient exact inference algorithms based on the combination of it column generation and pre-computed bounds on the model's cost structure. Improving worst-case performance is impossible. However, our method substantially speeds real-world, typical-case inference in chains and trees. Experiments show our method to be twice as fast as exact Viterbi for Wall Street Journal part-of-speech tagging and over thirteen times faster for a joint part-of-speed and named-entity-recognition task. Our algorithm is also extendable to new techniques for approximate inference, to faster two-best inference, and new opportunities for connections between inference and learning. David Belanger 0002, Alexandre Tachard Passos, Sebastian Riedel 0001, Andrew McCallum |
NIPS | 3 |
| 2012 | Combining joint models for biomedical event extractionabstractBACKGROUND: We explore techniques for performing model combination between the UMass and Stanford biomedical event extraction systems. Both sub-components address event extraction as a structured prediction problem, and use dual decomposition (UMass) and parsing algorithms (Stanford) to find the best scoring event structure. Our primary focus is on stacking where the predictions from the Stanford system are used as features in the UMass system. For comparison, we look at simpler model combination techniques such as intersection and union which require only the outputs from each system and combine them directly. RESULTS: First, we find that stacking substantially improves performance while intersection and union provide no significant benefits. Second, we investigate the graph properties of event structures and their impact on the combination of our systems. Finally, we trace the origins of events proposed by the stacked model to determine the role each system plays in different components of the output. We learn that, while stacking can propose novel event structures not seen in either base model, these events have extremely low precision. Removing these novel events improves our already state-of-the-art F1 to 56.6% on the test set of Genia (Task 1). Overall, the combined system formed via stacking ("FAUST") performed well in the BioNLP 2011 shared task. The FAUST system obtained 1st place in three out of four tasks: 1st place in Genia Task 1 (56.0% F1) and Task 2 (53.9%), 2nd place in the Epigenetics and Post-translational Modifications track (35.0%), and 1st place in the Infectious Diseases track (55.6%). CONCLUSION: We present a state-of-the-art event extraction system that relies on the strengths of structured prediction and model combination through stacking. Akin to results on other tasks, stacking outperforms intersection and union and leads to very strong results. The utility of model combination hinges on complementary views of the data, and we show that our sub-systems capture different graph properties of event structures. Finally, by removing low precision novel events, we show that performance from stacking can be further improved. David McClosky, Sebastian Riedel 0001, Mihai Surdeanu, Andrew McCallum, Christopher D. Manning |
BMC Bioinform. | 2 |
| 2011 | Fast and Robust Joint Models for Biomedical Event Extraction
Sebastian Riedel 0001, Andrew McCallum |
EMNLP | 1 |
| 2011 | Structured Relation Discovery using Generative Models
Limin Yao, Aria Haghighi, Sebastian Riedel 0001, Andrew McCallum |
EMNLP | 3 |
| 2011 | U-Compare bio-event meta-service: compatible BioNLP event extraction servicesabstractBACKGROUND: Bio-molecular event extraction from literature is recognized as an important task of bio text mining and, as such, many relevant systems have been developed and made available during the last decade. While such systems provide useful services individually, there is a need for a meta-service to enable comparison and ensemble of such services, offering optimal solutions for various purposes. RESULTS: We have integrated nine event extraction systems in the U-Compare framework, making them intercompatible and interoperable with other U-Compare components. The U-Compare event meta-service provides various meta-level features for comparison and ensemble of multiple event extraction systems. Experimental results show that the performance improvements achieved by the ensemble are significant. CONCLUSIONS: While individual event extraction systems themselves provide useful features for bio text mining, the U-Compare meta-service is expected to improve the accessibility to the individual systems, and to enable meta-level uses over multiple event extraction systems such as comparison and ensemble. Yoshinobu Kano, Jari Björne, Filip Ginter, Tapio Salakoski, Ekaterina Buyko, Udo Hahn, Kevin Cohen 0001, Karin Verspoor, Christophe Roeder, Lawrence Hunter, Halil Kilicoglu, Sabine Bergler, Sofie Van Landeghem, Thomas Van Parys, Yves Van de Peer, Makoto Miwa, Sophia Ananiadou, Mariana L. Neves, Alberto D. Pascual-Montano, Arzucan Özgür, Dragomir R. Radev, Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Jin-Dong Kim, Sampo Pyysalo, Tomoko Ohta, Jun'ichi Tsujii |
BMC Bioinform. | 22 |
| 2011 | Bio-molecular Event Extraction with Markov LogicabstractThis article presents a novel approach to event extraction from biological text using Markov Logic. It can be described by three design decisions: (1) instead of building a pipeline using local classifiers, we design and learn a joint probabilistic model over events in a sentence; (2) instead of developing specific inference and learning algorithms for our joint model, we apply Markov Logic, a general purpose Statistical Relation Learning language, for this task; (3) we represent events as relations over the token indices of a sentence, as opposed to structures that relate event entities to gene or protein mentions. In this article, we extend our original work by providing an error analysis for binding events. Moreover, we investigate the impact of different loss functions to precision, recall and F-measure. Finally, we show how to extract events of different types that share the same event clue. This extension allowed us to improve our performance our performance even further, leading to the third best scores for task 1 (in close range to the second place) and the best results for task 2 with a 14% point margin. Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Toshihisa Takagi, Jun'ichi Tsujii |
Comput. Intell. | 1 |
| 2010 | Collective Cross-Document Relation Extraction Without Labelled Data
Limin Yao, Sebastian Riedel 0001, Andrew McCallum |
EMNLP | 2 |
| 2010 | Relaxed Marginal Inference and its Application to Dependency Parsing
Sebastian Riedel 0001, David A. Smith |
HLT-NAACL | 1 |
| 2010 | Constraint-Driven Rank-Based Learning for Information Extraction
Sameer Singh 0001, Limin Yao, Sebastian Riedel 0001, Andrew McCallum |
HLT-NAACL | 3 |
| 2010 | Modeling Relations and Their Mentions without Labeled Text
Sebastian Riedel 0001, Limin Yao, Andrew McCallum |
ECML/PKDD (3) | 1 |
| 2010 | Inference by Minimizing Size, Divergence, or their Sum
Sebastian Riedel 0001, David A. Smith, Andrew McCallum |
UAI | 1 |
| 2009 | Jointly Identifying Temporal Relations with Markov Logic
Katsumasa Yoshikawa, Sebastian Riedel 0001, Masayuki Asahara, Yuji Matsumoto 0001 |
ACL/IJCNLP | 2 |
| 2009 | Jointly Identifying Predicates, Arguments and Senses using Markov Logic
Iván V. Meza, Sebastian Riedel 0001 |
HLT-NAACL | 2 |
| 2008 | Collective Semantic Role Labelling with Markov Logic
Sebastian Riedel 0001, Iván V. Meza |
CoNLL | 1 |
| 2008 | Accurate statistical spoken language understanding from limited development resourcesabstractRobust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction). Iván V. Meza, Sebastian Riedel 0001, Oliver Lemon |
ICASSP | 2 |
| 2008 | Improving the Accuracy and Efficiency of MAP Inference for Markov Logic
Sebastian Riedel 0001 |
UAI | 1 |
| 2007 | The CoNLL 2007 Shared Task on Dependency Parsing
Joakim Nivre, Johan Hall, Sandra Kübler, Ryan T. McDonald, Jens Nilsson 0001, Sebastian Riedel 0001, Deniz Yuret |
EMNLP-CoNLL | 6 |
| 2006 | Multi-lingual Dependency Parsing with Incremental Integer Linear Programming
Sebastian Riedel 0001, Ruken Cakici, Iván V. Meza |
CoNLL | 1 |
| 2006 | Incremental Integer Linear Programming for Non-projective Dependency Parsing
Sebastian Riedel 0001, James Clarke |
EMNLP | 1 |