VLDB 2026 Research / reviewers in the wild / expert
Douwe Kiela
dblp:136/9140
· DBLP profile ↗
76ranked-venue papers
14as first author
32since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 12 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Representational Instruction TuningabstractAll text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM-7B is among the top models on the Massive Text Embedding Benchmark (MTEB) and outperforms various models up to its size on a range of generative tasks. By scaling up further, GritLM-8x7B achieves even stronger generative performance while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm. Niklas Muennighoff, Hongjin Su, Liang Wang 0046, Nan Yang 0002, Furu Wei, Tao Yu 0009, Amanpreet Singh, Douwe Kiela |
ICLR | 8 |
| 2025 | OLMoE: Open Mixture-of-Experts Language ModelsabstractWe introduce OLMoE, a fully open, state-of-the-art language model leveraging sparse Mixture-of-Experts (MoE). OLMoE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMoE-1B-7B-Instruct. Our models outperform all available models with similar active parameters, even surpassing larger ones like Llama2-13B-Chat and DeepSeekMoE-16B. We present novel findings on MoE training, define and analyze new routing properties showing high specialization in our model, and open-source all our work: model weights, training data, code, and logs. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Pete Walsh 0001, Oyvind Tafjord, Nathan Lambert 0001, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, Dave Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi |
ICLR | 19 |
| 2025 | Great Models Think Alike and this Undermines AI OversightabstractAs Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as AI Oversight. We study how model similarity affects both aspects of AI oversight by proposing Chance Adjusted Probabilistic Agreement (CAPA)–a metric for LM similarity based on overlap in model mistakes. Using CAPA, we first show that LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results. Then, we study training on LM annotations, and find complementary knowledge between the weak supervisor and strong student model plays a crucial role in gains from weak-to-strong generalization. As model capabilities increase, it becomes harder to find their mistakes, and we might defer more to AI oversight. However, we observe a concerning trend–model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures. Our work underscores the importance of reporting and correcting for model similarity, especially in the emerging paradigm of AI oversight. Shashwat Goel, Joschka Strüber, Ilze Amanda Auzina, Karuna K. Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, Jonas Geiping |
ICML | 6 |
| 2025 | Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in AlignmentabstractAbstract Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes alignment a complicated procedure, sometimes producing subpar results. We study this and find that (i) preference data gives a better learning signal when the underlying responses are contrastive, and (ii) alignment objectives lead to better performance when they specify more control over the model during training. Based on these insights, we introduce Contrastive Learning from AI Revisions (CLAIR), a data-creation method which leads to more contrastive preference pairs, and Anchored Preference Optimization (APO), a controllable and more stable alignment objective. We align Llama-3-8B-Instruct using various comparable datasets and alignment objectives and measure MixEval-Hard scores, which correlate highly with human judgments. The CLAIR preferences lead to the strongest performance out of all datasets, and APO consistently outperforms less controllable objectives. Our best model, trained on 32K CLAIR preferences with APO, improves Llama-3-8B-Instruct by 7.65%, closing the gap with GPT4-turbo by 45%. Our code and datasets are available. Karel D'Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, Shikib Mehri |
Trans. Assoc. Comput. Linguistics | 7 |
| 2024 | Leveraging Diffusion Perturbations for Measuring Fairness in Computer VisionabstractComputer vision models have been known to encode harmful biases, leading to the potentially unfair treatment of historically marginalized groups, such as people of color. However, there remains a lack of datasets balanced along demographic traits that can be used to evaluate the downstream fairness of these models. In this work, we demonstrate that diffusion models can be leveraged to create such a dataset. We first use a diffusion model to generate a large set of images depicting various occupations. Subsequently, each image is edited using inpainting to generate multiple variants, where each variant refers to a different perceived race. Using this dataset, we benchmark several vision-language models on a multi-class occupation classification task. We find that images generated with non-Caucasian labels have a significantly higher occupation misclassification rate than images generated with Caucasian labels, and that several misclassifications are suggestive of racial biases. We measure a model’s downstream fairness by computing the standard deviation in the probability of predicting the true occupation label across the different identity groups. Using this fairness metric, we find significant disparities between the evaluated vision-and-language models. We hope that our work demonstrates the potential value of diffusion methods for fairness evaluations. Nicholas Lui, Bryan Chia, William Berrios, Candace Ross, Douwe Kiela |
AAAI | 5 |
| 2024 | I am a Strange Dataset: Metalinguistic Tests for Language ModelsabstractStatements involving metalinguistic selfreference ("This paper has six sections.")are prevalent in many domains.Can current large language models (LLMs) handle such language?In this paper, we present "I am a Strange Dataset", a new dataset for addressing this question.There are two subtasks: generation and verification.In generation, models continue statements like "The penultimate word in this sentence is" (where a correct continuation is "is").In verification, models judge the truth of statements like "The penultimate word in this sentence is sentence."(false).We also provide minimally different metalinguistic non-self-reference examples to complement the main dataset by probing for whether models can handle metalinguistic language at all.The dataset is hand-crafted by experts and validated by non-expert annotators.We test a variety of open-source LLMs (7B to 70B parameters) as well as closed-source LLMs through APIs.All models perform close to chance across both subtasks and even on the non-self-referential metalinguistic control data, though we find some steady improvement with model scale.GPT 4 is the only model to consistently do significantly better than chance, and it is still only in the 60% range, while our untrained human annotators score well in the 89-93% range.The dataset and evaluation toolkit are available at https://github.com/ TristanThrush/i-am-a-strange-dataset. Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, Douwe Kiela |
ACL (1) | 5 |
| 2024 | Anchor Points: Benchmarking Models with Much Fewer ExamplesabstractModern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior.Here, we suggest that model performance can be benchmarked and elucidated with much smaller evaluation sets.We first show that in six popular language classification benchmarks, model confidence in the correct class on many pairs of points is strongly correlated across models.We build upon this phenomenon to propose Anchor Point Selection, a technique to select small subsets of datasets that capture model behavior across the entire dataset.Anchor points reliably rank models: across 87 diverse language model-prompt pairs, evaluating models using 1-30 anchor points outperforms uniform sampling and other baselines at accurately ranking models.Moreover, just a dozen anchor points can be used to estimate model per-class predictions on all other points in a dataset with low error, sufficient for gauging where the model is likely to fail.Lastly, we present Anchor Point Maps for visualizing these insights and facilitating comparisons of the performance of different models on various regions within the dataset distribution. Rajan Vivek, Kawin Ethayarajh, Diyi Yang, Douwe Kiela |
EACL (1) | 4 |
| 2024 | Nearest Neighbor Normalization Improves Multimodal RetrievalabstractMultimodal models leverage large-scale pretraining to achieve strong but still imperfect performance on tasks such as image captioning, visual question answering, and cross-modal retrieval. In this paper, we present a simple and efficient method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization (NNN). We show an improvement on retrieval metrics in both text retrieval and image retrieval for all of the contrastive models that we tested (CLIP, BLIP, ALBEF, SigLIP, BEiT) and for both of the datasets that we used (MS-COCO and Flickr30k). NNN requires a reference database, but does not require any training on this database, and can even increase the retrieval accuracy of a model after finetuning. Neil Chowdhury, Franklin Wang, Sumedh Shenoy, Douwe Kiela, Sarah Schwettmann, Tristan Thrush |
EMNLP | 4 |
| 2024 | Model Alignment as Prospect Theoretic OptimizationabstractKahneman & Tversky’s $\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases—the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect theory literature. Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable. More broadly, our work suggests that there is no one HALO that is universally superior; the best loss depends on the inductive biases most appropriate for a given setting, an oft-overlooked consideration. Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Daniel Jurafsky, Douwe Kiela |
ICML | 5 |
| 2023 | Investigating Multi-source Active Learning for Natural Language InferenceabstractIn recent years, active learning has been successfully applied to an array of NLP tasks.However, prior work often assumes that training and test data are drawn from the same distribution.This is problematic, as in real-life settings data may stem from several sources of varying relevance and quality.We show that four popular active learning schemes fail to outperform random selection when applied to unlabelled pools comprised of multiple data sources on the task of natural language inference.We reveal that uncertainty-based strategies perform poorly due to the acquisition of collective outliers, i.e., hard-to-learn instances that hamper learning and generalization.When outliers are removed, strategies are found to recover and outperform random baselines.In further analysis, we find that collective outliers vary in form between sources, and show that hard-tolearn data is not always categorically harmful.Lastly, we leverage dataset cartography to introduce difficulty-stratified testing and find that different strategies are affected differently by example learnability and difficulty. Ard Snijders, Douwe Kiela, Aikaterini Margatina |
EACL | 2 |
| 2023 | OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text DocumentsabstractLarge multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELICS, we train on the dataset vision and language models of 9 and 80 billion parameters, IDEFICS-9B and IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code. Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh |
NeurIPS | 10 |
| 2023 | DataPerf: Benchmarks for Data-Centric AI DevelopmentabstractMachine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry. Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi |
NeurIPS | 13 |
| 2022 | FLAVA: A Foundational Language And Vision Alignment ModelabstractState-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal (contrastive) or multi-modal (with earlier fusion) but not both; and they often only target specific modalities or tasks. A promising direction would be to use a single holistic universal model, as a “foundation”, that targets all modalities at once-a true vision and language foundation model should be good at vision tasks, language tasks, and cross- and multi-modal vision and language tasks. We introduce FLAVA as such a model and demonstrate impressive performance on a wide range of 35 tasks spanning these target modalities. Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, Douwe Kiela |
CVPR | 7 |
| 2022 | Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityabstractWe present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly-but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set offine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground. Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, Candace Ross |
CVPR | 6 |
| 2022 | Perturbation Augmentation for Fairer NLPabstractUnwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets.In this work, we ask whether training on demographically perturbed data leads to fairer language models.We collect a large dataset of human annotated text perturbations and train a neural perturbation model, which we show outperforms heuristic alternatives.We find that (i) language models (LMs) pre-trained on demographically perturbed corpora are typically more fair, and (ii) LMs finetuned on perturbed GLUE datasets exhibit less demographic bias on downstream tasks, and (iii) fairness improvements do not come at the expense of performance on downstream tasks.Lastly, we discuss outstanding questions about how best to evaluate the (un)fairness of large language models.We hope that this exploration of neural demographic perturbation will help drive more improvement towards fairer NLP. Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, Adina Williams |
EMNLP | 5 |
| 2022 | Grounding, Meaning and Foundation Models: Adventures in Multimodal Machine LearningabstractIn this talk I will present a vision for acquiring perceptually grounded meaning in machines, as a key next challenge for natural language processing. I will cover some recent work that tries to improve how we do model evaluation in multimodal settings, focusing on the new Adversarial VQA and Winoground evaluation datasets. After that, I will talk about our latest large-scale vision and language "foundation model", called FLAVA: a single holistic universal transformer that targets all modalities at once and that shows impressive performance on a wide range of tasks. Douwe Kiela |
ACM Multimedia | 1 |
| 2022 | Models in the Loop: Aiding Crowdworkers with Generative Annotation AssistantsabstractMax Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, Douwe Kiela. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Max Bartolo, Tristan Thrush, Sebastian Riedel 0001, Pontus Stenetorp, Robin Jia, Douwe Kiela |
NAACL-HLT | 6 |
| 2021 | On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized StudyabstractDivyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Wen-tau Yih. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Divyansh Kaushik, Douwe Kiela, Zachary C. Lipton, Scott Yih |
ACL/IJCNLP (1) | 2 |
| 2021 | I like fish, especially dolphins: Addressing Contradictions in Dialogue ModelingabstractYixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, Jason Weston. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, Jason Weston |
ACL/IJCNLP (1) | 4 |
| 2021 | DynaSent: A Dynamic Benchmark for Sentiment AnalysisabstractChristopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Christopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela |
ACL/IJCNLP (1) | 4 |
| 2021 | Reservoir TransformersabstractSheng Shen, Alexei Baevski, Ari Morcos, Kurt Keutzer, Michael Auli, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sheng Shen 0001, Alexei Baevski, Ari S. Morcos, Kurt Keutzer, Michael Auli, Douwe Kiela |
ACL/IJCNLP (1) | 6 |
| 2021 | Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionabstractBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bertie Vidgen, Tristan Thrush, Zeerak Talat, Douwe Kiela |
ACL/IJCNLP (1) | 4 |
| 2021 | Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationabstractDespite recent progress, state-of-the-art question answering models remain vulnerable to a variety of adversarial attacks. While dynamic adversarial data collection, in which a human annotator tries to write examples that fool a model-in-the-loop, can improve model robustness, this process is expensive which limits the scale of the collected data. In this work, we are the first to use synthetic adversarial data generation to make question answering models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data. Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 0001, Pontus Stenetorp, Douwe Kiela |
EMNLP (1) | 6 |
| 2021 | Gradient-based Adversarial Attacks against Text TransformersabstractWe propose the first general-purpose gradientbased adversarial attack against transformer models.Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization.We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks, outperforming prior work in terms of adversarial success rate with matching imperceptibility as per automated and human evaluation.Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs. Chuan Guo 0001, Alexandre Sablayrolles, Hervé Jégou, Douwe Kiela |
EMNLP (1) | 4 |
| 2021 | What's Hidden in a One-layer Randomly Weighted Transformer?abstractWe demonstrate that, hidden within one-layer randomly weighted neural networks, there exist subnetworks that can achieve impressive performance, without ever modifying the weight initializations, on machine translation tasks.To find subnetworks for onelayer randomly weighted neural networks, we apply different binary masks to the same weight matrix to generate different layers.Hidden within a one-layer randomly weighted Transformer, we find that subnetworks that can achieve 29.45/17.29 BLEU on IWSLT14/WMT14.Using a fixed pretrained embedding layer, the previously found subnetworks are smaller than, but can match 98%/92% (34.14/25.24BLEU) of the performance of, a trained Transformer small/base on IWSLT14/WMT14.Furthermore, we demonstrate the effectiveness of larger and deeper transformers in this setting, as well as the impact of different initialization methods. 1 M = Top k (S), where Top k (S i,j ) = 1 S i,j in top k%, 0 else. Sheng Shen 0001, Zhewei Yao, Douwe Kiela, Kurt Keutzer, Michael W. Mahoney |
EMNLP (1) | 3 |
| 2021 | Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleabstractA possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines.In this paper, we propose a different explanation: MLMs succeed on downstream tasks mostly due to their ability to model higher-order word cooccurrence statistics.To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and we show that these models still achieve high accuracy after finetuning on many downstream tasks -including tasks specifically designed to be challenging for models that ignore word order.Our models also perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information.Overall, our results show that purely distributional information largely explains the success of pretraining, and they underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge. Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, Douwe Kiela |
EMNLP (1) | 6 |
| 2021 | Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz |
ICLR | 10 |
| 2021 | Rissanen Data Analysis: Examining Dataset Characteristics via Description LengthabstractWe introduce a method to determine if a certain capability helps to achieve an accurate model of given data. We view labels as being generated from the inputs by a program composed of subroutines with different capabilities, and we posit that a subroutine is useful if and only if the minimal program that invokes it is shorter than the one that does not. Since minimum program length is uncomputable, we instead estimate the labels’ minimum description length (MDL) as a proxy, giving us a theoretically-grounded method for analyzing dataset characteristics. We call the method Rissanen Data Analysis (RDA) after the father of MDL, and we showcase its applicability on a wide variety of settings in NLP, ranging from evaluating the utility of generating subquestions before answering a question, to analyzing the value of rationales and explanations, to investigating the importance of different parts of speech, and uncovering dataset gender bias. Ethan Perez, Douwe Kiela, Kyunghyun Cho |
ICML | 2 |
| 2021 | Dynabench: Rethinking Benchmarking in NLPabstractDouwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams |
NAACL-HLT | 1 |
| 2021 | Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingabstractWe introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality. Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela |
NeurIPS | 9 |
| 2021 | True Few-Shot Learning with Language ModelsabstractPretrained language models (LMs) perform well on many tasks even when learning from a few examples, but prior work uses many held-out examples to tune various aspects of learning, such as hyperparameters, training objectives, and natural language templates ("prompts"). Here, we evaluate the few-shot ability of LMs when such held-out examples are unavailable, a setting we call true few-shot learning. We test two model selection criteria, cross-validation and minimum description length, for choosing LM prompts and hyperparameters in the true few-shot setting. On average, both marginally outperform random selection and greatly underperform selection based on held-out examples. Moreover, selection criteria often prefer models that perform significantly worse than randomly-selected ones. We find similar results even when taking into account our uncertainty in a model's true performance during selection, as well as when varying the amount of computation and number of examples used for selection. Overall, our findings suggest that prior work significantly overestimated the true few-shot ability of LMs given the difficulty of few-shot model selection. Ethan Perez, Douwe Kiela, Kyunghyun Cho |
NeurIPS | 2 |
| 2021 | Human-Adversarial Visual Question AnsweringabstractPerformance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, it is clear that the problem is far from being solved. In order to stress test VQA models, we benchmark them against human-adversarial examples. Human subjects interact with a state-of-the-art VQA model, and for each image in the dataset, attempt to find a question where the model’s predicted answer is incorrect. We find that a wide range of state-of-the-art models perform poorly when evaluated on these examples. We conduct an extensive analysis of the collected adversarial examples and provide guidance on future research directions. We hope that this Adversarial VQA (AdVQA) benchmark can help drive progress in the field and advance the state of the art. Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, José Alberto López Magaña, Tristan Thrush, Wojciech Galuba, Devi Parikh, Douwe Kiela |
NeurIPS | 8 |
| 2020 | Generating Interactive Worlds with TextabstractProcedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangement of the elements. In this work, we investigate a machine learning approach for world creation using content from the multi-player text adventure game environment LIGHT (Urbanek et al. 2019). We introduce neural network based models to compositionally arrange locations, characters, and objects into a coherent whole. In addition to creating worlds based on existing elements, our models can generate new game content. Humans can also leverage our models to interactively aid in worldbuilding. We show that the game environments created with our approach are cohesive, diverse, and preferred by human evaluators compared to other machine learning based world construction algorithms. Angela Fan, Jack Urbanek, Pratik Ringshia, Emily Dinan, Emma Qian, Siddharth Karamcheti, Shrimai Prabhumoye, Douwe Kiela, Tim Rocktäschel, Arthur Szlam, Jason Weston |
AAAI | 8 |
| 2020 | Adversarial NLI: A New Benchmark for Natural Language UnderstandingabstractWe introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.We show that training models on this new dataset leads to state-of-the-art performance on a variety of popular NLI benchmarks, while posing a more difficult challenge with its new test set.Our analysis sheds light on the shortcomings of current state-of-theart models, and shows that non-expert annotators are successful at finding their weaknesses.The data collection method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate. Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, Douwe Kiela |
ACL | 6 |
| 2020 | Queens are Powerful too: Mitigating Gender Bias in Dialogue GenerationabstractModels often easily learn biases present in the training data, and their predictions directly reflect this bias.We analyze gender bias in dialogue data, and examine how this bias is actually amplified in subsequent generative chit-chat dialogue models.We measure gender bias in six existing dialogue datasets, and focus on the most biased one, the multiplayer text-based fantasy adventure dataset LIGHT (Urbanek et al., 2019), as a testbed for our bias mitigation techniques.The LIGHT dataset is highly imbalanced with respect to gender, containing predominantly male characters, likely because it is entirely collected by crowdworkers and reflects common biases that exist in fantasy or medieval settings.We consider three techniques to mitigate gender bias: counterfactual data augmentation, targeted data collection, and bias controlled training.We show that our proposed techniques mitigate gender bias in LIGHT by balancing the genderedness of generated dialogue utterances and are particularly effective in combination.We quantify performance using various evaluation methods-such as quantity of gendered words, a dialogue safety classifier, and human studies-all of which show that our models generate less gendered, but equally engaging chit-chat responses. Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, Jason Weston |
EMNLP (1) | 5 |
| 2020 | Multi-Dimensional Gender Bias ClassificationabstractMachine learning models are trained to find patterns in data.NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.In this work, we propose a novel, general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.In addition, we collect a new, crowdsourced evaluation benchmark.Distinguishing between gender bias along multiple dimensions enables us to train better and more fine-grained gender bias classifiers.We show our classifiers are valuable for a variety of applications, like controlling for gender bias in generative models, detecting gender bias in arbitrary text, and classifying text as offensive based on its genderedness. Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, Adina Williams |
EMNLP (1) | 5 |
| 2020 | Unsupervised Question Decomposition for Question AnsweringabstractWe aim to improve question answering (QA) by decomposing hard questions into simpler sub-questions that existing QA systems are capable of answering.Since labeling questions with decompositions is cumbersome, we take an unsupervised approach to produce sub-questions, also enabling us to leverage millions of questions from the internet.Specifically, we propose an algorithm for One-to-N Unsupervised Sequence transduction (ONUS) that learns to map one hard, multi-hop question to many simpler, singlehop sub-questions.We answer sub-questions with an off-the-shelf QA model and give the resulting answers to a recomposition model that combines them into a final answer.We show large QA improvements on HOTPOTQA over a strong baseline on the original, out-ofdomain, and multi-hop dev sets.ONUS automatically learns to decompose different kinds of questions, while matching the utility of supervised and heuristic decomposition methods for QA and exceeding those methods in fluency.Qualitatively, we find that using subquestions is promising for shedding light on why a QA system makes a prediction. 1 * KC was a part-time research scientist at Facebook AI Research while working on this paper.1 Our code, data, and pretrained models are available at https://github.com/facebookresearch/ UnsupervisedDecomposition. What profession do H. L. Mencken and Albert Camus have in common? Ethan Perez, Patrick S. H. Lewis, Scott Yih, Kyunghyun Cho, Douwe Kiela |
EMNLP (1) | 5 |
| 2020 | On the interaction between supervision and self-play in emergent communication
Ryan Lowe, Abhinav Gupta 0002, Jakob N. Foerster, Douwe Kiela, Joelle Pineau |
ICLR | 4 |
| 2020 | Learning Optimal Representations with the Decodable Information BottleneckabstractWe address the question of characterizing and finding optimal representations for supervised learning. Traditionally, this question has been tackled using the Information Bottleneck, which compresses the inputs while retaining information about the targets, in a decoder-agnostic fashion. In machine learning, however, our goal is not compression but rather generalization, which is intimately linked to the predictive family or decoder of interest (e.g. linear classifier). We propose the Decodable Information Bottleneck (DIB) that considers information retention and compression from the perspective of the desired predictive family. As a result, DIB gives rise to representations that are optimal in terms of expected test performance and can be estimated with guarantees. Empirically, we show that the framework can be used to enforce a small generalization gap on downstream classifiers and to predict the generalization ability of neural networks. Yann Dubois, Douwe Kiela, David J. Schwab, Ramakrishna Vedantam |
NeurIPS | 2 |
| 2020 | The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesabstractThis work proposes a new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes. It is constructed such that unimodal models struggle and only multimodal models can succeed: difficult examples (“benign confounders”) are added to the dataset to make it hard to rely on unimodal signals. The task requires subtle reasoning, yet is straightforward to evaluate as a binary classification problem. We provide baseline performance numbers for unimodal models, as well as for multimodal models with various degrees of sophistication. We find that state-of-the-art methods perform poorly compared to humans, illustrating the difficulty of the task and highlighting the challenge that this important problem poses to the community. Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, Davide Testuggine |
NeurIPS | 1 |
| 2020 | Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksabstractLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline. Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal 0001, Heinrich Küttler, Mike Lewis, Scott Yih, Tim Rocktäschel, Sebastian Riedel 0001, Douwe Kiela |
NeurIPS | 12 |
| 2019 | Analysis of Joint Multilingual Sentence Representations and Semantic K-Nearest Neighbor Graphs
Holger Schwenk, Douwe Kiela, Matthijs Douze |
AAAI | 2 |
| 2019 | Inferring Concept Hierarchies from Text Corpora via Hyperbolic EmbeddingsabstractWe consider the task of inferring “is-a” relationships from large text corpora. For this purpose, we propose a new method combining hyperbolic embeddings and Hearst patterns. This approach allows us to set appropriate constraints for inferring concept hierarchies from distributional contexts while also being able to predict missing “is-a”-relationships and to correct wrong extractions. Moreover – and in contrast with other methods – the hierarchical nature of hyperbolic space allows us to learn highly efficient representations and to improve the taxonomic consistency of the inferred hierarchies. Experimentally, we show that our approach achieves state-of-the-art performance on several commonly-used benchmarks. Matt Le 0001, Stephen Roller, Laetitia Meng-Papaxanthos, Douwe Kiela, Maximilian Nickel |
ACL (1) | 4 |
| 2019 | Emergent Linguistic Phenomena in Multi-Agent Communication GamesabstractLaura Harding Graesser, Kyunghyun Cho, Douwe Kiela. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Laura Graesser, Kyunghyun Cho, Douwe Kiela |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Countering Language Drift via Visual GroundingabstractJason Lee, Kyunghyun Cho, Douwe Kiela. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jason Lee 0002, Kyunghyun Cho, Douwe Kiela |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Finding Generalizable Evidence by Learning to Convince Q&A ModelsabstractEthan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ethan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Learning to Speak and Act in a Fantasy Text Adventure GameabstractJack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau 0001, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston |
EMNLP/IJCNLP (1) | 8 |
| 2019 | No Training Required: Exploring Random Encoders for Sentence Classification
John Wieting, Douwe Kiela |
ICLR (Poster) | 2 |
| 2019 | Hyperbolic Graph Neural NetworksabstractLearning from graph-structured data is an important task in machine learning and artificial intelligence, for which Graph Neural Networks (GNNs) have shown great promise. Motivated by recent advances in geometric representation learning, we propose a novel GNN architecture for learning representations on Riemannian manifolds with differentiable exponential and logarithmic maps. We develop a scalable algorithm for modeling the structural properties of graphs, comparing Euclidean and hyperbolic geometry. In our experiments, we show that hyperbolic GNNs can lead to substantial improvements on various benchmark datasets. Qi Liu 0049, Maximilian Nickel, Douwe Kiela |
NeurIPS | 3 |
| 2018 | Efficient Large-Scale Multi-Modal ClassificationabstractWhile the incipient internet was largely text-based, the modern digital world is becoming increasingly multi-modal. Here, we examine multi-modal classification where one modality is discrete, e.g. text, and the other is continuous, e.g. visual representations transferred from a convolutional neural network. In particular, we focus on scenarios where we have to be able to classify large quantities of data quickly. We investigate various methods for performing multi-modal fusion and analyze their trade-offs in terms of classification accuracy and computational efficiency. Our findings indicate that the inclusion of continuous information improves performance over text-only on a range of multi-modal classification tasks, even with simple fusion methods. In addition, we experiment with discretizing the continuous features in order to speed up and simplify the fusion process even further. Our results show that fusion with discretized features outperforms text-only classification, at a fraction of the computational cost of full multi-modal fusion, with the additional benefit of improved interpretability. Douwe Kiela, Edouard Grave, Armand Joulin, Tomás Mikolov |
AAAI | 1 |
| 2018 | Personalizing Dialogue Agents: I have a dog, do you have pets too?abstractSaizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston |
ACL (1) | 5 |
| 2018 | Dynamic Meta-Embeddings for Improved Sentence RepresentationsabstractWhile one of the first steps in many NLP systems is selecting what pre-trained word embeddings to use, we argue that such a step is better left for neural networks to figure out by themselves.To that end, we introduce dynamic meta-embeddings, a simple yet effective method for the supervised learning of embedding ensembles, which leads to stateof-the-art performance within the same model class on a variety of tasks.We subsequently show how the technique can be used to shed new light on the usage of word embeddings in NLP systems. Douwe Kiela, Changhan Wang, Kyunghyun Cho |
EMNLP | 1 |
| 2018 | Emergent Communication in a Multi-Modal, Multi-Step Referential Game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, Kyunghyun Cho |
ICLR (Poster) | 3 |
| 2018 | Emergent Translation in Multi-Agent Communication
Jason Lee 0002, Kyunghyun Cho, Jason Weston, Douwe Kiela |
ICLR (Poster) | 4 |
| 2018 | Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang 0001, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston |
ICLR (Poster) | 7 |
| 2018 | Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic GeometryabstractWe are concerned with the discovery of hierarchical relationships from large-scale unstructured similarity scores. For this purpose, we study different models of hyperbolic space and find that learning embeddings in the Lorentz model is substantially more efficient than in the Poincar{é}-ball model. We show that the proposed approach allows us to learn high-quality embeddings of large taxonomies which yield improvements over Poincar{é} embeddings, especially in low dimensions. Lastly, we apply our model to discover hierarchies in two real-world datasets: we show that an embedding in hyperbolic space can reveal important aspects of a company’s organizational structure as well as reveal historical relationships between language families. Maximilian Nickel, Douwe Kiela |
ICML | 2 |
| 2018 | SentEval: An Evaluation Toolkit for Universal Sentence Representations
Alexis Conneau, Douwe Kiela |
LREC | 2 |
| 2018 | Learning Visually Grounded Sentence RepresentationsabstractDouwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Douwe Kiela, Alexis Conneau, Allan Jabri, Maximilian Nickel |
NAACL-HLT | 1 |
| 2017 | Automatically Generating Rhythmic Verse with Neural NetworksabstractWe propose two novel methodologies for the automatic generation of rhythmic poetry in a variety of forms.The first approach uses a neural language model trained on a phonetic encoding to learn an implicit representation of both the form and content of English poetry.This model can effectively learn common poetic devices such as rhyme, rhythm and alliteration.The second approach considers poetry generation as a constraint satisfaction problem where a generative neural language model is tasked with learning a representation of content, and a discriminative weighted finite state machine constrains it on the basis of form.By manipulating the constraints of the latter model, we can generate coherent poetry with arbitrary forms and themes.A large-scale extrinsic evaluation demonstrated that participants consider machine-generated poems to be written by humans 54% of the time.In addition, participants rated a machinegenerated poem to be the most human-like amongst all evaluated. Jack Hopkins, Douwe Kiela |
ACL (1) | 2 |
| 2017 | Evaluation by Association: A Systematic Study of Quantitative Word Association EvaluationabstractRecent work on evaluating representation learning architectures in NLP has established a need for evaluation protocols based on subconscious cognitive measures rather than manually tailored intrinsic similarity and relatedness tasks.In this work, we propose a novel evaluation framework that enables large-scale evaluation of such architectures in the free word association (WA) task, which is firmly grounded in cognitive theories of human semantic representation.This evaluation is facilitated by the existence of large manually constructed repositories of word association data.In this paper, we (1) present a detailed analysis of the new quantitative WA evaluation protocol, (2) suggest new evaluation metrics for the WA task inspired by its direct analogy with information retrieval problems, (3) evaluate various state-of-the-art representation models on this task, and (4) discuss the relationship between WA and prior evaluations of semantic representation with well-known similarity and relatedness evaluation sets.We have made the WA evaluation toolkit publicly available. Ivan Vulic, Douwe Kiela, Anna Korhonen |
EACL (1) | 2 |
| 2017 | Supervised Learning of Universal Sentence Representations from Natural Language Inference DataabstractMany modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features.Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful.Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted.In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors (Kiros et al., 2015) on a wide range of transfer tasks.Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks.Our encoder is publicly available 1 . Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes |
EMNLP | 2 |
| 2017 | Grasping the Finer Point: A Supervised Similarity Network for Metaphor DetectionabstractThe ubiquity of metaphor in our everyday communication makes it an important problem for natural language understanding.Yet, the majority of metaphor processing systems to date rely on handengineered features and there is still no consensus in the field as to which features are optimal for this task.In this paper, we present the first deep learning architecture designed to capture metaphorical composition.Our results demonstrate that it outperforms the existing approaches in the metaphor identification task. Marek Rei, Luana Bulat, Douwe Kiela, Ekaterina Shutova |
EMNLP | 3 |
| 2017 | Poincaré Embeddings for Learning Hierarchical RepresentationsabstractRepresentation learning has become an invaluable approach for learning from symbolic data such as text and graphs. However, state-of-the-art embedding methods typically do not account for latent hierarchical structures which are characteristic for many complex symbolic datasets. In this work, we introduce a new approach for learning hierarchical representations of symbolic data by embedding them into hyperbolic space -- or more precisely into an n-dimensional Poincaré ball. Due to the underlying hyperbolic geometry, this allows us to learn parsimonious representations of symbolic data by simultaneously capturing hierarchy and similarity. We present an efficient algorithm to learn the embeddings based on Riemannian optimization and show experimentally that Poincaré embeddings can outperform Euclidean embeddings significantly on data with latent hierarchies, both in terms of representation capacity and in terms of generalization ability. Maximilian Nickel, Douwe Kiela |
NIPS | 2 |
| 2017 | HyperLex: A Large-Scale Evaluation of Graded Lexical EntailmentabstractWe introduce HyperLex—a data set and evaluation resource that quantifies the extent of the semantic category membership, that is, type-of relation, also known as hyponymy–hypernymy or lexical entailment (LE) relation between 2,616 concept pairs. Cognitive psychology research has established that typicality and category/class membership are computed in human semantic memory as a gradual rather than binary relation. Nevertheless, most NLP research and existing large-scale inventories of concept category membership (WordNet, DBPedia, etc.) treat category membership and LE as binary. To address this, we asked hundreds of native English speakers to indicate typicality and strength of category membership between a diverse range of concept pairs on a crowdsourcing platform. Our results confirm that category membership and LE are indeed more gradual than binary. We then compare these human judgments with the predictions of automatic systems, which reveals a huge gap between human performance and state-of-the-art LE, distributional and representation learning models, and substantial differences between the models themselves. We discuss a pathway for improving semantic models to overcome this discrepancy, and indicate future application areas for improved graded LE systems. Ivan Vulic, Daniela Gerz, Douwe Kiela, Felix Hill, Anna Korhonen |
Comput. Linguistics | 3 |
| 2017 | Learning Neural Audio Embeddings for Grounding Semantics in Auditory PerceptionabstractMulti-modal semantics, which aims to ground semantic representations in perception, has relied on feature norms or raw image data for perceptual input. In this paper we examine grounding semantic representations in raw auditory data, using standard evaluations for multi-modal semantics. After having shown the quality of such auditorily grounded representations, we show how they can be applied to tasks where auditory perception is relevant, including two unsupervised categorization experiments, and provide further analysis. We find that features transfered from deep neural networks outperform bag of audio words approaches. To our knowledge, this is the first work to construct multi-modal models from a combination of textual information and auditory information extracted from deep neural networks, and the first work to evaluate the performance of tri-modal (textual, visual and auditory) semantic models. Douwe Kiela, Stephen Clark |
J. Artif. Intell. Res. | 1 |
| 2017 | Visually Grounded and Textual Semantic Models Differentially Decode Brain Activity Associated with Concrete and Abstract NounsabstractImportant advances have recently been made using computational semantic models to decode brain activity patterns associated with concepts; however, this work has almost exclusively focused on concrete nouns. How well these models extend to decoding abstract nouns is largely unknown. We address this question by applying state-of-the-art computational models to decode functional Magnetic Resonance Imaging (fMRI) activity patterns, elicited by participants reading and imagining a diverse set of both concrete and abstract nouns. One of the models we use is linguistic, exploiting the recent word2vec skipgram approach trained on Wikipedia. The second is visually grounded, using deep convolutional neural networks trained on Google Images. Dual coding theory considers concrete concepts to be encoded in the brain both linguistically and visually, and abstract concepts only linguistically. Splitting the fMRI data according to human concreteness ratings, we indeed observe that both models significantly decode the most concrete nouns; however, accuracy is significantly greater using the text-based models for the most abstract nouns. More generally this confirms that current computational models are sufficiently advanced to assist in investigating the representational structure of abstract concepts in the brain. Andrew J. Anderson, Douwe Kiela, Stephen Clark, Massimo Poesio |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Robust Text Classification for Sparsely Labelled Data Using Multi-level EmbeddingsabstractThe conventional solution for handling sparsely labelled data is extensive feature engineering. This is time consuming and task and domain specific. We present a novel approach for learning embedded features that aims to alleviate this problem. Our approach jointly learns embeddings at different levels of granularity (word, sentence and document) along with the class labels. The intuition is that topic semantics represented by embeddings at multiple levels results in better classification. We evaluate this approach in unsupervised and semi-supervised settings on two sparsely labelled classification tasks, outperforming the handcrafted models and several embedding baselines. Simon Baker, Douwe Kiela, Anna Korhonen |
COLING | 2 |
| 2016 | Comparing Data Sources and Architectures for Deep Visual Representation Learning in SemanticsabstractMulti-modal distributional models learn grounded representations for improved performance in semantics.Deep visual representations, learned using convolutional neural networks, have been shown to achieve particularly high performance.In this study, we systematically compare deep visual representation learning techniques, experimenting with three well-known network architectures.In addition, we explore the various data sources that can be used for retrieving relevant images, showing that images from search engines perform as well as, or better than, those from manually crafted resources such as ImageNet.Furthermore, we explore the optimal number of images and the multi-lingual applicability of multi-modal semantics.We hope that these findings can serve as a guide for future research in the field. Douwe Kiela, Anita L. Vero, Stephen Clark |
EMNLP | 1 |
| 2016 | Vision and Feature Norms: Improving automatic feature norm learning through cross-modal mapsabstractProperty norms have the potential to aid a wide range of semantic tasks, provided that they can be obtained for large numbers of concepts. Recent work has focused on text as the main source of information for automatic property extraction. In this paper we examine property norm prediction from visual, rather than textual, data, using cross-modal maps learnt between property norm and visual spaces. We also investigate the importance of having a complete feature norm dataset, for both training and testing. Finally, we evaluate how these datasets and cross-modal maps can be used in an image retrieval task. Luana Bulat, Douwe Kiela, Stephen Clark |
HLT-NAACL | 2 |
| 2016 | Black Holes and White Rabbits: Metaphor Identification with Visual FeaturesabstractMetaphor is pervasive in our communication, which makes it an important problem for natural language processing (NLP).Numerous approaches to metaphor processing have thus been proposed, all of which relied on linguistic features and textual data to construct their models.Human metaphor comprehension is, however, known to rely on both our linguistic and perceptual experience, and vision can play a particularly important role when metaphorically projecting imagery across domains.In this paper, we present the first metaphor identification method that simultaneously draws knowledge from linguistic and visual data.Our results demonstrate that it outperforms linguistic and visual models in isolation, as well as being competitive with the best-performing metaphor identification methods, that rely on hand-crafted knowledge about domains and perception. Ekaterina Shutova, Douwe Kiela, Jean Maillard |
HLT-NAACL | 2 |
| 2015 | Multi- and Cross-Modal Semantics Beyond Vision: Grounding in Auditory PerceptionabstractMulti-modal semantics has relied on feature norms or raw image data for perceptual input.In this paper we examine grounding semantic representations in raw auditory data, using standard evaluations for multi-modal semantics, including measuring conceptual similarity and relatedness.We also evaluate cross-modal mappings, through a zero-shot learning task mapping between linguistic and auditory modalities.In addition, we evaluate multimodal representations on an unsupervised musical instrument clustering task.To our knowledge, this is the first work to combine linguistic and auditory information into multi-modal representations. Douwe Kiela, Stephen Clark |
EMNLP | 1 |
| 2015 | Specializing Word Embeddings for Similarity or RelatednessabstractWe demonstrate the advantage of specializing semantic word embeddings for either similarity or relatedness.We compare two variants of retrofitting and a joint-learning approach, and find that all three yield specialized semantic spaces that capture human intuitions regarding similarity and relatedness better than unspecialized spaces.We also show that using specialized spaces in NLP tasks and applications leads to clear improvements, for document classification and synonym selection, which rely on either similarity or relatedness but not both. Douwe Kiela, Felix Hill, Stephen Clark |
EMNLP | 1 |
| 2015 | Visual Bilingual Lexicon Induction with Transferred ConvNet FeaturesabstractThis paper is concerned with the task of bilingual lexicon induction using imagebased features.By applying features from a convolutional neural network (CNN), we obtain state-of-the-art performance on a standard dataset, obtaining a 79% relative improvement over previous work which uses bags of visual words based on SIFT features.The CNN image-based approach is also compared with state-of-the-art linguistic approaches to bilingual lexicon induction, even outperforming these for one of three language pairs on another standard dataset.Furthermore, we shed new light on the type of visual similarity metric to use for genuine similarity versus relatedness tasks, and experiment with using multiple layers from the same network in an attempt to improve performance. Douwe Kiela, Ivan Vulic, Stephen Clark |
EMNLP | 1 |
| 2015 | Unsupervised discovery of information structure in biomedical documentsabstractMOTIVATION: Information structure (IS) analysis is a text mining technique, which classifies text in biomedical articles into categories that capture different types of information, such as objectives, methods, results and conclusions of research. It is a highly useful technique that can support a range of Biomedical Text Mining tasks and can help readers of biomedical literature find information of interest faster, accelerating the highly time-consuming process of literature review. Several approaches to IS analysis have been presented in the past, with promising results in real-world biomedical tasks. However, all existing approaches, even weakly supervised ones, require several hundreds of hand-annotated training sentences specific to the domain in question. Because biomedicine is subject to considerable domain variation, such annotations are expensive to obtain. This makes the application of IS analysis across biomedical domains difficult. In this article, we investigate an unsupervised approach to IS analysis and evaluate the performance of several unsupervised methods on a large corpus of biomedical abstracts collected from PubMed. RESULTS: Our best unsupervised algorithm (multilevel-weighted graph clustering algorithm) performs very well on the task, obtaining over 0.70 F scores for most IS categories when applied to well-known IS schemes. This level of performance is close to that of lightly supervised IS methods and has proven sufficient to aid a range of practical tasks. Thus, using an unsupervised approach, IS could be applied to support a wide range of tasks across sub-domains of biomedicine. We also demonstrate that unsupervised learning brings novel insights into IS of biomedical literature and discovers information categories that are not present in any of the existing IS schemes. AVAILABILITY AND IMPLEMENTATION: The annotated corpus and software are available at http://www.cl.cam.ac.uk/∼dk427/bio14info.html. Douwe Kiela, Ulla Stenius, Anna Korhonen |
Bioinform. | 1 |
| 2014 | Learning Image Embeddings using Convolutional Neural Networks for Improved Multi-Modal SemanticsabstractWe construct multi-modal concept representations by concatenating a skip-gram linguistic representation vector with a visual concept representation vector computed using the feature extraction layers of a deep convolutional neural network (CNN) trained on a large labeled object recognition dataset.This transfer learning approach brings a clear performance gain over features based on the traditional bag-of-visual-word approach.Experimental results are reported on the WordSim353 and MEN semantic relatedness evaluation tasks.We use visual features computed using either ImageNet or ESP Game images. Douwe Kiela, Léon Bottou |
EMNLP | 1 |
| 2013 | Detecting Compositionality of Multi-Word Expressions using Nearest Neighbours in Vector Space ModelsabstractWe present a novel unsupervised approach to detecting the compositionality of multi-word expressions.We compute the compositionality of a phrase through substituting the constituent words with their "neighbours" in a semantic vector space and averaging over the distance between the original phrase and the substituted neighbour phrases.Several methods of obtaining neighbours are presented.The results are compared to existing supervised results and achieve state-of-the-art performance on a verb-object dataset of human compositionality ratings. Douwe Kiela, Stephen Clark |
EMNLP | 1 |