Desmond Elliott

dblp:46/7536 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-3112-7904ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Dynaword: From One-shot to Continuously Developed Datasets
abstract
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguously licensed sources restricting use, sharing, and derivative works; (2) static dataset releases that prevent community contributions and diminish longevity; and (3) quality assurance processes restricted to publishing teams rather than leveraging community expertise. To address these limitations, we introduce two contributions: the Dynaword approach and Danish Dynaword. The Dynaword approach is a framework for creating large-scale, open datasets that can be continuously updated through community collaboration. Danish Dynaword is a concrete implementation that validates this approach and demonstrates its potential. Danish Dynaword contains over four times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry and research. The repository includes light-weight tests to ensure data formatting, quality, and documentation, establishing a sustainable framework for ongoing community contributions and dataset evolution.
Kenneth C. Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan, Balázs Szabó, Márton Kardos, Kirsten Vad, Johan Heinsen, Andrea Blasi Núñez, Gianluca Barmina, Jacob Nielsen, Rasmus Larsen 0001, Rob van der Goot, Peter Bjerregaard Vahlstrup, Per Møldrup-Dalum, Desmond Elliott, Lukas Galke Poech, Peter Schneider-Kamp, Kristoffer L. Nielbo
LREC15
2026 The Second Workshop on Evaluation of Multimodal Generation
abstract
Multimodal generation and retrieval systems are increasingly central to modern information retrieval, powering retrieval-augmented generation (RAG), multimodal search, recommendation, and knowledge intensive applications. Despite rapid progress in multimodal large language models (MLLMs), robust and principled evaluation of multimodal generation and retrieval remains a major open challenge for the IR community. This workshop aims to foster discussions and research efforts by bringing together researchers and practitioners in information retrieval, natural language processing, computer vision, and multimodal AI. Our goal is to establish evaluation methods for multimodal research and advance research efforts in this direction.
Wei Zhang 0098, Xiang Dai 0001, Sarvnaz Karimi, Desmond Elliott, Biaoyan Fang, Mong Yuan Sim
SIGIR4
2026 ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
abstract
Reasoning over sequences of images remains a challenge for multimodal large language models (MLLMs). While recent models incorporate multi-image data during pretraining, they still struggle to recognize sequential structures, often treating images independently. This work introduces ImageChain, a framework that enhances MLLMs with sequential reasoning capabilities over image data by modeling visual sequences as a multi-turn conversation. In ImageChain, images are interleaved with corresponding textual descriptions to form a controlled dialogue that explicitly captures temporal dependencies and narrative progression. Our method optimizes for the task of next-scene description, where the model generates a context-aware description of an upcoming scene based on preceding visual and textual cues. We demonstrate that our approach improves performance on the next-scene description task – achieving an average improvement from 3.7% to 19% in SimRate, a metric that quantifies semantic similarity to human-annotated ground truths. Moreover, ImageChain achieves robust zero-shot out-of-domain performance in applications ranging from comics to robotics. Extensive experiments validate that instruction-tuning in a multimodal, multi-turn conversation design is key to bridging the gap between static image understanding and temporally-aware reasoning.1
Danae Sanchez Villegas, Ingo Ziegler, Desmond Elliott
WACV3
2025 Multilingual Pretraining for Pixel Language Models
abstract
Pixel language models operate directly on images of rendered text, eliminating the need for a fixed vocabulary.While these models have demonstrated strong capabilities for downstream cross-lingual transfer, multilingual pretraining remains underexplored.We introduce PIXEL-M4, a model pretrained on four visually and linguistically diverse languages: English, Hindi, Ukrainian, and Simplified Chinese.Multilingual evaluations on semantic and syntactic tasks show that PIXEL-M4 outperforms an English-only counterpart on non-Latin scripts.Word-level probing analyses confirm that PIXEL-M4 captures rich linguistic features, even in languages not seen during pretraining.Furthermore, an analysis of its hidden representations shows that multilingual pretraining yields a semantic embedding space closely aligned across the languages used for pretraining.This work demonstrates that multilingual pretraining substantially enhances the capability of pixel language models to effectively support a diverse set of languages.
Ilker Kesen, Jonas F. Lotz, Ingo Ziegler, Phillip Rust, Desmond Elliott
EMNLP5
2025 CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
abstract
Abstract Building high-quality datasets for specialized tasks is a time-consuming and resource-intensive process that often requires specialized domain knowledge. We propose Corpus Retrieval and Augmentation for Fine-Tuning (CRAFT), a method for generating synthetic datasets, given a small number of user-written few-shots that demonstrate the task to be performed. Given these examples, CRAFT uses large-scale public web-crawled corpora and similarity-based document retrieval to find other relevant human-written documents. Lastly, instruction-tuned large language models (LLMs) augment the retrieved documents into custom-formatted task samples, which then can be used for fine- tuning. We demonstrate that CRAFT can efficiently generate large-scale task-specific training datasets for four diverse tasks: biology, medicine, and commonsense question-answering (QA), as well as summarization. Our experiments show that CRAFT-based models outperform or match general LLMs on QA tasks, while exceeding models trained on human-curated summarization data by 46 preference points. CRAFT outperforms other synthetic dataset generation methods such as Self- and Evol-Instruct, and remains robust even when the quality of the initial few-shots varies.
Ingo Ziegler, Abdullatif Köksal, Desmond Elliott, Hinrich Schütze
Trans. Assoc. Comput. Linguistics3
2024 Understanding Retrieval Robustness for Retrieval-augmented Image Captioning
abstract
Recent advances in retrieval-augmented models for image captioning highlight the benefit of retrieving related captions for efficient, lightweight models with strong domain-transfer capabilities.While these models demonstrate the success of retrieval augmentation, retrieval models are still far from perfect in practice: the retrieved information can sometimes mislead the model, resulting in incorrect generation and worse performance.In this paper, we analyze the robustness of a retrieval-augmented captioning model SMALLCAP.Our analysis shows that the model is sensitive to tokens that appear in the majority of the retrieved captions, and the input attribution shows that those tokens are likely copied into the generated output.Given these findings, we propose to train the model by sampling retrieved captions from more diverse sets.This decreases the chance that the model learns to copy majority tokens, and improves both in-domain and cross-domain performance.
Wenyan Li 0001, Jiaang Li 0002, Rita Ramos, Raphael Tang, Desmond Elliott
ACL (1)5
2024 The Role of Data Curation in Image Captioning
abstract
Image captioning models are typically trained by treating all samples equally, neglecting to account for mismatched or otherwise difficult data points.In contrast, recent work has shown the effectiveness of training models by scheduling the data using curriculum learning strategies.This paper contributes to this direction by actively curating difficult samples in datasets without increasing the total number of samples.We explore the effect of using three data curation methods within the training process: complete removal of a sample, caption replacement, or image replacement via a text-to-image generation model.Experiments on the Flickr30K and COCO datasets with the BLIP and BEiT-3 models demonstrate that these curation methods do indeed yield improved image captioning models, underscoring their efficacy.
Wenyan Li 0001, Jonas F. Lotz, Chen Qiu 0005, Desmond Elliott
EACL (1)4
2024 FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
abstract
Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott
EMNLP12
2024 Sequential Compositional Generalization in Multimodal Models
abstract
Semih Yagcioglu, Osman Batur İnce, Aykut Erdem, Erkut Erdem, Desmond Elliott, Deniz Yuret. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Semih Yagcioglu, Osman Batur Ince, Aykut Erdem, Erkut Erdem, Desmond Elliott, Deniz Yuret
NAACL-HLT5
2023 Smallcap: Lightweight Image Captioning Prompted with Retrieval Augmentation
abstract
Recent advances in image captioning have focused on scaling the data and model size, substantially increasing the cost of pretraining and finetuning. As an alternative to large models, we present Smallcap, which generates a caption conditioned on an input image and related captions retrieved from a datastore. Our model is lightweight and fast to train, as the only learned parameters are in newly introduced cross-attention layers between a pre-trained CLIP encoder and GPT-2 decoder. Smallcap can transfer to new domains without additional finetuning and can exploit large-scale data in a training-free fashion since the contents of the datastore can be readily replaced. Our experiments show that Smallcap, trained only on COCO, has competitive performance on this benchmark, and also transfers to other domains without retraining, solely through retrieval from target-domain data. Further improvement is achieved through the training-free exploitation of diverse human-labeled and web data, which proves to be effective for a range of domains, including the nocaps benchmark, designed to test generalization to unseen visual concepts.11Code: https://github.com/RitaRamo/smallcap.
Rita Ramos, Bruno Martins 0001, Desmond Elliott, Yova Kementchedjhieva
CVPR3
2023 Retrieval-augmented Image Captioning
abstract
Inspired by retrieval-augmented language generation and pretrained Vision and Language (V&L) encoders, we present a new approach to image captioning that generates sentences given the input image and a set of captions retrieved from a datastore, as opposed to the image alone.The encoder in our model jointly processes the image and retrieved captions using a pretrained V&L BERT, while the decoder attends to the multimodal encoder representations, benefiting from the extra textual evidence from the retrieved captions.Experimental results on the COCO dataset show that image captioning can be effectively formulated from this new perspective.Our model, named EXTRA, benefits from using captions retrieved from the training dataset, and it can also benefit from using an external dataset without the need for retraining.Ablation studies show that retrieving a sufficient number of captions (e.g., k=5) can improve captioning quality.Our work contributes towards using pretrained V&L encoders for generative tasks, instead of standard classification tasks.
Rita Ramos, Desmond Elliott, Bruno Martins 0001
EACL2
2023 PHD: Pixel-Based Language Modeling of Historical Documents
abstract
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein
EMNLP3
2023 Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models
abstract
Pretrained machine learning models are known to perpetuate and even amplify existing biases in data, which can result in unfair outcomes that ultimately impact user experience.Therefore, it is crucial to understand the mechanisms behind those prejudicial biases to ensure that model performance does not result in discriminatory behaviour toward certain groups or populations.In this work, we define gender bias as our case study.We quantify bias amplification in pretraining and after fine-tuning on three families of vision-and-language models.We investigate the connection, if any, between the two learning stages, and evaluate how bias amplification reflects on model performance.Overall, we find that bias amplification in pretraining and after fine-tuning are independent.We then examine the effect of continued pretraining on gender-neutral data, finding that this reduces group disparities, i.e., promotes fairness, on VQAv2 and retrieval tasks without significantly compromising task performance.
Laura Cabello Piqueras, Emanuele Bugliarello, Stephanie Brandl, Desmond Elliott
EMNLP4
2023 Text Rendering Strategies for Pixel Language Models
abstract
Pixel-based language models process text rendered as images, which allows them to handle any script, making them a promising approach to open vocabulary language modelling.However, recent approaches use text renderers that produce a large set of almost-equivalent input patches, which may prove sub-optimal for downstream tasks, due to redundancy in the input representations.In this paper, we investigate four approaches to rendering text in the PIXEL model (Rust et al., 2023), and find that simple character bigram rendering brings improved performance on sentence-level tasks without compromising performance on tokenlevel or multilingual tasks.This new rendering strategy also makes it possible to train a more compact model with only 22M parameters that performs on par with the original 86M parameter model.Our analyses show that character bigram rendering leads to a consistently better model but with an anisotropic patch embedding space, driven by a patch frequency bias, highlighting the connections between image patchand tokenization-based language models.
Jonas F. Lotz, Elizabeth Salesky, Phillip Rust, Desmond Elliott
EMNLP4
2023 Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, Desmond Elliott
ICLR6
2022 Date Recognition in Historical Parish Records
Laura Cabello Piqueras, Constanza Fierro, Jonas F. Lotz, Phillip Rust, Joen Rommedahl, Jeppe Klok Due, Christian Igel, Desmond Elliott, Carsten B. Pedersen, Israfel Salazar, Anders Søgaard
ICFHR8
2022 IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages
abstract
Reliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community.
Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic
ICML5
2021 The Role of Syntactic Planning in Compositional Image Captioning
abstract
Image captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different distributions of images.Recently, Nikolaus et al. (2019) introduced a dataset to assess compositional generalization in image captioning, where models are evaluated on their ability to describe images with unseen adjective-noun and noun-verb compositions.In this work, we investigate different methods to improve compositional generalization by planning the syntactic structure of a caption.Our experiments show that jointly modeling tokens and syntactic tags enhances generalization in both RNNand Transformer-based models, while also improving performance on standard metrics.
Emanuele Bugliarello, Desmond Elliott
EACL2
2021 Visually Grounded Reasoning across Languages and Cultures
abstract
The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems.
Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott
EMNLP (1)6
2021 Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
abstract
Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities.We propose a diagnostic method based on cross-modal input ablation to assess the extent to which these models actually integrate cross-modal information.This method involves ablating inputs from one modality, either entirely or selectively based on cross-modal grounding alignments, and evaluating the model prediction performance on the other modality.Model performance is measured by modality-specific tasks that mirror the model pretraining objectives (e.g.masked language modelling for text).Models that have learned to construct cross-modal representations using both modalities are expected to perform worse when inputs are missing from a modality.We find that recently proposed models have much greater relative difficulty predicting text when visual information is ablated, compared to predicting visual object categories when text is ablated, indicating that these models are not symmetrically cross-modal.
Stella Frank, Emanuele Bugliarello, Desmond Elliott
EMNLP (1)3
2021 Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs
abstract
Abstract Large-scale pretraining and task-specific fine- tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI. These models can be categorized into either single-stream or dual-stream encoders. We study the differences between these two categories, and show how they can be unified under a single theoretical framework. We then conduct controlled experiments to discern the empirical differences between five vision and language BERTs. Our experiments show that training data and hyperparameters are responsible for most of the differences between the reported results, but they also reveal that the embedding layer plays a crucial role in these massive models.
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott
Trans. Assoc. Comput. Linguistics4
2020 The Sensitivity of Language Models and Humans to Winograd Schema Perturbations
abstract
Large-scale pretrained language models are the major driving force behind recent improvements in performance on the Winograd Schema Challenge, a widely employed test of commonsense reasoning ability.We show, however, with a new diagnostic dataset, that these models are sensitive to linguistic perturbations of the Winograd examples that minimally affect human understanding.Our results highlight interesting differences between humans and language models: language models are more sensitive to number or gender alternations and synonym replacements than humans, and humans are more stable and consistent in their predictions, maintain a much higher absolute performance, and perform better on non-associative instances than associative ones.Overall, humans are correct more often than out-of-the-box models, and the models are sometimes right for the wrong reasons.Finally, we show that fine-tuning on a large, task-specific dataset can offer a solution to these issues.
Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, Anders Søgaard
ACL5
2020 On Forgetting to Cite Older Papers: An Analysis of the ACL Anthology
abstract
The field of natural language processing is experiencing a period of unprecedented growth, and with it a surge of published papers.This represents an opportunity for us to take stock of how we cite the work of other researchers, and whether this growth comes at the expense of "forgetting" about older literature.In this paper, we address this question through bibliographic analysis.We analyze the age of outgoing citations in papers published at selected ACL venues between 2010 and 2019, finding that there is indeed a tendency for recent papers to cite more recent work, but the rate at which papers older than 15 years are cited has remained relatively stable.
Marcel Bollmann, Desmond Elliott
ACL2
2020 CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
abstract
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon
ACL5
2020 Multimodal machine translation through visuals and speech
abstract
Abstract Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area are spoken language translation, image-guided translation, and video-guided translation, which exploit audio and visual modalities, respectively. These tasks are distinguished from their monolingual counterparts of speech recognition, image captioning, and video captioning by the requirement of models to generate outputs in a different language. This survey reviews the major data resources for these tasks, the evaluation campaigns concentrated around them, the state of the art in end-to-end and pipeline approaches, and also the challenges in performance evaluation. The paper concludes with a discussion of directions for future research in these areas: the need for more expansive and challenging datasets, for targeted evaluations of model performance, and for multimodality in both the input and output space.
Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, Jörg Tiedemann
Mach. Transl.5
2019 Compositional Generalization in Image Captioning
abstract
Image captioning models are usually evaluated on their ability to describe a held-out set of images, not on their ability to generalize to unseen concepts.We study the problem of compositional generalization, which measures how well a model composes unseen combinations of concepts when describing images.Stateof-the-art image captioning models show poor generalization performance on this task.We propose a multi-task model to address the poor performance, that combines caption generation and image-sentence ranking, and uses a decoding mechanism that re-ranks the captions according their similarity to the image.This model is substantially better at generalizing to unseen combinations of concepts compared to state-of-the-art captioning models.
Mitja Nikolaus, Mostafa Abdou, Matthew Lamm, Rahul Aralikatte, Desmond Elliott
CoNLL5
2019 Adversarial Removal of Demographic Attributes Revisited
abstract
Maria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Maria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard
EMNLP/IJCNLP (1)4
2018 Measuring the Diversity of Automatic Image Descriptions
abstract
Automatic image description systems typically produce generic sentences that only make use of a small subset of the vocabulary available to them. In this paper, we consider the production of generic descriptions as a lack of diversity in the output, which we quantify using established metrics and two new metrics that frame image description as a word recall task. This framing allows us to evaluate system performance on the head of the vocabulary, as well as on the long tail, where system performance degrades. We use these metrics to examine the diversity of the sentences generated by nine state-of-the-art systems on the MS COCO data set. We find that the systems trained with maximum likelihood objectives produce less diverse output than those trained with additional adversarial objectives. However, the adversarially-trained models only produce more types from the head of the vocabulary and not the tail. Besides vocabulary-based methods, we also look at the compositional capacity of the systems, specifically their ability to create compound nouns and prepositional phrases of different lengths. We conclude that there is still much room for improvement, and offer a toolkit to measure progress towards the goal of generating more diverse image descriptions.
Emiel van Miltenburg, Desmond Elliott, Piek Vossen
COLING2
2018 Lessons Learned in Multilingual Grounded Language Learning
abstract
Recent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language.Here, we investigate in detail which conditions affect the performance of this type of grounded language learning model.We show that multilingual training improves over bilingual training, and that low-resource languages benefit from training with higher-resource languages.We demonstrate that a multilingual model can be trained equally well on either translations or comparable sentence pairs, and that annotating the same set of images in multiple language enables further improvements via an additional caption-caption ranking objective.
Ákos Kádár, Desmond Elliott, Marc-Alexandre Côté, Grzegorz Chrupala, Afra Alishahi
CoNLL2
2018 Adversarial Evaluation of Multimodal Machine Translation
abstract
The promise of combining vision and language in multimodal machine translation is that systems will produce better translations by leveraging the image data.However, inconsistent results have lead to uncertainty about whether the images actually improve translation quality.We present an adversarial evaluation method to directly examine the utility of the image data in this task.Our evaluation measures whether multimodal translation systems perform better given either the congruent image or a random incongruent image, in addition to the correct source language sentence.We find that two out of three publicly available systems are sensitive to this perturbation of the data, and recommend that all systems pass this evaluation in the future.* Work carried out at the University of Edinburgh.Two dogs play with an orange toy in tall grass. ModelZwei Hunde spielen im hohen Gras mit einem orangen Spielzeug.
Desmond Elliott
EMNLP1
2018 Talking about other people: an endless range of possibilities
abstract
Image description datasets, such as Flickr30K and MS COCO, show a high degree of variation in the ways that crowdworkers talk about the world.Although this gives us a rich and diverse collection of data to work with, it also introduces uncertainty about how the world should be described.This paper shows the extent of this uncertainty in the PEOPLE domain.We present a taxonomy of different ways to talk about other people.This taxonomy serves as a reference point to think about how other people should be described, and can be used to classify and compute statistics about labels applied to people.
Emiel van Miltenburg, Desmond Elliott, Piek Vossen
INLG2
2018 Assessing multilingual multimodal image description: Studies of native speaker preferences and translator choices
abstract
Abstract Two studies on multilingual multimodal image description provide empirical evidence towards two questions at the core of the task: (i) whether target language speakers prefer descriptions generated directly in their native language, as compared to descriptions translated from a different language; (ii) whether images improve human translation of descriptions. These results provide guidance for future work in multimodal natural language processing by first showing that on the whole, translations are not distinguished from native language descriptions, and second delineating and quantifying the information gained from the image during the human translation task.
Stella Frank, Desmond Elliott, Lucia Specia
Nat. Lang. Eng.2
2017 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures (Extended Abstract)
abstract
Automatic image description generation is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the known approaches based on how they conceptualise this problem and provide a review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image-text datasets and the evaluation measures that have been developed to assess the quality of machine-generated descriptions. Finally we explore future directions in the area of automatic image description.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
IJCAI3
2017 Imagination Improves Multimodal Translation
abstract
We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.
Desmond Elliott, Ákos Kádár
IJCNLP(1)1
2017 Cross-linguistic differences and similarities in image descriptions
abstract
Automatic image description systems are commonly trained and evaluated on large image description datasets.Recently, researchers have started to collect such datasets for languages other than English.An unexplored question is how different these datasets are from English and, if there are any differences, what causes them to differ.This paper provides a crosslinguistic comparison of Dutch, English, and German image descriptions.We find that these descriptions are similar in many respects, but the familiarity of crowd workers with the subjects of the images has a noticeable influence on description specificity.
Emiel van Miltenburg, Desmond Elliott, Piek Vossen
INLG2
2016 1 Million Captioned Dutch Newspaper Images
Desmond Elliott, Martijn Kleppe
LREC1
2016 A Corpus of Images and Text in Online News
Laura Hollink, Adriatik Bedjeti, Martin van Harmelen, Desmond Elliott
LREC4
2016 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
abstract
Automatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the existing approaches based on how they conceptualize this problem, viz., models that cast description as either generation problem or as a retrieval problem over a visual or multimodal representational space. We provide a detailed review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image datasets and the evaluation measures that have been developed to assess the quality of machine-generated image descriptions. Finally we extrapolate future directions in the area of automatic image description generation.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
J. Artif. Intell. Res.3
2015 Describing Images using Inferred Visual Dependency Representations
abstract
Desmond Elliott, Arjen de Vries. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Desmond Elliott, Arjen P. de Vries
ACL (1)1
2014 Query-by-Example Image Retrieval using Visual Dependency Representations
Desmond Elliott, Victor Lavrenko, Frank Keller
COLING1
2013 Image Description using Visual Dependency Representations
abstract
Describing the main event of an image involves identifying the objects depicted and predicting the relationships between them.Previous approaches have represented images as unstructured bags of regions, which makes it difficult to accurately predict meaningful relationships between regions.In this paper, we introduce visual dependency representations to capture the relationships between the objects in an image, and hypothesize that this representation can improve image description.We test this hypothesis using a new data set of region-annotated images, associated with visual dependency representations and gold-standard descriptions.We describe two template-based description generation models that operate over visual dependency representations.In an image description task, we find that these models outperform approaches that rely on object proximity or corpus information to generate descriptions on both automatic measures and on human judgements.
Desmond Elliott, Frank Keller
EMNLP1
2010 Finding and filtering information for children
abstract
Children face several challenges when using information access systems. These include formulating queries, judging the relevance of documents, and focusing attention on interface cues, such as query suggestions, while typing queries. It has also been shown that children want a personalised Web experience and prefer content presented to them that matches their long-term entertainment and education needs. To this end, we have developed an interaction-based information filtering system to address these challenges.
Desmond Elliott, Richard Glassey, Tamara Polajnar, Leif Azzopardi
SIGIR1
2009 A proactive personalised retrieval system
abstract
We present a personalised retrieval system that captures explicit relevance feedback to build an evolving user profile with multiple aspects. The user profile is used to proactively retrieve results between search sessions to support multi-session search tasks. This approach to supporting users with their multi-session search tasks is evaluated in a between-subjects multiple time-series study with ten subjects performing two simulated work situation tasks over five sessions. System interaction data shows that subjects using the personalised retrieval system issue fewer queries and interact with fewer results than subjects using a baseline system. The interaction data also shows a trend of subjects interacting with the proactively retrieved results in the personalised retrieval system.
Desmond Elliott, Joemon M. Jose
CIKM1
2009 Aspect-based video browsing - A user study
abstract
In this paper, we present a user study on a novel video search interface based on the concept of aspect browsing. We aim to confirm whether automatically suggesting new aspects can increase the performance of an aspect-based browser. The proposed strategy is to assist the user in exploratory video search by actively suggesting new query terms and video shots. We use a clustering technique to identify potential aspects and use the results to propose suggestions to the user to help them in their search task. We evaluate this approach by analysing the users' perception and by exploiting the log files.
Frank Hopfgartner, Thierry Urruty, David Hannah, Desmond Elliott, Joemon M. Jose
ICME4