EDBT 2026 Demo / reviewers in the wild / expert
Emanuele Bugliarello
dblp:241/9497
· DBLP profile ↗
18ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0002-2999-7081ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human ratingabstractWhile text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt.
While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human annotations, evaluation task) has not been systematically measured.
We find that looking at only *one slice of data*, i.e. one set of capabilities or human annotations, is not enough to obtain stable conclusions that generalise to new conditions or slices when evaluating T2I models or alignment metrics.
We address this by introducing an evaluation suite of $>$100K annotations across four human annotation templates that comprehensively evaluates models' capabilities across a range of methods for gathering human annotations and comparing models.
In particular, we propose (1) a carefully curated set of prompts -- *Gecko2K*; (2) a statistically grounded method of comparing T2I models; and (3) how to systematically evaluate metrics under three *evaluation tasks* -- *model ordering, pair-wise instance scoring, point-wise instance scoring*.
Using this evaluation suite, we evaluate a wide range of metrics and find that a metric may do better in one setting but worse in another.
As a result, we introduce a new, interpretable auto-eval metric that is consistently better correlated with human ratings than such existing metrics on our evaluation suite--across different human templates and evaluation settings--and on TIFA160. Olivia Wiles, Isabela Albuquerque, Ivana Kajic, Su Wang 0001, Emanuele Bugliarello, Yasumasa Onoe, Pinelopi Papalampidi, Ira Ktena, Christopher Knutsen, Cyrus Rashtchian, Anant Nawalgaria, Jordi Pont-Tuset, Aida Nematzadeh |
ICLR | 6 |
| 2024 | No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsabstractWe study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the common filtering of training data to English image-text pairs disadvantages communities of lower socioeconomic status and negatively impacts cultural understanding. Notably, this performance gap is not captured by - and even at odds with - the currently popular evaluation metrics derived from the Western-centric ImageNet and COCO datasets. Second, pretraining with global, unfiltered data before fine-tuning on English content can improve cultural understanding without sacrificing performance on said popular benchmarks. Third, we introduce the task of geo-localization as a novel evaluation metric to assess cultural diversity in VLMs. Our work underscores the value of using diverse data to create more inclusive multimodal systems and lays the groundwork for developing VLMs that better represent global perspectives. Angeline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 0038, Andreas Steiner 0001, Xiaohua Zhai, Ibrahim Alabdulmohsin |
NeurIPS | 3 |
| 2023 | Measuring Progress in Fine-grained Vision-and-Language UnderstandingabstractEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Emanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh |
ACL (1) | 1 |
| 2023 | Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingabstractRecent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data.In particular, we propose two pretraining approaches to contextualise visual entities in a multimodal setup.With verbalised scene graphs, we transform visual relation triplets into structured captions, and treat them as additional image descriptions.With masked relation prediction, we further encourage relating entities from image regions with visually masked contexts.When applied to strong baselines pretrained on large amounts of Web data, zero-shot evaluations on both coarse-grained and fine-grained tasks show the efficacy of our methods in learning multimodal representations from weakly-supervised relations data. Emanuele Bugliarello, Aida Nematzadeh, Lisa Anne Hendricks |
EMNLP | 1 |
| 2023 | Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language ModelsabstractPretrained machine learning models are known to perpetuate and even amplify existing biases in data, which can result in unfair outcomes that ultimately impact user experience.Therefore, it is crucial to understand the mechanisms behind those prejudicial biases to ensure that model performance does not result in discriminatory behaviour toward certain groups or populations.In this work, we define gender bias as our case study.We quantify bias amplification in pretraining and after fine-tuning on three families of vision-and-language models.We investigate the connection, if any, between the two learning stages, and evaluate how bias amplification reflects on model performance.Overall, we find that bias amplification in pretraining and after fine-tuning are independent.We then examine the effect of continued pretraining on gender-neutral data, finding that this reduces group disparities, i.e., promotes fairness, on VQAv2 and retrieval tasks without significantly compromising task performance. Laura Cabello Piqueras, Emanuele Bugliarello, Stephanie Brandl, Desmond Elliott |
EMNLP | 2 |
| 2023 | Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, Desmond Elliott |
ICLR | 3 |
| 2023 | StoryBench: A Multifaceted Benchmark for Continuous Story VisualizationabstractGenerating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark for video generation requires data annotated over time, which contrasts with the single caption used often in video datasets. To fill this gap, we collect comprehensive human annotations on three existing datasets, and introduce StoryBench: a new, challenging multi-task benchmark to reliably evaluate forthcoming text-to-video models. Our benchmark includes three video generation tasks of increasing difficulty: action execution, where the next action must be generated starting from a conditioning video; story continuation, where a sequence of actions must be executed starting from a conditioning video; and story generation, where a video must be generated from only text prompts. We evaluate small yet strong text-to-video baselines, and show the benefits of training on story-like data algorithmically generated from existing video captions. Finally, we establish guidelines for human evaluation of video stories, and reaffirm the need of better automatic metrics for video generation. StoryBench aims at encouraging future research efforts in this exciting new area. Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mohammad Taghi Saffar, Han Zhang 0010, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kindermans, Paul Voigtlaender |
NeurIPS | 1 |
| 2022 | Challenges and Strategies in Cross-Cultural NLPabstractDaniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard |
ACL (1) | 7 |
| 2022 | IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesabstractReliable evaluation benchmarks designed for replicability and comprehensiveness have driven progress in machine learning. Due to the lack of a multilingual benchmark, however, vision-and-language research has mostly focused on English language tasks. To fill this gap, we introduce the Image-Grounded Language Understanding Evaluation benchmark. IGLUE brings together{—}by both aggregating pre-existing datasets and creating new ones{—}visual question answering, cross-modal retrieval, grounded reasoning, and grounded entailment tasks across 20 diverse languages. Our benchmark enables the evaluation of multilingual multimodal models for transfer learning, not only in a zero-shot setting, but also in newly defined few-shot learning setups. Based on the evaluation of the available state-of-the-art models, we find that translate-test transfer is superior to zero-shot transfer and that few-shot learning is hard to harness for many tasks. Moreover, downstream performance is partially explained by the amount of available unlabelled textual data for pretraining, and only weakly by the typological distance of target{–}source languages. We hope to encourage future research efforts in this area by releasing the benchmark to the community. Emanuele Bugliarello, Fangyu Liu 0001, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, Ivan Vulic |
ICML | 1 |
| 2022 | Mostra: A Flexible Balancing Framework to Trade-off User, Artist and Platform Objectives for Music SequencingabstractWe consider the task of sequencing tracks on music streaming platforms where the goal is to maximise not only user satisfaction, but also artist- and platform-centric objectives, needed to ensure long-term health and sustainability of the platform. Grounding the work across four objectives: Sat, Discovery, Exposure and Boost, we highlight the need and the potential to trade-off performance across these objectives, and propose Mostra, a Set Transformer-based encoder-decoder architecture equipped with submodular multi-objective beam search decoding. The proposed model affords system designers the power to balance multiple goals, and dynamically control the impact on one objective to satisfy other objectives. Through extensive experiments on data from a large-scale music streaming platform, we present insights on the trade-offs that exist across different objectives, and demonstrate that the proposed framework leads to a superior, just-in-time balancing across the various metrics of interest. Emanuele Bugliarello, Rishabh Mehrotra, James Kirk, Mounia Lalmas-Roelleke |
WWW | 1 |
| 2021 | On Language Models for CreolesabstractCreole languages such as Nigerian Pidgin English and Haitian Creole are under-resourced and largely ignored in the NLP literature.Creoles typically result from the fusion of a foreign language with multiple local languages, and what grammatical and lexical features are transferred to the creole is a complex process (Sessarego, 2020).While creoles are generally stable, the prominence of some features may be much stronger with certain demographics or in some linguistic situations (Winford, 1999;Patrick, 1999).This paper makes several contributions: We collect existing corpora and release models for Haitian Creole, Nigerian Pidgin English, and Singaporean Colloquial English.We evaluate these models on intrinsic and extrinsic tasks.Motivated by the above literature, we compare standard language models with distributionally robust ones and find that, somewhat surprisingly, the standard language models are superior to the distributionally robust ones.We investigate whether this is an effect of overparameterization or relative distributional stability, and find that the difference persists in the absence of over-parameterization, and that drift is limited, confirming the relative stability of creole languages. Heather C. Lent, Emanuele Bugliarello, Miryam de Lhoneux, Chen Qiu 0005, Anders Søgaard |
CoNLL | 2 |
| 2021 | The Role of Syntactic Planning in Compositional Image CaptioningabstractImage captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different distributions of images.Recently, Nikolaus et al. (2019) introduced a dataset to assess compositional generalization in image captioning, where models are evaluated on their ability to describe images with unseen adjective-noun and noun-verb compositions.In this work, we investigate different methods to improve compositional generalization by planning the syntactic structure of a caption.Our experiments show that jointly modeling tokens and syntactic tags enhances generalization in both RNNand Transformer-based models, while also improving performance on standard metrics. Emanuele Bugliarello, Desmond Elliott |
EACL | 1 |
| 2021 | Visually Grounded Reasoning across Languages and CulturesabstractThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems. Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott |
EMNLP (1) | 2 |
| 2021 | Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal TransformersabstractPretrained vision-and-language BERTs aim to learn representations that combine information from both modalities.We propose a diagnostic method based on cross-modal input ablation to assess the extent to which these models actually integrate cross-modal information.This method involves ablating inputs from one modality, either entirely or selectively based on cross-modal grounding alignments, and evaluating the model prediction performance on the other modality.Model performance is measured by modality-specific tasks that mirror the model pretraining objectives (e.g.masked language modelling for text).Models that have learned to construct cross-modal representations using both modalities are expected to perform worse when inputs are missing from a modality.We find that recently proposed models have much greater relative difficulty predicting text when visual information is ablated, compared to predicting visual object categories when text is ablated, indicating that these models are not symmetrically cross-modal. Stella Frank, Emanuele Bugliarello, Desmond Elliott |
EMNLP (1) | 2 |
| 2021 | Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTsabstractAbstract Large-scale pretraining and task-specific fine- tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI. These models can be categorized into either single-stream or dual-stream encoders. We study the differences between these two categories, and show how they can be unified under a single theoretical framework. We then conduct controlled experiments to discern the empirical differences between five vision and language BERTs. Our experiments show that training data and hyperparameters are responsible for most of the differences between the reported results, but they also reveal that the embedding layer plays a crucial role in these massive models. Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | It's Easier to Translate out of English than into it: Measuring Neural Translation Difficulty by Cross-Mutual InformationabstractThe performance of neural machine translation systems is commonly evaluated in terms of BLEU. However, due to its reliance on target language properties and generation, the BLEU metric does not allow an assessment of which translation directions are more difficult to model. In this paper, we propose cross-mutual information (XMI): an asymmetric information-theoretic metric of machine translation difficulty that exploits the probabilistic nature of most neural machine translation models. XMI allows us to better evaluate the difficulty of translating text into the target language while controlling for the difficulty of the target-side generation component independent of the translation task. We then present the first systematic and controlled study of cross-lingual translation difficulties using modern neural translation systems. Code for replicating our experiments is available online at https://github.com/e-bug/nmt-difficulty. Emanuele Bugliarello, Sabrina J. Mielke, Antonios Anastasopoulos, Ryan Cotterell, Naoaki Okazaki |
ACL | 1 |
| 2020 | Enhancing Machine Translation with Dependency-Aware Self-AttentionabstractMost neural machine translation models only rely on pairs of parallel sentences, assuming syntactic information is automatically learned by an attention mechanism.In this work, we investigate different approaches to incorporate syntactic knowledge in the Transformer model and also propose a novel, parameter-free, dependency-aware self-attention mechanism that improves its translation quality, especially for long sentences and in low-resource scenarios.We show the efficacy of each approach on WMT English↔German and English→Turkish, and WAT English→Japanese translation tasks. Emanuele Bugliarello, Naoaki Okazaki |
ACL | 1 |
| 2019 | Matrix Completion in the Unit Hypercube via Structured Matrix FactorizationabstractSeveral complex tasks that arise in organizations can be simplified by mapping them into a matrix completion problem. In this paper, we address a key challenge faced by our company: predicting the efficiency of artists in rendering visual effects (VFX) in film shots. We tackle this challenge by using a two-fold approach: first, we transform this task into a constrained matrix completion problem with entries bounded in the unit interval [0,1]; second, we propose two novel matrix factorization models that leverage our knowledge of the VFX environment. Our first approach, expertise matrix factorization (EMF), is an interpretable method that structures the latent factors as weighted user-item interplay. The second one, survival matrix factorization (SMF), is instead a probabilistic model for the underlying process defining employees' efficiencies. We show the effectiveness of our proposed models by extensive numerical tests on our VFX dataset and two additional datasets with values that are also bounded in the [0,1] interval. Emanuele Bugliarello, Swayambhoo Jain, Vineeth Rakesh |
IJCAI | 1 |