Aida Nematzadeh

dblp:153/9556 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 11 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 8 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating
abstract
While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human annotations, evaluation task) has not been systematically measured. We find that looking at only *one slice of data*, i.e. one set of capabilities or human annotations, is not enough to obtain stable conclusions that generalise to new conditions or slices when evaluating T2I models or alignment metrics. We address this by introducing an evaluation suite of $>$100K annotations across four human annotation templates that comprehensively evaluates models' capabilities across a range of methods for gathering human annotations and comparing models. In particular, we propose (1) a carefully curated set of prompts -- *Gecko2K*; (2) a statistically grounded method of comparing T2I models; and (3) how to systematically evaluate metrics under three *evaluation tasks* -- *model ordering, pair-wise instance scoring, point-wise instance scoring*. Using this evaluation suite, we evaluate a wide range of metrics and find that a metric may do better in one setting but worse in another. As a result, we introduce a new, interpretable auto-eval metric that is consistently better correlated with human ratings than such existing metrics on our evaluation suite--across different human templates and evaluation settings--and on TIFA160.
Olivia Wiles, Isabela Albuquerque, Ivana Kajic, Su Wang 0001, Emanuele Bugliarello, Yasumasa Onoe, Pinelopi Papalampidi, Ira Ktena, Christopher Knutsen, Cyrus Rashtchian, Anant Nawalgaria, Jordi Pont-Tuset, Aida Nematzadeh
ICLR14
2024 A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 Frames
abstract
Understanding long, real-world videos requires modeling of long-range visual dependencies. To this end, we explore video-first architectures, building on the common paradigm of transferring large-scale, image–text models to video via shallow temporal fusion. However, we expose two limitations to the approach: (1) decreased spatial capabilities, likely due to poor video–language alignment in standard video datasets, and (2) higher memory consumption, bottlenecking the number of frames that can be processed. To mitigate the memory bottleneck, we systematically analyze the memory/accuracy trade-off of various efficient methods: factorized attention, parameter-efficient image-to-video adaptation, input masking, and multi-resolution patchification. Surprisingly, simply masking large portions of the video (up to 75%) during contrastive pre-training proves to be one of the most robust ways to scale encoders to videos up to 4.3 minutes at 1 FPS. Our simple approach for training long video-to-text models, which scales to 1B parameters, does not add new architectural complexity and is able to outperform the popular paradigm of using much larger LLMs as an information aggregator over segment-based information on benchmarks with long-range temporal dependencies (YouCook2, EgoSchema).
Pinelopi Papalampidi, Skanda Koppula, Shreya Pathak, Justin Chiu, Joseph Heyward, Viorica Patraucean, Antoine Miech, Andrew Zisserman, Aida Nematzadeh
CVPR10
2024 Evaluating Numerical Reasoning in Text-to-Image Models
abstract
Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the most advanced models have only rudimentary numerical skills. Specifically, their ability to correctly generate an exact number of objects in an image is limited to small numbers, it is highly dependent on the context the number term appears in, and it deteriorates quickly with each successive number. We also demonstrate that models have poor understanding of linguistic quantifiers (such as “few” or “as many as”), the concept of zero, and struggle with more advanced concepts such as fractional representations. We bundle prompts, generated images and human annotations into GeckoNum, a novel benchmark for evaluation of numerical reasoning.
Ivana Kajic, Olivia Wiles, Isabela Albuquerque, Su Wang 0001, Jordi Pont-Tuset, Aida Nematzadeh
NeurIPS7
2023 Measuring Progress in Fine-grained Vision-and-Language Understanding
abstract
Emanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Emanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh
ACL (1)5
2023 Evaluating Visual Number Discrimination in Deep Neural Networks
Ivana Kajic, Aida Nematzadeh
CogSci2
2023 MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot Prompting
abstract
Oscar Mañas, Pau Rodriguez Lopez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, Aishwarya Agrawal. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Oscar Mañas, Pau Rodríguez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, Aishwarya Agrawal
EACL4
2023 Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining
abstract
Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data.In particular, we propose two pretraining approaches to contextualise visual entities in a multimodal setup.With verbalised scene graphs, we transform visual relation triplets into structured captions, and treat them as additional image descriptions.With masked relation prediction, we further encourage relating entities from image regions with visually masked contexts.When applied to strong baselines pretrained on large amounts of Web data, zero-shot evaluations on both coarse-grained and fine-grained tasks show the efficacy of our methods in learning multimodal representations from weakly-supervised relations data.
Emanuele Bugliarello, Aida Nematzadeh, Lisa Anne Hendricks
EMNLP2
2022 A Systematic Investigation of Commonsense Knowledge in Large Language Models
abstract
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xiang Li 0069, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh
EMNLP6
2022 Flamingo: a Visual Language Model for Few-Shot Learning
abstract
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan
NeurIPS21
2021 Mutual Exclusivity as Competition in Cross-situational Word Learning
Zahra Shekarchi, Aida Nematzadeh, Thomas L. Griffiths 0001, Suzanne Stevenson
CogSci2
2021 Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
abstract
Abstract Recently, multimodal transformer models have gained popularity because their performance on downstream tasks suggests they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks, we study three important factors that can impact the quality of learned representations: pretraining data, the attention mechanism, and loss functions. By pretraining models on six datasets, we observe that dataset noise and language similarity to our downstream task are important indicators of model performance. Through architectural analysis, we learn that models with a multimodal attention mechanism can outperform deeper models with modality-specific attention mechanisms. Finally, we show that successful contrastive losses used in the self-supervised learning literature do not yield similar performance gains when used in multimodal transformers.
Lisa Anne Hendricks, John Mellor, Rosália G. Schneider, Jean-Baptiste Alayrac, Aida Nematzadeh
Trans. Assoc. Comput. Linguistics5
2020 Learning to Segment Actions from Observation and Narration
abstract
We apply a generative segmental model of task structure, guided by narration, to action segmentation in video.We focus on unsupervised and weakly-supervised settings where no action labels are known during training.Despite its simplicity, our model performs competitively with previous work on a dataset of naturalistic instructional videos.Our model allows us to vary the sources of supervision used in training, and we find that both task structure and narrative language provide large benefits in segmentation quality.
Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, Aida Nematzadeh
ACL6
2020 Tracing the Emergence of Gendered Language in Childhood
Ben Prystawski, Erin Grant, Aida Nematzadeh, Spike W. S. Lee, Suzanne Stevenson, Yang Xu 0023
CogSci3
2020 Visual Grounding in Video for Unsupervised Word Translation
abstract
There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -- it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -- all without any parallel corpora and simply by watching many videos of people speaking while doing things.
Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira 0001, Phil Blunsom, Andrew Zisserman
CVPR3
2018 Learning Hierarchical Visual Representations in Deep Neural Networks Using Hierarchical Linguistic Labels
Joshua C. Peterson, Paul Soulos, Aida Nematzadeh, Thomas L. Griffiths 0001
CogSci3
2018 Evaluating Theory of Mind in Question Answering
abstract
We propose a new dataset for evaluating question answering models with respect to their capacity to reason about beliefs.Our tasks are inspired by theory-of-mind experiments that examine whether children are able to reason about the beliefs of others, in particular when those beliefs differ from reality.We evaluate a number of recent neural models with memory augmentation.We find that all fail on our tasks, which require keeping track of inconsistent states of the world; moreover, the models' accuracy decreases notably when random sentences are introduced to the tasks at test. 1
Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, Thomas L. Griffiths 0001
EMNLP1
2017 How Can Memory-Augmented Neural Networks Pass a False-Belief Task?
Erin Grant, Aida Nematzadeh, Thomas L. Griffiths 0001
CogSci2
2017 Calculating Probabilities Simplifies Word Learning
Aida Nematzadeh, Barend Beekhuizen, Suzanne Stevenson
CogSci1
2017 Evaluating Vector-Space Models of Word Representation, or, The Unreasonable Effectiveness of Counting Words Near Other Words
Aida Nematzadeh, Stephan C. Meylan, Thomas L. Griffiths 0001
CogSci1
2016 The Interaction of Memory and Attention in Novel Word Generalization: A Computational Investigation
Erin Grant, Aida Nematzadeh, Suzanne Stevenson
CogSci2
2016 Simple Search Algorithms on Semantic Networks Learned from Language Use
Aida Nematzadeh, Filip Miscevic, Suzanne Stevenson
CogSci1
2015 A Computational Account of Novel Word Generalization
Aida Nematzadeh, Erin Grant, Suzanne Stevenson
CogSci1
2015 A Computational Cognitive Model of Novel Word Generalization
abstract
A key challenge in vocabulary acquisition is learning which of the many possible meanings is appropriate for a word. The word generalization problem refers to how children associate a word such as dog with a meaning at the appropriate category level in a taxonomy of objects, such as Dalma-tians, dogs, or animals. We present the first computational study of word general-ization integrated within a word-learning model. The model simulates child and adult patterns of word generalization in a word-learning task. These patterns arise due to the interaction of type and token frequencies in the input data, an influence often observed in people’s generalization of linguistic categories. 1
Aida Nematzadeh, Erin Grant, Suzanne Stevenson
EMNLP1
2014 Structural Differences in the Semantic Networks of Simulated Word Learners
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci1
2014 A Cognitive Model of Semantic Network Learning
abstract
Child semantic development includes learning the meaning of words as well as the semantic relations among words.A presumed outcome of semantic development is the formation of a semantic network that reflects this knowledge.We present an algorithm for simultaneously learning word meanings and gradually growing a semantic network, which adheres to the cognitive plausibility requirements of incrementality and limited computations.We demonstrate that the semantic connections among words in addition to their context is necessary in forming a semantic network that resembles an adult's semantic knowledge.
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
EMNLP1
2013 Word Learning in the Wild: What Natural Data Can Tell Us
Barend Beekhuizen, Afsaneh Fazly, Aida Nematzadeh, Suzanne Stevenson
CogSci3
2013 Desirable Difficulty in Learning: A Computational Investigation
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci1
2012 Interaction of Word Learning and Semantic Category Formation in Late Talking
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci1
2011 A Computational Study of Late Talking in Word-Meaning Acquisition
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci1