VLDB 2026 Research / reviewers in the wild / expert
Aida Nematzadeh
dblp:153/9556
· DBLP profile ↗
29ranked-venue papers
11as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 11 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 8 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human ratingabstractWhile text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt.
While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human annotations, evaluation task) has not been systematically measured.
We find that looking at only *one slice of data*, i.e. one set of capabilities or human annotations, is not enough to obtain stable conclusions that generalise to new conditions or slices when evaluating T2I models or alignment metrics.
We address this by introducing an evaluation suite of $>$100K annotations across four human annotation templates that comprehensively evaluates models' capabilities across a range of methods for gathering human annotations and comparing models.
In particular, we propose (1) a carefully curated set of prompts -- *Gecko2K*; (2) a statistically grounded method of comparing T2I models; and (3) how to systematically evaluate metrics under three *evaluation tasks* -- *model ordering, pair-wise instance scoring, point-wise instance scoring*.
Using this evaluation suite, we evaluate a wide range of metrics and find that a metric may do better in one setting but worse in another.
As a result, we introduce a new, interpretable auto-eval metric that is consistently better correlated with human ratings than such existing metrics on our evaluation suite--across different human templates and evaluation settings--and on TIFA160. Olivia Wiles, Isabela Albuquerque, Ivana Kajic, Su Wang 0001, Emanuele Bugliarello, Yasumasa Onoe, Pinelopi Papalampidi, Ira Ktena, Christopher Knutsen, Cyrus Rashtchian, Anant Nawalgaria, Jordi Pont-Tuset, Aida Nematzadeh |
ICLR | 14 |
| 2024 | A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 FramesabstractUnderstanding long, real-world videos requires modeling of long-range visual dependencies. To this end, we explore video-first architectures, building on the common paradigm of transferring large-scale, image–text models to video via shallow temporal fusion. However, we expose two limitations to the approach: (1) decreased spatial capabilities, likely due to poor video–language alignment in standard video datasets, and (2) higher memory consumption, bottlenecking the number of frames that can be processed. To mitigate the memory bottleneck, we systematically analyze the memory/accuracy trade-off of various efficient methods: factorized attention, parameter-efficient image-to-video adaptation, input masking, and multi-resolution patchification. Surprisingly, simply masking large portions of the video (up to 75%) during contrastive pre-training proves to be one of the most robust ways to scale encoders to videos up to 4.3 minutes at 1 FPS. Our simple approach for training long video-to-text models, which scales to 1B parameters, does not add new architectural complexity and is able to outperform the popular paradigm of using much larger LLMs as an information aggregator over segment-based information on benchmarks with long-range temporal dependencies (YouCook2, EgoSchema). Pinelopi Papalampidi, Skanda Koppula, Shreya Pathak, Justin Chiu, Joseph Heyward, Viorica Patraucean, Antoine Miech, Andrew Zisserman, Aida Nematzadeh |
CVPR | 10 |
| 2024 | Evaluating Numerical Reasoning in Text-to-Image ModelsabstractText-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the most advanced models have only rudimentary numerical skills. Specifically, their ability to correctly generate an exact number of objects in an image is limited to small numbers, it is highly dependent on the context the number term appears in, and it deteriorates quickly with each successive number. We also demonstrate that models have poor understanding of linguistic quantifiers (such as “few” or “as many as”), the concept of zero, and struggle with more advanced concepts such as fractional representations. We bundle prompts, generated images and human annotations into GeckoNum, a novel benchmark for evaluation of numerical reasoning. Ivana Kajic, Olivia Wiles, Isabela Albuquerque, Su Wang 0001, Jordi Pont-Tuset, Aida Nematzadeh |
NeurIPS | 7 |
| 2023 | Measuring Progress in Fine-grained Vision-and-Language UnderstandingabstractEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Emanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks, Aida Nematzadeh |
ACL (1) | 5 |
| 2023 | Evaluating Visual Number Discrimination in Deep Neural Networks
Ivana Kajic, Aida Nematzadeh |
CogSci | 2 |
| 2023 | MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot PromptingabstractOscar Mañas, Pau Rodriguez Lopez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, Aishwarya Agrawal. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Oscar Mañas, Pau Rodríguez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, Aishwarya Agrawal |
EACL | 4 |
| 2023 | Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingabstractRecent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data.In particular, we propose two pretraining approaches to contextualise visual entities in a multimodal setup.With verbalised scene graphs, we transform visual relation triplets into structured captions, and treat them as additional image descriptions.With masked relation prediction, we further encourage relating entities from image regions with visually masked contexts.When applied to strong baselines pretrained on large amounts of Web data, zero-shot evaluations on both coarse-grained and fine-grained tasks show the efficacy of our methods in learning multimodal representations from weakly-supervised relations data. Emanuele Bugliarello, Aida Nematzadeh, Lisa Anne Hendricks |
EMNLP | 2 |
| 2022 | A Systematic Investigation of Commonsense Knowledge in Large Language ModelsabstractXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xiang Li 0069, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh |
EMNLP | 6 |
| 2022 | Flamingo: a Visual Language Model for Few-Shot LearningabstractBuilding models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan |
NeurIPS | 21 |
| 2021 | Mutual Exclusivity as Competition in Cross-situational Word Learning
Zahra Shekarchi, Aida Nematzadeh, Thomas L. Griffiths 0001, Suzanne Stevenson |
CogSci | 2 |
| 2021 | Decoupling the Role of Data, Attention, and Losses in Multimodal TransformersabstractAbstract Recently, multimodal transformer models have gained popularity because their performance on downstream tasks suggests they learn rich visual-linguistic representations. Focusing on zero-shot image retrieval tasks, we study three important factors that can impact the quality of learned representations: pretraining data, the attention mechanism, and loss functions. By pretraining models on six datasets, we observe that dataset noise and language similarity to our downstream task are important indicators of model performance. Through architectural analysis, we learn that models with a multimodal attention mechanism can outperform deeper models with modality-specific attention mechanisms. Finally, we show that successful contrastive losses used in the self-supervised learning literature do not yield similar performance gains when used in multimodal transformers. Lisa Anne Hendricks, John Mellor, Rosália G. Schneider, Jean-Baptiste Alayrac, Aida Nematzadeh |
Trans. Assoc. Comput. Linguistics | 5 |
| 2020 | Learning to Segment Actions from Observation and NarrationabstractWe apply a generative segmental model of task structure, guided by narration, to action segmentation in video.We focus on unsupervised and weakly-supervised settings where no action labels are known during training.Despite its simplicity, our model performs competitively with previous work on a dataset of naturalistic instructional videos.Our model allows us to vary the sources of supervision used in training, and we find that both task structure and narrative language provide large benefits in segmentation quality. Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, Aida Nematzadeh |
ACL | 6 |
| 2020 | Tracing the Emergence of Gendered Language in Childhood
Ben Prystawski, Erin Grant, Aida Nematzadeh, Spike W. S. Lee, Suzanne Stevenson, Yang Xu 0023 |
CogSci | 3 |
| 2020 | Visual Grounding in Video for Unsupervised Word TranslationabstractThere are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -- it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -- all without any parallel corpora and simply by watching many videos of people speaking while doing things. Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira 0001, Phil Blunsom, Andrew Zisserman |
CVPR | 3 |
| 2018 | Learning Hierarchical Visual Representations in Deep Neural Networks Using Hierarchical Linguistic Labels
Joshua C. Peterson, Paul Soulos, Aida Nematzadeh, Thomas L. Griffiths 0001 |
CogSci | 3 |
| 2018 | Evaluating Theory of Mind in Question AnsweringabstractWe propose a new dataset for evaluating question answering models with respect to their capacity to reason about beliefs.Our tasks are inspired by theory-of-mind experiments that examine whether children are able to reason about the beliefs of others, in particular when those beliefs differ from reality.We evaluate a number of recent neural models with memory augmentation.We find that all fail on our tasks, which require keeping track of inconsistent states of the world; moreover, the models' accuracy decreases notably when random sentences are introduced to the tasks at test. 1 Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, Thomas L. Griffiths 0001 |
EMNLP | 1 |
| 2017 | How Can Memory-Augmented Neural Networks Pass a False-Belief Task?
Erin Grant, Aida Nematzadeh, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2017 | Calculating Probabilities Simplifies Word Learning
Aida Nematzadeh, Barend Beekhuizen, Suzanne Stevenson |
CogSci | 1 |
| 2017 | Evaluating Vector-Space Models of Word Representation, or, The Unreasonable Effectiveness of Counting Words Near Other Words
Aida Nematzadeh, Stephan C. Meylan, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2016 | The Interaction of Memory and Attention in Novel Word Generalization: A Computational Investigation
Erin Grant, Aida Nematzadeh, Suzanne Stevenson |
CogSci | 2 |
| 2016 | Simple Search Algorithms on Semantic Networks Learned from Language Use
Aida Nematzadeh, Filip Miscevic, Suzanne Stevenson |
CogSci | 1 |
| 2015 | A Computational Account of Novel Word Generalization
Aida Nematzadeh, Erin Grant, Suzanne Stevenson |
CogSci | 1 |
| 2015 | A Computational Cognitive Model of Novel Word GeneralizationabstractA key challenge in vocabulary acquisition is learning which of the many possible meanings is appropriate for a word. The word generalization problem refers to how children associate a word such as dog with a meaning at the appropriate category level in a taxonomy of objects, such as Dalma-tians, dogs, or animals. We present the first computational study of word general-ization integrated within a word-learning model. The model simulates child and adult patterns of word generalization in a word-learning task. These patterns arise due to the interaction of type and token frequencies in the input data, an influence often observed in people’s generalization of linguistic categories. 1 Aida Nematzadeh, Erin Grant, Suzanne Stevenson |
EMNLP | 1 |
| 2014 | Structural Differences in the Semantic Networks of Simulated Word Learners
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson |
CogSci | 1 |
| 2014 | A Cognitive Model of Semantic Network LearningabstractChild semantic development includes learning the meaning of words as well as the semantic relations among words.A presumed outcome of semantic development is the formation of a semantic network that reflects this knowledge.We present an algorithm for simultaneously learning word meanings and gradually growing a semantic network, which adheres to the cognitive plausibility requirements of incrementality and limited computations.We demonstrate that the semantic connections among words in addition to their context is necessary in forming a semantic network that resembles an adult's semantic knowledge. Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson |
EMNLP | 1 |
| 2013 | Word Learning in the Wild: What Natural Data Can Tell Us
Barend Beekhuizen, Afsaneh Fazly, Aida Nematzadeh, Suzanne Stevenson |
CogSci | 3 |
| 2013 | Desirable Difficulty in Learning: A Computational Investigation
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson |
CogSci | 1 |
| 2012 | Interaction of Word Learning and Semantic Category Formation in Late Talking
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson |
CogSci | 1 |
| 2011 | A Computational Study of Late Talking in Word-Meaning Acquisition
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson |
CogSci | 1 |