VLDB 2026 Research / reviewers in the wild / expert
Sandro Pezzelle
dblp:182/2260
· DBLP profile ↗
18ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-3969-7445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Tools to Teammates: Evaluating LLMs in Multi-Session Coding InteractionsabstractNathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici |
ACL (1) | 7 |
| 2024 | Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic ProcessesabstractThere is an intricate relation between the properties of an image and how humans behave while describing the image.This behavior shows ample variation, as manifested in human signals such as eye movements and when humans start to describe the image.Despite the value of such signals of visuo-linguistic variation, they are virtually disregarded in the training of current pretrained models, which motivates further investigation.Using a corpus of Dutch image descriptions with concurrently collected eye-tracking data, we explore the nature of the variation in visuo-linguistic signals, and find that they correlate with each other.Given this result, we hypothesize that variation stems partly from the properties of the images, and explore whether image representations encoded by pretrained vision encoders can capture such variation.Our results indicate that pretrained models do so to a weak-to-moderate degree, suggesting that the models lack biases about what makes a stimulus complex for humans and what leads to variations in human outputs. Ece Takmaz, Sandro Pezzelle, Raquel Fernández |
EACL (1) | 2 |
| 2023 | Dealing with Semantic Underspecification in Multimodal NLPabstractIntelligent systems that aim at mastering language as humans do must deal with its semantic underspecification, namely, the possibility for a linguistic signal to convey only part of the information needed for communication to succeed.Consider the usages of the pronoun they, which can leave the gender and number of its referent(s) underspecified.Semantic underspecification is not a bug but a crucial language feature that boosts its storage and processing efficiency.Indeed, human speakers can quickly and effortlessly integrate semanticallyunderspecified linguistic signals with a wide range of non-linguistic information, e.g., the multimodal context, social or cultural conventions, and shared knowledge.Standard NLP models have, in principle, no or limited access to such extra information, while multimodal systems grounding language into other modalities, such as vision, are naturally equipped to account for this phenomenon.However, we show that they struggle with it, which could negatively affect their performance and lead to harmful consequences when used for applications.In this position paper, we argue that our community should be aware of semantic underspecification if it aims to develop language technology that can successfully interact with human users.We discuss some applications where mastering it is crucial and outline a few directions toward achieving this goal. Sandro Pezzelle |
ACL (1) | 1 |
| 2023 | A Psycholinguistic Analysis of BERT's Representations of CompoundsabstractThis work studies the semantic representations learned by BERT for compounds, that is, expressions such as sunlight or bodyguard.We build on recent studies that explore semantic information in Transformers at the word level and test whether BERT aligns with human semantic intuitions when dealing with expressions (e.g., sunlight) whose overall meaning depends-to a various extent-on the semantics of the constituent words (sun, light).We leverage a dataset that includes human judgments on two psycholinguistic measures of compound semantic analysis: lexeme meaning dominance (LMD; quantifying the weight of each constituent toward the compound meaning) and semantic transparency (ST; evaluating the extent to which the compound meaning is recoverable from the constituents' semantics).We show that BERT-based measures moderately align with human intuitions, especially when using contextualized representations, and that LMD is overall more predictable than ST.Contrary to the results reported for 'standard' words, higher, more contextualized layers are the best at representing compound meaning.These findings shed new light on the abilities of BERT in dealing with fine-grained semantic phenomena.Moreover, they can provide insights into how speakers represent compounds. Lars Buijtelaar, Sandro Pezzelle |
EACL | 2 |
| 2023 | When Language Models Fall in Love: Animacy Processing in Transformer Language ModelsabstractAnimacy-whether an entity is alive and sentient-is fundamental to cognitive processing, impacting areas such as memory, vision, and language.However, animacy is not always expressed directly in language: in English it often manifests indirectly, in the form of selectional constraints on verbs and adjectives.This poses a potential issue for transformer language models (LMs): they often train only on text, and thus lack access to extralinguistic information from which humans learn about animacy.We ask: how does this impact LMs' animacy processing-do they still behave as humans do?We answer this question using open-source LMs.Like previous studies, we find that LMs behave much like humans when presented with entities whose animacy is typical.However, we also show that even when presented with stories about atypically animate entities, such as a peanut in love, LMs adapt: they treat these entities as animate, though they do not adapt as well as humans.Even when the context indicating atypical animacy is very short, LMs pick up on subtle clues and change their behavior.We conclude that despite the limited signal through which LMs can learn about animacy, they are indeed sensitive to the relevant lexical semantic nuances available in English. Michael Hanna 0001, Yonatan Belinkov, Sandro Pezzelle |
EMNLP | 3 |
| 2023 | The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal ModelsabstractDespite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction.In this work, we explore to what extent they handle basic linguistic constructions-active-passive voice, coordination, and relative clauses-that even preschool children can typically master.We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities.We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings.Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples.Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting.This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities. Xinyi Chen 0005, Raquel Fernández, Sandro Pezzelle |
EMNLP | 3 |
| 2023 | GROOViST: A Metric for Grounding Objects in Visual StorytellingabstractA proper evaluation of stories generated for a sequence of images-the task commonly referred to as visual storytelling-must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding.In this work, we focus on evaluating the degree of grounding, that is, the extent to which a story is about the entities shown in the images.We analyze current metrics, both designed for this purpose and for general vision-text alignment.Given their observed shortcomings, we propose a novel evaluation tool, GROOViST, that accounts for cross-modal dependencies, temporal misalignments (the fact that the order in which entities appear in the story and the image sequence may not match), and human intuitions on visual grounding.An additional advantage of GROOViST is its modular design, where the contribution of each component can be assessed and interpreted individually. Aditya K. Surikuchi, Sandro Pezzelle, Raquel Fernández |
EMNLP | 2 |
| 2022 | Time Alignment between Gaze and Speech in Image Descriptions: Exploring Theories of Linearization
Ece Takmaz, Sandro Pezzelle, Raquel Fernández |
CogSci | 2 |
| 2021 | EaSe: A Diagnostic Tool for VQA based on Answer DiversityabstractWe propose EASE, a simple diagnostic tool for Visual Question Answering (VQA) which quantifies the difficulty of an image, question sample.EASE is based on the pattern of answers provided by multiple annotators to a given question.In particular, it considers two aspects of the answers: (i) their Entropy; (ii) their Semantic content.First, we prove the validity of our diagnostic to identify samples that are easy/hard for state-of-art VQA models.Second, we show that EASE can be successfully used to select the most-informative samples for training/fine-tuning. Crucially, only information that is readily available in any VQA dataset is used to compute its scores. 1 Shailza Jolly, Sandro Pezzelle, Moin Nabi |
NAACL-HLT | 2 |
| 2021 | Word Representation Learning in Multimodal Pre-Trained Transformers: An Intrinsic EvaluationabstractAbstract This study carries out a systematic intrinsic evaluation of the semantic representations learned by state-of-the-art pre-trained multimodal Transformers. These representations are claimed to be task-agnostic and shown to help on many downstream language-and-vision tasks. However, the extent to which they align with human semantic intuitions remains unclear. We experiment with various models and obtain static word representations from the contextualized ones they learn. We then evaluate them against the semantic judgments provided by human speakers. In line with previous evidence, we observe a generalized advantage of multimodal representations over language- only ones on concrete word pairs, but not on abstract ones. On the one hand, this confirms the effectiveness of these models to align language and vision, which results in better semantic representations for concepts that are grounded in images. On the other hand, models are shown to follow different representation learning patterns, which sheds some light on how and when they perform multimodal integration. Sandro Pezzelle, Ece Takmaz, Raquel Fernández |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Asking questions with a big impact: Adapting to other interpretations of gradable adjectives
Sandro Pezzelle, Raquel Fernández |
CogSci | 1 |
| 2020 | Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational ContextsabstractDialogue participants often refer to entities or situations repeatedly within a conversation, which contributes to its cohesiveness.Subsequent references exploit the common ground accumulated by the interlocutors and hence have several interesting properties, namely, they tend to be shorter and reuse expressions that were effective in previous mentions.In this paper, we tackle the generation of first and subsequent references in visually grounded dialogue.We propose a generation model that produces referring utterances grounded in both the visual and the conversational context.To assess the referring effectiveness of its output, we also implement a reference resolution system.Our experiments and analyses show that the model produces better, more effective referring utterances than a model not grounded in the dialogue context, and generates subsequent references that exhibit linguistic patterns akin to humans. Ece Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair, Raquel Fernández |
EMNLP (1) | 3 |
| 2020 | Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human GazeabstractWhen speakers describe an image, they tend to look at objects before mentioning them.In this paper, we investigate such sequential crossmodal alignment by modelling the image description generation process computationally.We take as our starting point a state-of-theart image captioning system and develop several model variants that exploit information from human gaze patterns recorded during language production.In particular, we propose the first approach to image description generation where visual processing is modelled sequentially.Our experiments and analyses confirm that better descriptions can be obtained by exploiting gaze-driven attention and shed light on human cognitive processes by comparing different ways of aligning the gaze modality with language production.We find that processing gaze data sequentially leads to descriptions that are better aligned to those produced by speakers, more diverse, and more naturalparticularly when gaze is encoded with a dedicated recurrent component. Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández |
EMNLP (1) | 2 |
| 2019 | Is the Red Square Big? MALeViC: Modeling Adjectives Leveraging Visual ContextsabstractSandro Pezzelle, Raquel Fernández. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sandro Pezzelle, Raquel Fernández |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Comparatives, Quantifiers, Proportions: a Multi-Task Model for the Learning of Quantities from VisionabstractComunicació presentada a la Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2018), celebrada els dies 1 a 6 de juny de 2018 a Nova Orleans, Estats Units d'Amèrica. Sandro Pezzelle, Ionut Sorodoc, Raffaella Bernardi |
NAACL-HLT | 1 |
| 2018 | Learning quantification from images: A structured neural architectureabstractAbstract Major advances have recently been made in merging language and vision representations. Most tasks considered so far have confined themselves to the processing of objects and lexicalised relations amongst objects (content words). We know, however, that humans (even pre-school children) can abstract over raw multimodal data to perform certain types of higher level reasoning, expressed in natural language byfunction words. A case in point is given by their ability to learn quantifiers, i.e. expressions likefew,someandall. From formal semantics and cognitive linguistics, we know that quantifiers are relations over sets which, as a simplification, we can see as proportions. For instance, inmost fish are red,mostencodes the proportion of fish which are red fish. In this paper, we study how well current neural network strategies model such relations. We propose a task where, given an image and a query expressed by an object–property pair, the system must return a quantifier expressing which proportions of the queried object have the queried property. Our contributions are twofold. First, we show that the best performance on this task involves coupling state-of-the-art attention mechanisms with a network architecture mirroring the logical structure assigned to quantifiers by classic linguistic formalisation. Second, we introduce a new balanced dataset of image scenarios associated with quantification queries, which we hope will foster further research in this area. Ionut Sorodoc, Sandro Pezzelle, Aurélie Herbelot, Mariella Dimiccoli |
Nat. Lang. Eng. | 2 |
| 2017 | FOIL it! Find One mismatch between Image and Language captionabstractRavi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi |
ACL (1) | 2 |
| 2016 | The LAMBADA dataset: Word prediction requiring a broad discourse contextabstractDenis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández |
ACL (1) | 6 |