EDBT 2026 Demo / reviewers in the wild / expert
Diego Frassinelli
dblp:176/0216
· DBLP profile ↗
18ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-1517-2185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fruitcakes and Cupcakes Emerging from Noise: The ComposiGen Dataset of Compounds and Their Compositionality
Jule Godbersen, Sinan Kurtyigit, Emma Raimundo Schulz, Tonmoy Rakshit, Diego Frassinelli, Sabine Schulte im Walde, Carina Silberer |
LREC | 5 |
| 2026 | Resource-Lean Lexicon Induction for German Dialects
Robert Litschko, Barbara Plank, Diego Frassinelli |
LREC | 3 |
| 2026 | I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
Shijia Zhou, Saif M. Mohammad, Barbara Plank, Diego Frassinelli |
LREC | 4 |
| 2026 | Literally Concrete or Figuratively Abstract? Multilingual Concreteness Norms for Verb-Object ExpressionsabstractAbstract While existing concreteness norms primarily target words in isolation, little attention has been paid to concreteness in context. To address this, we systematically collect multilingual concreteness ratings using Best-Worst Scaling (BWS) for 5,814 verb-direct object noun expressions in three languages with different degrees of resource availability: English, German, and Slovene. We identify consistent patterns where the concreteness of verb-noun combinations is more strongly influenced by the nominal object than the verb. Through comparative analyses on an English subset, we demonstrate that BWS guarantees more reliable concreteness judgments than traditional rating scales. Expanding beyond our human-generated data, we use traditional and LLM-based automatic extrapolation methods to generate a large-scale multilingual resource of over 430,000 expressions. Additionally, we conduct a study examining the interaction between concreteness and literal vs. figurative judgments for a subset of 1,800 expressions in all three languages, along with example usage sentences. Our findings show that lower concreteness ratings correlate with figurative language, thus reinforcing the link between abstractness and figurativeness. All resources are available from https://github.com/urbikn/multilingual-concreteness-vo. Urban Knuples, Diego Frassinelli, Alexander Fraser 0001, Sabine Schulte im Walde |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies Between Model Predictions and Human Responses in VQAabstractLarge vision-language models struggle to accurately predict responses provided by multiple human annotators, particularly when those responses exhibit high uncertainty. In this study, we focus on a Visual Question Answering (VQA) task and comprehensively evaluate how well the output of the state-of-the-art vision-language model correlates with the distribution of human responses. To do so, we categorize our samples based on their levels (low, medium, high) of human uncertainty in disagreement (HUD) and employ, not only accuracy, but also three new human-correlated metrics for the first time in VQA, to investigate the impact of HUD. We also verify the effect of common calibration and human calibration (Baan et al. 2022) on the alignment of models and humans. Our results show that even BEiT3, currently the best model for this task, struggles to capture the multi-label distribution inherent in diverse human responses. Additionally, we observe that the commonly used accuracy-oriented calibration technique adversely affects BEiT3’s ability to capture HUD, further widening the gap between model predictions and human distributions. In contrast, we show the benefits of calibrating models towards human distributions for VQA, to better align model confidence with human uncertainty. Our findings highlight that for VQA, the alignment between human responses and model predictions is understudied and is an important target for future studies. Diego Frassinelli, Barbara Plank |
AAAI | 2 |
| 2025 | Disentangling Subjectivity and Uncertainty for Hate Speech Annotation and Modeling using GazeabstractÖzge Alacam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Özge Alaçam, Sanne Hoeken, Andreas Säuberli, Hannes Gröner, Diego Frassinelli, Sina Zarrieß, Barbara Plank |
EMNLP | 5 |
| 2025 | AbsVis - Benchmarking How Humans and Vision-Language Models "See" Abstract Concepts in ImagesabstractAbstract concepts like mercy and peace often lack clear visual grounding, and thus challenge humans and models to provide suitable image representations. To address this challenge, we introduce AbsVis – a dataset of 675 images annotated with 14,175 concept–explanation attributions from humans and two Vision-Language Models (VLMs: Qwen and LLaVA), where each concept is accompanied by a textual explanation. We compare human and VLM attributions in terms of diversity, abstractness, and alignment, and find that humans attribute more varied concepts. AbsVis also includes 2,680 human preference judgments evaluating the quality of a subset of these annotations, showing that overlapping concepts (attributed by both humans and VLMs) are most preferred. Explanations clarify and strengthen the perceived attributions, both from humans and VLMs. Explanations clarify and strengthen the perceived attributions, both from human and VLMs. Finally, we show that VLMs can approximate human preferences and use them to fine-tune VLMs via Direct Preference Optimization (DPO), yielding improved alignments with preferred concept–explanation pairs. Tarun Tater, Diego Frassinelli, Sabine Schulte im Walde |
EMNLP | 2 |
| 2024 | GRIT: A Dataset of Group Reference Recognition in ItalianabstractFor the analysis of political discourse a reliable identification of group references, i.e., linguistic components that refer to individuals or groups of people, is useful. However, the task of automatically recognizing group references has not yet gained much attention within NLP. To address this gap, we introduce GRIT (Group Reference for Italian), a large-scale, multi-domain manually annotated dataset for group reference recognition in Italian. GRIT represents a new resource for automatic and generalizable recognition of group references. With this dataset, we aim to establish group reference recognition as a valid classification task, which extends the domain of Named Entity Recognition by expanding its focus to literal and figurative mentions of social groups. We verify the potential of achieving automated group reference recognition for Italian through an experiment employing a fine-tuned BERT model. Our experimental results substantiate the validity of the task, implying a huge potential for applying automated systems to multiple fields of analysis, such as political text or social media analysis. Sergio E. Zanotto, Qi Yu 0007, Miriam Butt, Diego Frassinelli |
LREC/COLING | 4 |
| 2024 | Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categoriesabstractThe visual representation of a concept varies significantly depending on its meaning and the context where it occurs; this poses multiple challenges both for vision and multimodal models. Our study focuses on concreteness, a well-researched lexical-semantic variable, using it as a case study to examine the variability in visual representations. We rely on images associated with approximately 1,000 abstract and concrete concepts extracted from two different datasets: Bing and YFCC. Our goals are: (i) evaluate whether visual diversity in the depiction of concepts can reliably distinguish between concrete and abstract concepts; (ii) analyze the variability of visual features across multiple images of the same concept through a nearest neighbor analysis; and (iii) identify challenging factors contributing to this variability by categorizing and annotating images. Our findings indicate that for classifying images of abstract versus concrete concepts, a combination of basic visual features such as color and texture is more effective than features extracted by more complex models like Vision Transformer (ViT). However, ViTs show better performances in the nearest neighbor analysis, emphasizing the need for a careful selection of visual features when analyzing conceptual variables through modalities other than text. Tarun Tater, Sabine Schulte im Walde, Diego Frassinelli |
EMNLP | 3 |
| 2024 | Generalizable Sarcasm Detection is Just Around the Corner, of Course!abstractHyewon Jang, Diego Frassinelli. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hyewon Jang, Diego Frassinelli |
NAACL-HLT | 2 |
| 2023 | Intended and Perceived Sarcasm Between Close Friends: What Triggers Sarcasm and What Gets Conveyed?
Hyewon Jang, Bettina Braun, Diego Frassinelli |
CogSci | 3 |
| 2023 | Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness ContinuumabstractHumans tend to strongly agree on ratings on a scale for extreme cases (e.g., a CAT is judged as very concrete), but judgements on mid-scale words exhibit more disagreement.Yet, collected rating norms are heavily exploited across disciplines.Our study focuses on concreteness ratings and (i) implements correlations and supervised classification to identify salient multimodal characteristics of mid-scale words, and (ii) applies a hard clustering to identify patterns of systematic disagreement across raters.Our results suggest to either fine-tune or filter midscale target words before utilising them. Urban Knuples, Diego Frassinelli, Sabine Schulte im Walde |
CoNLL | 2 |
| 2021 | Electrophysiological signatures of multimodal comprehension in second language
Diego Frassinelli, Jyrki Tuomainen, Sebastian Klavinskis-whiting, Gabriella Vigliocco |
CogSci | 3 |
| 2020 | Interpreting Attention Models with Human Visual Attention in Machine Reading ComprehensionabstractWhile neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention.In this paper, we propose a new method that leverages eye-tracking data to investigate the relationship between human visual attention and neural attention in machine reading comprehension.To this end, we introduce a novel 23 participant eye tracking dataset -MQA-RC, in which participants read movie plots and answered pre-defined questions.We compare state of the art networks based on long shortterm memory (LSTM), convolutional neural models (CNN) and XLNet Transformer architectures.We find that higher similarity to human attention and performance significantly correlates to the LSTM and CNN models.However, we show this relationship does not hold true for the XLNet models -despite the fact that the XLNet performs best on this challenging task.Our results suggest that different architectures seem to learn rather different neural attention strategies and similarity of neural to human attention does not guarantee best performance. Ekta Sood, Simon Tannert, Diego Frassinelli, Andreas Bulling, Ngoc Thang Vu |
CoNLL | 3 |
| 2015 | Cumulative Contextual Facilitation in Word Activation and Processing: Evidence from Distributional Modelling
Diego Frassinelli, Frank Keller |
CogSci | 1 |
| 2013 | The Effect of Incremental Context on Conceptual Processing: Evidence from Visual World and Reading Experiments
Diego Frassinelli, Frank Keller, Christoph Scheepers |
CogSci | 1 |
| 2012 | The Plausibility of Semantic Properties Generated by a Distributional Model: Evidence from a Visual World Experiment
Diego Frassinelli, Frank Keller |
CogSci | 1 |
| 2012 | Concepts in context: Evidence from a feature-norming study
Diego Frassinelli, Alessandro Lenci |
CogSci | 1 |