VLDB 2026 Research / reviewers in the wild / expert
Elia Bruni
dblp:18/11004
· DBLP profile ↗
28ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | iVISPAR - An Interactive Visual-Spatial Reasoning Benchmark for VLMsabstractVision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. iVISPAR is based on a variant of the sliding tile puzzle—a classic problem that demands logical planning, spatial awareness, and multi-step reasoning. The benchmark supports visual 3D, 2D, and text-based input modalities, enabling comprehensive assessments of VLMs’ planning and reasoning skills. We evaluate a broad suite of state-of-the-art open-source and closed-source VLMs, comparing their performance while also providing optimal path solutions and a human baseline to assess the task’s complexity and feasibility for humans. Results indicate that while VLMs perform better on 2D tasks compared to 3D or text-based settings, they struggle with complex spatial configurations and consistently fall short of human performance, illustrating the persistent challenge of visual alignment. This underscores critical gaps in current VLM capabilities, highlighting their limitations in achieving human-level cognition. Project website: https://microcosm.ai/ivispar. Julius Mayer 0001, Mohamad Ballout, Serwan Jassim, Farbod N. Nezami, Elia Bruni |
EMNLP | 5 |
| 2024 | Interpretability of Language Models via Task SpacesabstractThe usual way to interpret language models (LMs) is to test their performance on different benchmarks and subsequently infer their internal processes.In this paper, we present an alternative approach, concentrating on the quality of LM processing, with a focus on their language abilities.To this end, we construct 'linguistic task spaces' -representations of an LM's language conceptualisation -that shed light on the connections LMs draw between language phenomena.Task spaces are based on the interactions of the learning signals from different linguistic phenomena, which we assess via a method we call 'similarity probing'.To disentangle the learning signals of linguistic phenomena, we further introduce a method called 'fine-tuning via gradient differentials' (FTGD).We apply our methods to language models of three different scales and find that larger models generalise better to overarching general concepts for linguistic tasks, making better use of their shared structure.Further, the distributedness of linguistic processing increases with pre-training through increased parameter sharing between related linguistic tasks.The overall generalisation patterns are mostly stable throughout training and not marked by incisive stages, potentially explaining the lack of successful curriculum strategies for LMs. Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes |
ACL (1) | 3 |
| 2024 | Context Shapes Emergent Communication about Concepts at Different Levels of AbstractionabstractWe study the communication of concepts at different levels of abstraction and in different contexts in an agent-based, interactive reference game. While playing the concept-level reference game, the neural network agents develop a communication system from scratch. We use a novel symbolic dataset that disentangles concept type (ranging from specific to generic) and context (ranging from fine to coarse) to study the influence of these factors on the emerging language. We compare two game scenarios: one in which speaker agents have access to context information (context-aware) and one in which the speaker agents do not have access to context information (context-unaware). First, we find that the agents learn higher-level concepts from the object inputs alone. Second, an analysis of the emergent communication system shows that only context-aware agents learn to communicate efficiently by adapting their messages to the context conditions and relying on context for unambiguous reference. Crucially, this behavior is not explicitly incentivized by the game, but efficient communication emerges and is driven by the availability of context alone. The emerging language we observe is reminiscent of evolutionary pressures on human languages and highlights the pivotal role of context in a communication system. Kristina Kobrock, Xenia Ohmer, Elia Bruni, Nicole Gotzner |
LREC/COLING | 3 |
| 2024 | GRASP: A Novel Benchmark for Evaluating Language GRounding and Situated Physics Understanding in Multimodal Language Models
Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, Elia Bruni |
IJCAI | 6 |
| 2024 | From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense ConsistencyabstractAbstract The staggering pace with which the capabilities of large language models (LLMs) are increasing, as measured by a range of commonly used natural language understanding (NLU) benchmarks, raises many questions regarding what “understanding” means for a language model and how it compares to human understanding. This is especially true since many LLMs are exclusively trained on text, casting doubt on whether their stellar benchmark performances are reflective of a true understanding of the problems represented by these benchmarks, or whether LLMs simply excel at uttering textual forms that correlate with what someone who understands the problem would say. In this philosophically inspired work, we aim to create some separation between form and meaning, with a series of tests that leverage the idea that world understanding should be consistent across presentational modes—inspired by Fregean senses—of the same meaning. Specifically, we focus on consistency across languages as well as paraphrases. Taking GPT-3.5 as our object of study, we evaluate multisense consistency across five different languages and various tasks. We start the evaluation in a controlled setting, asking the model for simple facts, and then proceed with an evaluation on four popular NLU benchmarks. We find that the model’s multisense consistency is lacking and run several follow-up analyses to verify that this lack of consistency is due to a sense-dependent task understanding. We conclude that, in this aspect, the understanding of LLMs is still quite far from being consistent and human-like, and deliberate on how this impacts their utility in the context of learning about human language and understanding. Xenia Ohmer, Elia Bruni, Dieuwke Hupkes |
Comput. Linguistics | 2 |
| 2023 | Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learningabstractFinding the best way of adapting pre-trained language models to a task is a big challenge in current NLP.Just like the previous generation of task-tuned models (TT), models that are adapted to tasks via in-context-learning (ICL) are robust in some setups but not in others.Here, we present a detailed analysis of which design choices cause instabilities and inconsistencies in LLM predictions.First, we show how spurious correlations between input distributions and labels -a known issue in TT models -form only a minor problem for prompted models.Then, we engage in a systematic, holistic evaluation of different factors that have been found to influence predictions in a prompting setup.We test all possible combinations of a range of factors on both vanilla and instructiontuned (IT) LLMs of different scale and statistically analyse the results to show which factors are the most influential, interactive or stable.Our results show which factors can be used without precautions and which should be avoided or handled with care in most settings. Lucas Weber, Elia Bruni, Dieuwke Hupkes |
CoNLL | 2 |
| 2022 | The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case StudyabstractObtaining human-like performance in NLP is often argued to require compositional generalisation.Whether neural networks exhibit this ability is usually studied by training models on highly compositional synthetic data.However, compositionality in natural language is much more complex than the rigid, arithmeticlike version such data adheres to, and artificial compositionality tests thus do not allow us to determine how neural models deal with more realistic forms of compositionality.In this work, we re-instantiate three compositionality tests from the literature and reformulate them for neural machine translation (NMT).Our results highlight that: i) unfavourably, models trained on more data are more compositional; ii) models are sometimes less compositional than expected, but sometimes more, exemplifying that different levels of compositionality are required, and models are not always able to modulate between them correctly; iii) some of the non-compositional behaviours are mistakes, whereas others reflect the natural variation in data.Apart from an empirical study, our work is a call to action: we should rethink the evaluation of compositionality in neural networks and develop benchmarks using real data to evaluate compositionality on natural language, where composing meaning is not as straightforward as doing the math. 1 Verna Dankers, Elia Bruni, Dieuwke Hupkes |
ACL (1) | 2 |
| 2022 | Emergence of Hierarchical Reference Systems in Multi-agent CommunicationabstractIn natural language, referencing objects at different levels of specificity is a fundamental pragmatic mechanism for efficient communication in context. We develop a novel communication game, the hierarchical reference game, to study the emergence of such reference systems in artificial agents. We consider a simplified world, in which concepts are abstractions over a set of primitive attributes (e.g., color, style, shape). Depending on how many attributes are combined, concepts are more general (“circle”) or more specific (“red dotted circle”). Based on the context, the agents have to communicate at different levels of this hierarchy. Our results show that the agents learn to play the game successfully and can even generalize to novel concepts. To achieve abstraction, they use implicit (omitting irrelevant information) and explicit (indicating that attributes are irrelevant) strategies. In addition, the compositional structure underlying the concept hierarchy is reflected in the emergent protocols, indicating that the need to develop hierarchical reference systems supports the emergence of compositionality. Xenia Ohmer, Marko Duda, Elia Bruni |
COLING | 3 |
| 2021 | Co-evolution of language and agents in referential gamesabstractReferential games offer a grounded learning environment for neural agents which accounts for the fact that language is functionally used to communicate.However, they do not take into account a second constraint considered to be fundamental for the shape of human language: that it must be learnable by new language learners.Cogswell et al. (2019) introduced cultural transmission within referential games through a changing population of agents to constrain the emerging language to be learnable.However, the resulting languages remain inherently biased by the agents' underlying capabilities.In this work, we introduce Language Transmission Simulator to model both cultural and architectural evolution in a population of agents.As our core contribution, we empirically show that the optimal situation is to take into account also the learning biases of the language learners and thus let language and agents coevolve.When we allow the agent population to evolve through architectural evolution, we achieve across the board improvements on all considered metrics and surpass the gains made with cultural transmission.These results stress the importance of studying the underlying agent architecture and pave the way to investigate the co-evolution of language and agent in language emergence studies. Gautier Dagan, Dieuwke Hupkes, Elia Bruni |
EACL | 3 |
| 2021 | Language Modelling as a Multi-Task ProblemabstractIn this paper, we propose to study language modelling as a multi-task problem, bringing together three strands of research: multitask learning, linguistics, and interpretability.Based on hypotheses derived from linguistic theory, we investigate whether language models adhere to learning principles of multi-task learning during training.To showcase the idea, we analyse the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items (NPIs).Our experiments demonstrate that a multi-task setting naturally emerges within the objective of the more general task of language modelling.We argue that this insight is valuable for multitask learning, linguistics and interpretability research and can lead to exciting new findings in all three domains. Lucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke Hupkes |
EACL | 3 |
| 2020 | Modelling Form-Meaning Systematicity with Linguistic and Visual Features
Arie Soeteman, E. Dario Gutiérrez, Elia Bruni, Ekaterina Shutova |
AAAI | 3 |
| 2020 | Location Attention for Extrapolation to Longer SequencesabstractNeural networks are surprisingly good at interpolating and perform remarkably well when the training set examples resemble those in the test set.However, they are often unable to extrapolate patterns beyond the seen data, even when the abstractions required for such patterns are simple.In this paper, we first review the notion of extrapolation, why it is important, and how one could hope to tackle it.We then focus on a specific type of extrapolation, which is especially useful for natural language processing: generalization to sequences longer than those seen during training.We hypothesize that models with a separate contentand location-based attention are more likely to extrapolate than those with common attention mechanisms.We empirically support our claim for recurrent seq2seq models with our proposed attention on variants of the Lookup Table task.This sheds light on some striking failures of neural models for sequences and on possible methods to approaching such issues. Yann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia Bruni |
ACL | 4 |
| 2020 | The Grammar of Emergent LanguagesabstractIn this paper, we consider the syntactic properties of languages emerged in referential games, using unsupervised grammar induction (UGI) techniques originally designed to analyse natural language.We show that the considered UGI techniques are appropriate to analyse emergent languages and we then study if the languages that emerge in a typical referential game setup exhibit syntactic structure, and to what extent this depends on the maximum message length and number of symbols that the agents are allowed to use.Our experiments demonstrate that a certain message length and vocabulary size are required for structure to emerge, but they also illustrate that more sophisticated game scenarios are required to obtain syntactic properties more akin to those observed in human language.We argue that UGI techniques should be part of the standard toolkit for analysing emergent languages and release a comprehensive library to facilitate such analysis for future researchers. Oskar van der Wal, Silvan de Boer, Elia Bruni, Dieuwke Hupkes |
EMNLP (1) | 3 |
| 2020 | Compositionality Decomposed: How do Neural Networks Generalise? (Extended Abstract)abstractDespite a multitude of empirical studies, little consensus exists on whether neural networks are able to generalise compositionally. As a response to this controversy, we present a set of tests that provide a bridge between, on the one hand, the vast amount of linguistic and philosophical theory about compositionality of language and, on the other, the successful neural models of language. We collect different interpretations of compositionality and translate them into five theoretically grounded tests for models that are formulated on a task-independent level. To demonstrate the usefulness of this evaluation paradigm, we instantiate these five tests on a highly compositional data set which we dub PCFG SET, apply the resulting tests to three popular sequence-to-sequence models and provide an in-depth analysis of the results. Dieuwke Hupkes, Verna Dankers, Mathijs Mul, Elia Bruni |
IJCAI | 4 |
| 2020 | Compositionality Decomposed: How do Neural Networks Generalise?abstractDespite a multitude of empirical studies, little consensus exists on whether neural networks are able to generalise compositionally, a controversy that, in part, stems from a lack of agreement about what it means for a neural model to be compositional. As a response to this controversy, we present a set of tests that provide a bridge between, on the one hand, the vast amount of linguistic and philosophical theory about compositionality of language and, on the other, the successful neural models of language. We collect different interpretations of compositionality and translate them into five theoretically grounded tests for models that are formulated on a task-independent level. In particular, we provide tests to investigate (i) if models systematically recombine known parts and rules (ii) if models can extend their predictions beyond the length they have seen in the training data (iii) if models’ composition operations are local or global (iv) if models’ predictions are robust to synonym substitutions and (v) if models favour rules or exceptions during training. To demonstrate the usefulness of this evaluation paradigm, we instantiate these five tests on a highly compositional data set which we dub PCFG SET and apply the resulting tests to three popular sequence-to-sequence models: a recurrent, a convolution-based and a transformer model. We provide an in-depth analysis of the results, which uncover the strengths and weaknesses of these three architectures and point to potential areas of improvement. Dieuwke Hupkes, Verna Dankers, Mathijs Mul, Elia Bruni |
J. Artif. Intell. Res. | 4 |
| 2019 | The PhotoBook Dataset: Building Common Ground through Visually-Grounded DialogueabstractThis paper introduces the PhotoBook dataset, a large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.Taking inspiration from seminal work on dialogue analysis, we propose a data-collection task formulated as a collaborative game prompting two online participants to refer to images utilising both their visual context as well as previously established referring expressions.We provide a detailed description of the task setup and a thorough analysis of the 2,500 dialogues collected.To further illustrate the novel features of the dataset, we propose a baseline model for reference resolution which uses a simple method to take into account shared information accumulated in a reference chain.Our results show that this information is particularly important to resolve later descriptions and underline the need to develop more sophisticated models of common ground in dialogue interaction. 1 Janosch Haber, Tim Baumgärtner, Ece Takmaz, Lieke Gelderloos, Elia Bruni, Raquel Fernández |
ACL (1) | 5 |
| 2018 | Ask No More: Deciding when to guess in referential visual dialogueabstractOur goal is to explore how the abilities brought in by a dialogue manager can be included in end-to-end visually grounded conversational agents. We make initial steps towards this general goal by augmenting a task-oriented visual dialogue model with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess. Our analyses show that adding a decision making component produces dialogues that are less repetitive and that include fewer unnecessary questions, thus potentially leading to more efficient and less unnatural interactions. Ravi Shekhar, Tim Baumgärtner, Aashish Venkatesh, Elia Bruni, Raffaella Bernardi, Raquel Fernández |
COLING | 4 |
| 2017 | Adversarial evaluation for open-domain dialogue generationabstractWe investigate the potential of adversarial evaluation methods for open-domain dialogue generation systems, comparing the performance of a discriminative agent to that of humans on the same task.Our results show that the task is hard, both for automated models and humans, but that a discriminative agent can learn patterns that lead to above-chance performance. Elia Bruni, Raquel Fernández |
SIGDIAL Conference | 1 |
| 2017 | Gland segmentation in colon histology images: The glas challenge contest
Korsuk Sirinukunwattana, Josien P. W. Pluim, Hao Chen 0011, Xiaojuan Qi 0001, Pheng-Ann Heng, Li Yang Wang, Bogdan J. Matuszewski, Elia Bruni, Urko Sanchez, Anton Böhm, Olaf Ronneberger, Bassem Ben Cheikh, Daniel Racoceanu, Philipp Kainz, Michael Pfeiffer 0001, Martin Urschler, David R. J. Snead, Nasir M. Rajpoot |
Medical Image Anal. | 9 |
| 2015 | Affective Analysis of Professional and Amateur Abstract Paintings Using Statistical Analysis and Art TheoryabstractWhen artists express their feelings through the artworks they create, it is believed that the resulting works transform into objects with “emotions” capable of conveying the artists' mood to the audience. There is little to no dispute about this belief: Regardless of the artwork, genre, time, and origin of creation, people from different backgrounds are able to read the emotional messages. This holds true even for the most abstract paintings. Could this idea be applied to machines as well? Can machines learn what makes a work of art “emotional”? In this work, we employ a state-of-the-art recognition system to learn which statistical patterns are associated with positive and negative emotions on two different datasets that comprise professional and amateur abstract artworks. Moreover, we analyze and compare two different annotation methods in order to establish the ground truth of positive and negative emotions in abstract art. Additionally, we use computer vision techniques to quantify which parts of a painting evoke positive and negative emotions. We also demonstrate how the quantification of evidence for positive and negative emotions can be used to predict which parts of a painting people prefer to focus on. This method opens new opportunities of research on why a specific painting is perceived as emotional at global and local scales. Andreza Sartori, Victoria Yanulevskaya, Almila Akdag Salah, Jasper R. R. Uijlings, Elia Bruni, Nicu Sebe |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2014 | Is this a wampimuk? Cross-modal mapping between distributional semantics and the visual worldabstractFollowing up on recent work on establishing a mapping between vector-based semantic embeddings of words and the visual representations of the corresponding objects from natural images, we first present a simple approach to cross-modal vector-based semantics for the task of zero-shot learning, in which an image of a previously unseen object is mapped to a linguistic representation denoting its word. We then introduce fast mapping, a challenging and more cognitively plausible variant of the zero-shot task, in which the learner is exposed to new objects and the corresponding words in very limited linguistic contexts. By combining prior linguistic and visual knowledge acquired about words and their objects, as well as exploiting the limited new evidence available, the learner must learn to associate new objects with words. Our results on this task pave the way to realistic simulations of how children or robots could use existing knowledge to bootstrap grounded semantic knowledge about new concepts. Angeliki Lazaridou, Elia Bruni, Marco Baroni |
ACL (1) | 2 |
| 2014 | Automatic Personality and Interaction Style Recognition from Facebook Profile PicturesabstractIn this paper, we address the issue of personality and interaction style recognition from profile pictures in Facebook. We recruited volunteers among Facebook users and collected a dataset of profile pictures, labeled with gold standard self-assessed personality and interaction style labels. Then, we exploited a bag-of-visual-words technique to extract features from pictures. Finally, different machine learning approaches were used to test the effectiveness of these features in predicting personality and interaction style traits. Our good results show that this task is very promising, because profile pictures convey a lot of information about a user and are directly connected to impression formation and identity management. Fabio Celli, Elia Bruni, Bruno Lepri |
ACM Multimedia | 2 |
| 2014 | Multimodal Distributional SemanticsabstractDistributional semantic models derive computational representations of word meaning from the patterns of co-occurrence of words in text. Such models have been a success story of computational linguistics, being able to provide reliable estimates of semantic relatedness for the many semantic tasks requiring them. However, distributional models extract meaning information exclusively from text, which is an extremely impoverished basis compared to the rich perceptual sources that ground human semantic knowledge. We address the lack of perceptual grounding of distributional models by exploiting computer vision techniques that automatically identify discrete visual words in images, so that the distributional representation of a word can be extended to also encompass its co-occurrence with the visual words of images it is associated with. We propose a flexible architecture to integrate text- and image-based distributional information, and we show in a set of empirical tests that our integrated model is superior to the purely text-based approach, and it provides somewhat complementary semantic information with respect to the latter. Elia Bruni, Nam-Khanh Tran, Marco Baroni |
J. Artif. Intell. Res. | 1 |
| 2013 | Of Words, Eyes and Brains: Correlating Image-Based Distributional Semantic Models with Neural Representations of ConceptsabstractTraditional distributional semantic models extract word meaning representations from cooccurrence patterns of words in text corpora.Recently, the distributional approach has been extended to models that record the cooccurrence of words with visual features in image collections.These image-based models should be complementary to text-based ones, providing a more cognitively plausible view of meaning grounded in visual perception.In this study, we test whether image-based models capture the semantic patterns that emerge from fMRI recordings of the neural signal.Our results indicate that, indeed, there is a significant correlation between image-based and brain-based semantic similarities, and that image-based models complement text-based ones, so that the best correlations are achieved when the two modalities are combined.Despite some unsatisfactory, but explained outcomes (in particular, failure to detect differential association of models with brain areas), the results show, on the one hand, that imagebased distributional semantic models can be a precious new tool to explore semantic representation in the brain, and, on the other, that neural data can be used as the ultimate test set to validate artificial semantic models in terms of their cognitive plausibility. Andrew J. Anderson, Elia Bruni, Ulisse Bordignon, Massimo Poesio, Marco Baroni |
EMNLP | 2 |
| 2012 | Distributional Semantics in Technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, Nam-Khanh Tran |
ACL (1) | 1 |
| 2012 | Distributional semantics with eyes: using image analysis to improve computational representations of word meaningabstractThe current trend in image analysis and multimedia is to use information extracted from text and text processing techniques to help vision-related tasks, such as automated image annotation and generating semantically rich descriptions of images. In this work, we claim that image analysis techniques can "return the favor" to the text processing community and be successfully used for a general-purpose representation of word meaning. We provide evidence that simple low-level visual features can enrich the semantic representation of word meaning with information that cannot be extracted from text alone, leading to improvement in the core task of estimating degrees of semantic relatedness between words, as well as providing a new, perceptually-enhanced angle on word semantics. Additionally, we show how distinguishing between a concept and its context in images can improve the quality of the word meaning representations extracted from images. Elia Bruni, Jasper R. R. Uijlings, Marco Baroni, Nicu Sebe |
ACM Multimedia | 1 |
| 2012 | In the eye of the beholder: employing statistical analysis and eye tracking for analyzing abstract paintingsabstractMost artworks are explicitly created to evoke a strong emotional response. During the centuries there were several art movements which employed different techniques to achieve emotional expressions conveyed by artworks. Yet people were always consistently able to read the emotional messages even from the most abstract paintings. Can a machine learn what makes an artwork emotional? In this work, we consider a set of 500 abstract paintings from Museum of Modern and Contemporary Art of Trento and Rovereto (MART), where each painting was scored as carrying a positive or negative response on a Likert scale of 1-7. We employ a state-of-the-art recognition system to learn which statistical patterns are associated with positive and negative emotions. Additionally, we dissect the classification machinery to determine which parts of an image evokes what emotions. This opens new opportunities to research why a specific painting is perceived as emotional. We also demonstrate how quantification of evidence for positive and negative emotions can be used to predict the way in which people observe paintings. Victoria Yanulevskaya, Jasper R. R. Uijlings, Elia Bruni, Andreza Sartori, Elisa Zamboni, Francesca Bacci, David Melcher, Nicu Sebe |
ACM Multimedia | 3 |
| 2012 | Automatic Analysis of Multimodal Requirements: A Research Preview
Elia Bruni, Alessio Ferrari 0001, Norbert Seyff, Gabriele Tolomei |
REFSQ | 1 |