VLDB 2026 Research / reviewers in the wild / expert
Gemma Boleda
dblp:43/2178
· DBLP profile ↗
36ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0001-6140-7080ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Representation and self-supervised learning · 28% Language models and text generation · 26% Information extraction and text analysis · 21% | |
| Theoretical computer science
2 papers |
Information theory · 98% Logic in computer science · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 21 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
0.6 | 1 | 2022 | Communication breakdown: On the low mutual intelligibility between human and neural captioning · EMNLP 2022 |
Multimedia analysis and retrieval
image retrieval |
0.6 | 1 | 2022 | Communication breakdown: On the low mutual intelligibility between human and neural captioning · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.5 | 2 | 2018 | How to represent a word and predict it, too: improving tied architectures for language modelling · EMNLP 2018 Distributional vectors encode referential attributes · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.4 | 1 | 2020 | Probing for Referential Information in Language Models · ACL 2020 |
Machine learning › Representation and self-supervised learning
probing |
0.4 | 1 | 2020 | Probing for Referential Information in Language Models · ACL 2020 |
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings |
0.4 | 1 | 2019 | Putting Words in Context: LSTM Language Models and Lexical Ambiguity · ACL (1) 2019 |
Natural language and speech › Language models and text generation
neural language model |
0.3 | 1 | 2018 | How to represent a word and predict it, too: improving tied architectures for language modelling · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
distributional semantics |
0.3 | 2 | 2012 | First Order vs. Higher Order Modification in Distributional Semantics · EMNLP-CoNLL 2012 Distributional Semantics in Technicolor · ACL (1) 2012 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2016 | Convolutional Neural Network Language Models · EMNLP 2016 |
Natural language and speech › Language models and text generation › language modeling
word prediction |
0.2 | 1 | 2016 | The LAMBADA dataset: Word prediction requiring a broad discourse context · ACL (1) 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
attribute extraction |
0.2 | 1 | 2015 | Distributional vectors encode referential attributes · EMNLP 2015 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2022 | Communication breakdown: On the low mutual intelligibility between human and neural captioning · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation › structured regularization
weight tying |
0.1 | 1 | 2018 | How to represent a word and predict it, too: improving tied architectures for language modelling · EMNLP 2018 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2016 | Convolutional Neural Network Language Models · EMNLP 2016 |
Natural language and speech › Language models and text generation
natural language understanding |
0.1 | 1 | 2016 | The LAMBADA dataset: Word prediction requiring a broad discourse context · ACL (1) 2016 |
Machine learning › Learning paradigms
multi-label classification |
0.1 | 1 | 2007 | Modelling Polysemy in Adjective Classes by Multi-Label Classification · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis › word sense disambiguation
polysemy |
0.1 | 1 | 2007 | Modelling Polysemy in Adjective Classes by Multi-Label Classification · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 1 | 2007 | Modelling Polysemy in Adjective Classes by Multi-Label Classification · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.1 | 1 | 2007 | Modelling Polysemy in Adjective Classes by Multi-Label Classification · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis › natural language semantics › computational semantics
semantic role assignment |
0.0 | 1 | 2004 | The Influence of Argument Structure on Semantic Role Assignment · EMNLP 2004 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.0 | 1 | 2004 | The Influence of Argument Structure on Semantic Role Assignment · EMNLP 2004 |
Methods — techniques the papers use, named apart from their topics
zero-shot evaluation · 1.1human evaluation · 1.1probing · 0.8information-theoretic measures · 0.8information-theoretic measure · 0.8transformer · 0.4LSTM · 0.4LSTM language model · 0.4word2vec · 0.3tied architecture · 0.3dataset construction · 0.2convolutional neural network · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On the Use of Language and Vision Models for Cognitive Science: The Case of Naming Norms
Andreas Mädebach, Eleonora Gualdoni, Gemma Boleda |
CogSci | 4 |
| 2024 | Why do objects have many names? A study on word informativeness in language use and lexical systemsabstractHuman lexicons contain many different words that speakers can use to refer to the same object, e.g., purple or magenta for the same shade of color.On the one hand, studies on language use have explored how speakers adapt their referring expressions to successfully communicate in context, without focusing on properties of the lexical system.On the other hand, studies in language evolution have discussed how competing pressures for informativeness and simplicity shape lexical systems, without tackling in-context communication.We aim at bridging the gap between these traditions, and explore why a soft mapping between referents and words is a good solution for communication, by taking into account both in-context communication and the structure of the lexicon.We propose a simple measure of informativeness for words and lexical systems, grounded in a visual space, and analyze color naming data for English and Mandarin Chinese.We conclude that optimal lexical systems are those where multiple words can apply to the same referent, conveying different amounts of information.Such systems allow speakers to maximize communication accuracy and minimize the amount of information they convey when communicating about referents in contexts. Eleonora Gualdoni, Gemma Boleda |
EMNLP | 2 |
| 2023 | Quantifying informativeness of names in visual space
Eleonora Gualdoni, Charles Kemp, Yang Xu 0023, Gemma Boleda |
CogSci | 4 |
| 2023 | The Impact of Familiarity on Naming Variation: A Study on Object Naming in Mandarin ChineseabstractDifferent speakers often produce different names for the same object or entity (e.g., "woman" vs. "tourist" for a female tourist).The reasons behind variation in naming are not well understood.We create a Language and Vision dataset for Mandarin Chinese that provides an average of 20 names for 1319 naturalistic images, and investigate how familiarity with a given kind of object relates to the degree of naming variation it triggers across subjects.We propose that familiarity influences naming variation in two competing ways: increasing familiarity can either expand vocabulary, leading to higher variation, or promote convergence on conventional names, thereby reducing variation.We find evidence for both factors being at play.Our study illustrates how computational resources can be used to address research questions in Cognitive Science. Yunke He, Xixian Liao, Jialing Liang, Gemma Boleda |
CoNLL | 4 |
| 2022 | Woman or tennis player? Visual typicality and lexical frequency affect variation in object naming
Eleonora Gualdoni, Thomas Brochhagen, Andreas Mädebach, Gemma Boleda |
CogSci | 4 |
| 2022 | Effects of task and visual context on referring expressions using natural scenes
Andreas Mädebach, Ekaterina Torubarova, Eleonora Gualdoni, Gemma Boleda |
CogSci | 4 |
| 2022 | Communication breakdown: On the low mutual intelligibility between human and neural captioningabstractWe compare the 0-shot performance of a neural caption-based image retriever when given as input either human-produced captions or captions generated by a neural captioner.We conduct this comparison on the recently introduced IMAGECODE data-set (Krojer et al., 2022), which contains hard distractors nearly identical to the images to be retrieved.We find that the neural retriever has much higher performance when fed neural rather than human captions, despite the fact that the former, unlike the latter, were generated without awareness of the distractors that make the task hard.Even more remarkably, when the same neural captions are given to human subjects, their retrieval performance is almost at chance level.Our results thus add to the growing body of evidence that, even when the "language" of neural models resembles English, this superficial resemblance might be deeply misleading. Roberto Dessì, Eleonora Gualdoni, Francesca Franzon, Gemma Boleda, Marco Baroni |
EMNLP | 4 |
| 2021 | Does referent predictability affect the choice of referential form? A computational approach using masked coreference resolutionabstractIt is often posited that more predictable parts of a speaker's meaning tend to be made less explicit, for instance using shorter, less informative words.Studying these dynamics in the domain of referring expressions has proven difficult, with existing studies, both psycholinguistic and corpus-based, providing contradictory results.We test the hypothesis that speakers produce less informative referring expressions (e.g., pronouns vs. full noun phrases) when the context is more informative about the referent, using novel computational estimates of referent predictability.We obtain these estimates training an existing coreference resolution system for English on a new task, masked coreference resolution, giving us a probability distribution over referents that is conditioned on the context but not the referring expression.The resulting system retains standard coreference resolution performance while yielding a better estimate of human-derived referent predictability than previous attempts.A statistical analysis of the relationship between model output and mention form supports the hypothesis that predictability affects the form of a mention, both its morphosyntactic type and its length. Laura Aina, Xixian Liao, Gemma Boleda, Matthijs Westera |
CoNLL | 3 |
| 2020 | Probing for Referential Information in Language ModelsabstractLanguage models keep track of complex linguistic information about the preceding context -including, e.g., syntactic relations in a sentence.We investigate whether they also capture information beneficial for resolving pronominal anaphora in English.We analyze two state of the art models with LSTM and Transformer architectures, respectively, using probe tasks on a coreference annotated corpus.Our hypothesis is that language models will capture grammatical properties of anaphora (such as agreement between a pronoun and its antecedent), but not semantico-referential information (the fact that pronoun and antecedent refer to the same entity).Instead, we find evidence that models capture referential aspects to some extent -though they are still much better at grammar.The Transformer outperforms the LSTM in all analyses, and exhibits in particular better semantico-referential abilities. Ionut Sorodoc, Kristina Gulordava, Gemma Boleda |
ACL | 3 |
| 2020 | Modeling word interpretation with deep language models: The interaction between expectations and lexical information
Laura Aina, Thomas Brochhagen, Gemma Boleda |
CogSci | 3 |
| 2020 | Deep daxes: Mutual exclusivity arises through both learning biases and pragmatic strategies in neural networks
Kristina Gulordava, Thomas Brochhagen, Gemma Boleda |
CogSci | 3 |
| 2020 | Humans Meet Models on Object Naming: A New Dataset and AnalysisabstractWe release ManyNames v2 (MN v2), a verified version of an object naming dataset that contains dozens of valid names per object for 25K images.We analyze issues in the data collection method originally employed, standard in Language & Vision (L&V), and find that the main source of noise in the data comes from simulating a naming context solely from an image with a target object marked with a bounding box, which causes subjects to sometimes disagree regarding which object is the target.We also find that both the degree of this uncertainty in the original data and the amount of true naming variation in MN v2 differs substantially across object domains.We use MN v2 to analyze a popular L&V model and demonstrate its effectiveness on the task of object naming.However, our fine-grained analysis reveals that what appears to be human-like model behavior is not stable across domains, e.g., the model confuses people and clothing objects much more frequently than humans do.We also find that standard evaluations underestimate the actual effectiveness of the naming model: on the single-label names of the original dataset (Visual Genome), it obtains -27% accuracy points than on MN v2, that includes all valid object names. Carina Silberer, Sina Zarrieß, Matthijs Westera, Gemma Boleda |
COLING | 4 |
| 2020 | Object Naming in Language and Vision: A Survey and a New DatasetabstractPeople choose particular names for objects, such as dog or puppy for a given dog. Object naming has been studied in Psycholinguistics, but has received relatively little attention in Computational Linguistics. We review resources from Language and Vision that could be used to study object naming on a large scale, discuss their shortcomings, and create a new dataset that affords more opportunities for analysis and modeling. Our dataset, ManyNames, provides 36 name annotations for each of 25K objects in images selected from VisualGenome. We highlight the challenges involved and provide a preliminary analysis of the ManyNames data, showing that there is a high level of agreement in naming, on average. At the same time, the average number of name types associated with an object is much higher in our dataset than in existing corpora for Language and Vision, such that ManyNames provides a rich resource for studying phenomena like hierarchical variation (chihuahua vs. dog), which has been discussed at length in the theoretical literature, and other less well studied phenomena like cross-classification (cake vs. dessert). Carina Silberer, Sina Zarrieß, Gemma Boleda |
LREC | 3 |
| 2019 | Putting Words in Context: LSTM Language Models and Lexical AmbiguityabstractIn neural network models of language, words are commonly represented using contextinvariant representations (word embeddings) which are then put in context in the hidden layers.Since words are often ambiguous, representing the contextually relevant information is not trivial.We investigate how an LSTM language model deals with lexical ambiguity in English, designing a method to probe its hidden representations for lexical and contextual information about words.We find that both types of information are represented to a large extent, but also that there is room for improvement for contextual information. Laura Aina, Kristina Gulordava, Gemma Boleda |
ACL (1) | 3 |
| 2018 | How to represent a word and predict it, too: improving tied architectures for language modellingabstractRecent state-of-the-art neural language models share the representations of words given by the input and output mappings.We propose a simple modification to these architectures that decouples the hidden state from the word embedding prediction.Our architecture leads to comparable or better results compared to previous tied models and models without tying, with a much smaller number of parameters.We also extend our proposal to word2vec models, showing that tying is appropriate for general word prediction tasks. Kristina Gulordava, Laura Aina, Gemma Boleda |
EMNLP | 3 |
| 2017 | "Show Me the Cup": Reference with Continuous Representations
Marco Baroni, Gemma Boleda, Sebastian Padó |
CICLing (1) | 2 |
| 2017 | Talking about the world with a distributed model
Gemma Boleda |
INLG | 1 |
| 2016 | The LAMBADA dataset: Word prediction requiring a broad discourse contextabstractDenis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández |
ACL (1) | 8 |
| 2016 | Convolutional Neural Network Language ModelsabstractConvolutional Neural Networks (CNNs) have shown to yield very strong results in several Computer Vision tasks. Their application to language has received much less attention, and it has mainly focused on static classification tasks, such as sentence classification for Sentiment Analysis or relation extraction. In this work, we study the application of CNNs to language modeling, a dynamic, sequential prediction task that needs models to capture local as well as long-range dependency information. Our contribution is twofold. First, we show that CNNs achieve 11-26% better absolute performance than feed-forward neural\nlanguage models, demonstrating their potential for language representation even in sequential tasks. As for recurrent models, our model outperforms RNNs but is below state of the art LSTM models. Second, we gain some understanding of the behavior of the model, showing that CNNs in language act as feature detectors at a high level of abstraction, like in Computer Vision, and that the model can profitably use information from as far as 16 words before the target. Ngoc-Quan Pham, Germán Kruszewski, Gemma Boleda |
EMNLP | 3 |
| 2016 | Formal Distributional Semantics: Introduction to the Special IssueabstractFormal Semantics and Distributional Semantics are two very influential semantic frameworks in Computational Linguistics. Formal Semantics is based on a symbolic tradition and centered around the inferential properties of language. Distributional Semantics is statistical and data-driven, and focuses on aspects of meaning related to descriptive content. The two frameworks are complementary in their strengths, and this has motivated interest in combining them into an overarching semantic framework: a “Formal Distributional Semantics.” Given the fundamentally different natures of the two paradigms, however, building an integrative framework poses significant theoretical and engineering challenges. The present issue of Computational Linguistics advances the state of the art in Formal Distributional Semantics; this introductory article explains the motivation behind it and summarizes the contributions of previous work on the topic, providing the necessary background for the articles that follow. Gemma Boleda, Aurélie Herbelot |
Comput. Linguistics | 1 |
| 2015 | Distributional vectors encode referential attributesabstractDistributional methods have proven to excel at capturing fuzzy, graded aspects of meaning (Italy is more similar to Spain than to Germany).In contrast, it is difficult to extract the values of more specific attributes of word referents from distributional representations, attributes of the kind typically found in structured knowledge bases (Italy has 60 million inhabitants).In this paper, we pursue the hypothesis that distributional vectors also implicitly encode referential attributes.We show that a standard supervised regression model is in fact sufficient to retrieve such attributes to a reasonable degree of accuracy: When evaluated on the prediction of both categorical and numeric attributes of countries and cities, the model consistently reduces baseline error by 30%, and is not far from the upper bound.Further analysis suggests that our model is able to "objectify" distributional representations for entities, anchoring them more firmly in the external world in measurable ways. Abhijeet Gupta, Gemma Boleda, Marco Baroni, Sebastian Padó |
EMNLP | 2 |
| 2014 | Inclusive yet Selective: Supervised Distributional Hypernymy Detection
Stephen Roller, Katrin Erk, Gemma Boleda |
COLING | 3 |
| 2012 | Distributional Semantics in Technicolor
Elia Bruni, Gemma Boleda, Marco Baroni, Nam-Khanh Tran |
ACL (1) | 2 |
| 2012 | First Order vs. Higher Order Modification in Distributional Semantics
Gemma Boleda, Eva Maria Vecchi, Miquel Cornudella, Louise McNally |
EMNLP-CoNLL | 1 |
| 2012 | Modeling Regular Polysemy: A Study on the Semantic Classification of Catalan AdjectivesabstractWe present a study on the automatic acquisition of semantic classes for Catalan adjectives from distributional and morphological information, with particular emphasis on polysemous adjectives. The aim is to distinguish and characterize broad classes, such as qualitative (gran ‘big’) and relational (pulmonar ‘pulmonary’) adjectives, as well as to identify polysemous adjectives such as econòmic (‘economic ∣ cheap’). We specifically aim at modeling regular polysemy, that is, types of sense alternations that are shared across lemmata. To date, both semantic classes for adjectives and regular polysemy have only been sparsely addressed in empirical computational linguistics. Two main specific questions are tackled in this article. First, what is an adequate broad semantic classification for adjectives? We provide empirical support for the qualitative and relational classes as defined in theoretical work, and uncover one type of adjective that has not received enough attention, namely, the event-related class. Second, how is regular polysemy best modeled in computational terms? We present two models, and argue that the second one, which models regular polysemy in terms of simultaneous membership to multiple basic classes, is both theoretically and empirically more adequate than the first one, which attempts to identify independent polysemous classes. Our best classifier achieves 69.1% accuracy, against a 51% baseline. Gemma Boleda, Sabine Schulte im Walde, Toni Badia |
Comput. Linguistics | 1 |
| 2010 | Language Technology Challenges of a 'Small' Language (Catalan)
Maite Melero, Gemma Boleda, Montse Cuadros, Cristina España-Bonet, Lluís Padró 0001, Martí Quixal, Carlos Rodríguez Penagos, Roser Saurí |
LREC | 2 |
| 2010 | ADN-Classifier: Automatically Assigning Denotation Types to Nominalizations
Aina Peris, Mariona Taulé, Gemma Boleda, Horacio Rodríguez |
LREC | 3 |
| 2010 | Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus
Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró 0001, German Rigau |
LREC | 2 |
| 2010 | Annotation and Representation of a Diachronic Corpus of Spanish
Cristina Sánchez Marco, Gemma Boleda, Josep Maria Fontana, Judith Domingo |
LREC | 2 |
| 2010 | The Database of Catalan Adjectives
Roser Sanromà, Gemma Boleda |
LREC | 2 |
| 2008 | Evaluation of a Machine Translation System for Low Resource Languages: METIS-II
Vincent Vandeghinste, Peter Dirix, Ineke Schuurman, Stella Markantonatou, Sokratis Sofianopoulos, Marina Vassiliou, Olga Yannoutsou, Toni Badia, Maite Melero, Gemma Boleda, Michael Carl, Paul Schmidt |
LREC | 10 |
| 2007 | Modelling Polysemy in Adjective Classes by Multi-Label Classification
Gemma Boleda, Sabine Schulte im Walde, Toni Badia |
EMNLP-CoNLL | 1 |
| 2004 | Acquisition of Semantic Classes for Adjectives from Distributional Evidence
Gemma Boleda, Toni Badia, Eloi Batlle |
COLING | 1 |
| 2004 | The Influence of Argument Structure on Semantic Role Assignment
Sebastian Padó, Gemma Boleda |
EMNLP | 2 |
| 2003 | Clustering Adjectives for Class Discovery
Gemma Boleda, Laura Alonso Alemany |
EACL | 1 |
| 2002 | CATCG: a general purpose parsing tool applied
Alex Alsina, Toni Badia, Gemma Boleda, Stefan Bott, Angel Gil, Martí Quixal, Oriol Valentín |
LREC | 3 |