VLDB 2026 Research / reviewers in the wild / expert
Barbara McGillivray
dblp:10/8162
· DBLP profile ↗
14ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-3426-8200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Linguistic Knowledge Graphs for Sense Prediction: A Case-study on Latin
Eleonora Ghizzota, Paola Marongiu, Pierpaolo Basile, Stefano Ferilli, Barbara McGillivray |
LREC | 5 |
| 2025 | From Detection to Explanation: Effective Learning Strategies for LLMs in Online Abusive Language ResearchabstractAbusive language detection relies on understanding different levels of intensity, expressiveness and targeted groups, which requires commonsense reasoning, world knowledge and linguistic nuances that evolve over time. Here, we frame the problem as a knowledge-guided learning task, and demonstrate that LLMs’ implicit knowledge without an accurate strategy is not suitable for multi-class detection nor explanation generation. We publicly release GLlama Alarm, the knowledge-Guided version of Llama-2 instruction fine-tuned for multi-class abusive language detection and explanation generation. By being fine-tuned on structured explanations and external reliable knowledge sources, our model mitigates bias and generates explanations that are relevant to the text and coherent with human reasoning, with an average 48.76% better alignment with human judgment according to our expert survey. Chiara Di Bonaventura, Lucia Siciliani, Pierpaolo Basile, Albert Meroño-Peñuela, Barbara McGillivray |
COLING | 5 |
| 2024 | Language Pivoting from Parallel Corpora for Word Sense Disambiguation of Historical Languages: A Case Study on LatinabstractWord Sense Disambiguation (WSD) is an important task in NLP, which serves the purpose of automatically disambiguating a polysemous word with its most likely sense in context. Recent studies have advanced the state of the art in this task, but most of the work has been carried out on contemporary English or other modern languages, leaving challenges posed by low-resource languages and diachronic change open. Although the problem with low-resource languages has recently been mitigated by using existing multilingual resources to propagate otherwise expensive annotations from English to other languages, such techniques have hitherto not been applied to historical languages such as Latin. In this work, we make the following two major contributions. First, we test such a strategy on a historical language and propose a new approach in this framework which makes use of existing bilingual corpora instead of native English datasets. Second, we fine-tune a Latin WSD model on the data produced and achieve state-of-the-art results on a standard benchmark for the task. Finally, we release the dataset generated with our approach, which is the largest dataset for Latin WSD to date. This work opens the door to further research, as our approach can be used for different historical and, generally, under-resourced languages. Iacopo Ghinassi, Simone Tedeschi, Paola Marongiu, Roberto Navigli, Barbara McGillivray |
LREC/COLING | 5 |
| 2023 | Towards a Conversational Web? A Benchmark for Analysing Semantic Change with Conversational Knowledge Bots and Linked Open Data
Florentina Armaselu, Elena Apostol, Christian Chiarcos, Anas Fahad Khan, Chaya Liebeskind, Barbara McGillivray, Ciprian-Octavian Truica, Andrius Utka, Giedre Valunaite Oleskeviciene |
LDK | 6 |
| 2023 | Workflow Reversal and Data Wrangling in Multilingual Diachronic Analysis and Linguistic Linked Open Data Modelling
Florentina Armaselu, Barbara McGillivray, Chaya Liebeskind, Giedre Valunaite Oleskeviciene, Andrius Utka, Daniela Gîfu, Anas Fahad Khan, Elena Apostol, Ciprian-Octavian Truica |
LDK | 2 |
| 2023 | Graph Databases for Diachronic Language Data Modelling
Barbara McGillivray, Pierluigi Cassotti, Davide Di Pierro 0001, Paola Marongiu, Anas Fahad Khan, Stefano Ferilli, Pierpaolo Basile |
LDK | 1 |
| 2021 | DWUG: A large Resource of Diachronic Word Usage Graphs in Four LanguagesabstractWord meaning is notoriously difficult to capture, both synchronically and diachronically.In this paper, we describe the creation of the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments.We describe in detail the multi-round incremental annotation process, the choice for a clustering algorithm to group usages into senses, and possible -diachronic and synchronic -uses for this dataset. Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, Barbara McGillivray |
EMNLP (1) | 5 |
| 2021 | HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use CaseabstractThe paper proposes an interdisciplinary approach including methods from disciplines such as history of concepts, linguistics, natural language processing (NLP) and Semantic Web, to create a comparative framework for detecting semantic change in multilingual historical corpora and generating diachronic ontologies as linguistic linked open data (LLOD). Initiated as a use case (UC4.2.1) within the COST Action Nexus Linguarum, European network for Web-centred linguistic data science, the study will explore emerging trends in knowledge extraction, analysis and representation from linguistic data science, and apply the devised methodology to datasets in the humanities to trace the evolution of concepts from the domain of socio-cultural transformation. The paper will describe the main elements of the methodological framework and preliminary planning of the intended workflow. Florentina Armaselu, Elena Apostol, Anas Fahad Khan, Chaya Liebeskind, Barbara McGillivray, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene |
LDK | 5 |
| 2020 | Living Machines: A study of atypical animacyabstractMariona Coll Ardanuy, Federico Nanni, Kaspar Beelen, Kasra Hosseini, Ruth Ahnert, Jon Lawrence, Katherine McDonough, Giorgia Tolfo, Daniel CS Wilson, Barbara McGillivray. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Mariona Coll Ardanuy, Federico Nanni, Kaspar Beelen, Kasra Hosseini, Ruth Ahnert, Jon Lawrence, Katherine McDonough, Giorgia Tolfo, Daniel C. S. Wilson, Barbara McGillivray |
COLING | 10 |
| 2020 | Assessing the Impact of OCR Quality on Downstream NLP TasksabstractA growing volume of heritage data is being digitized and made available as text via optical character recognition (OCR). Scholars and libraries are increasingly using OCR-generated text for retrieval and analysis. However, the process of creating text through OCR introduces varying degrees of error to the text. The impact of these errors on natural language processing (NLP) tasks has only been partially studied. We perform a series of extrinsic assessment tasks — sentence segmentation, named entity recognition, dependency parsing, information retrieval, topic modelling and neural language model fine-tuning — using popular, out-of-the-box tools in order to quantify the impact of OCR quality on these tasks. We find a consistent impact resulting from OCR errors on our downstream tasks with some tasks more irredeemably harmed by OCR errors. Based on these results, we offer some preliminary guidelines for working with text produced through OCR. Daniel van Strien, Kaspar Beelen, Mariona Coll Ardanuy, Kasra Hosseini, Barbara McGillivray, Giovanni Colavizza |
ICAART (1) | 5 |
| 2020 | Urban Dictionary Embeddings for Slang NLP ApplicationsabstractThe choice of the corpus on which word embeddings are trained can have a sizable effect on the learned representations, the types of analyses that can be performed with them, and their utility as features for machine learning models. To contribute to the existing sets of pre-trained word embeddings, we introduce and release the first set of word embeddings trained on the content of Urban Dictionary, a crowd-sourced dictionary for slang words and phrases. We show that although these embeddings are trained on fewer total tokens (by at least an order of magnitude compared to most popular pre-trained embeddings), they have high performance across a range of common word embedding evaluations, ranging from semantic similarity to word clustering tasks. Further, for some extrinsic tasks such as sentiment analysis and sarcasm detection where we expect to require some knowledge of colloquial language on social media data, initializing classifiers with the Urban Dictionary Embeddings resulted in improved performance compared to initializing with a range of other well-known, pre-trained embeddings that are order of magnitude larger in size. Steven R. Wilson 0001, Walid Magdy, Barbara McGillivray, Venkata Rama Kiran Garimella, Gareth Tyson |
LREC | 3 |
| 2019 | Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word EmbeddingsabstractPhilippa Shoemark, Farhana Ferdousi Liza, Dong Nguyen, Scott Hale, Barbara McGillivray. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Philippa Shoemark, Farhana Ferdousi Liza, Dong Nguyen 0002, Scott A. Hale, Barbara McGillivray |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Exploiting the Web for Semantic Change Detection
Pierpaolo Basile, Barbara McGillivray |
DS | 2 |
| 2008 | Unsupervised Acquisition of Verb Subcategorization Frames from Shallow-Parsed Corpora
Alessandro Lenci, Barbara McGillivray, Simonetta Montemagni, Vito Pirrelli |
LREC | 2 |