John P. McCrae

dblp:80/7108 · also John Philip McCrae · DBLP profile ↗
← Back
35ranked-venue papers in the field
9as first author
17since 2021 · last 2025
0000-0002-7227-1331ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 19 (4 first)Other / Interdisciplinary · 12 (5 first)Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 MOOC on Linguistic Linked Data
Jorge Gracia, Slavko Zitnik, Maxim Ionov, Christian Chiarcos, Dagmar Gromann, Francesco Mambrini, Marco Passarotti, Armando Stellato, John P. McCrae, Gilles Sérasset, Andon Tchechmedjiev, Sara Carvalho, Penny Labropoulou, Rute Costa
ESWC (2)9
2025 Benchmarking Hindi Term Extraction in Education: A Dataset and Analysis
abstract
This paper introduces the HTEC HindiTerm Extraction Dataset 2.0, a resourcedesigned to support terminology extractionand classification tasks within the education domain. HTEC 2.0 has been developed with the objective of providing a high-quality benchmark dataset for the evaluation of term recognition and classification methodologies in Hindi educationaldiscourse. The dataset consists of 97 documents sourced from Hindi Wikipedia, covering a diverse range of topics relevant tothe education sector. Within these documents, 1,702 terms have been manuallyannotated where each term is defined as asingle-word or multi-word expression thatconveys a domain-specific meaning. Theannotated terms in HTEC 2.0 are systematically categorized into seven distinct classes.Furthermore, this paper outlines the development of annotation guidelines, detailingthe criteria used to determine term boundaries and category assignments. By offeringa structured dataset with clearly definedterm classifications, HTEC 2.0 serves as avaluable resource for researchers workingon terminology extraction, domain-specificnamed entity recognition, and text classification in Hindi.
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK3
2025 Cuaċ: Fast and Small Universal Representations of Corpora
abstract
The increasing size and diversity of corpora in natural language processing requires highly efficient processing frameworks. Building on the universal corpus format, Teanga, we present Cuaċ, a format for the compact representation of corpora. We describe this methodology based on short-string compression and indexing techniques and show that the files created with this methodology are similar to compressed human-readable serializations and can be further compressed using lossless compression. We also show that this introduces no computational penalty on the time to process files. This methodology aims to speed up natural language processing pipelines and is the basis for a fast database system for corpora.
John P. McCrae, Bernardo Stearns, Alamgir Munir Qazi, Shubhanker Banerjee, Atul Kr. Ojha
LDK1
2025 When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection
abstract
The proliferation of misinformation necessitates robust yet computationally efficient fact verification systems. While current state-of-the-art approaches leverage Large Language Models (LLMs) for generating explanatory rationales, these methods face significant computational barriers and hallucination risks in real-world deployments. We present DeReC (Dense Retrieval Classification), a lightweight framework that demonstrates how general-purpose text embeddings can effectively replace autoregressive LLM-based approaches in fact verification tasks. By combining dense retrieval with specialized classification, our system achieves better accuracy while being significantly more efficient. DeReC outperforms explanation-generating LLMs in efficiency, reducing runtime by 95% on RAWFC (23 minutes 36 seconds compared to 454 minutes 12 seconds) and by 92% on LIAR-RAW (134 minutes 14 seconds compared to 1692 minutes 23 seconds), showcasing its effectiveness across varying dataset sizes. On the RAWFC dataset, DeReC achieves an F1 score of 65.58%, surpassing the state-of-the-art method L-Defense (61.20%). Our results demonstrate that carefully engineered retrieval-based systems can match or exceed LLM performance in specialized tasks while being significantly more practical for real-world deployment.
Alamgir Munir Qazi, John P. McCrae, Jamal Abdul Nasir
LDK2
2025 Empowering Recommender Systems using Automatically Generated Knowledge Graphs and Reinforcement Learning
abstract
Personalized recommender systems play a crucial role in direct marketing, particularly in financial services, where delivering relevant content can enhance customer engagement and promote informed decision-making. This study explores interpretable knowledge graph (KG)-based recommender systems by proposing two distinct approaches for personalized article recommendations within a multinational financial services firm. The first approach leverages Reinforcement Learning (RL) to traverse a KG constructed from both structured (tabular) and unstructured (textual) data, enabling interpretability through Path Directed Reasoning (PDR). The second approach employs the XGBoost algorithm, with post-hoc explainability techniques such as SHAP and ELI5 to enhance transparency. By integrating machine learning with automatically generated KGs, our methods not only improve recommendation accuracy but also provide interpretable insights, facilitating more informed decision-making in customer relationship management.
Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, John P. McCrae, János A Perge, Shovon Sengupta, Paul Buitelaar
LDK5
2024 Large Language Models for Few-Shot Automatic Term Extraction
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
NLDB (1)3
2023 MG2P: An Empirical Study Of Multilingual Training for Manx G2P
Shubhanker Banerjee, Bharathi Raja Chakravarthi, John P. McCrae
LDK3
2023 The Cardamom Workbench for Historical and Under-Resourced Languages
Adrian Doyle, Theodorus Fransen, Bernardo Stearns, John P. McCrae, Oksana Dereza, Priya Rani
LDK4
2023 PICKD: In-Situ Prompt Tuning for Knowledge-Grounded Dialogue Generation
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae
PAKDD (4)4
2023 Documenting the Open Multilingual Wordnet
abstract
In this project note we describe our work to make better documentation for the Open Multilingual Wordnet (OMW), a platform integrating many open wordnets.This includes the documentation of the OMW website itself as well as of semantic relations used by the component wordnets.Some of this documentation work was done with the support of the Google Season of Docs.The OMW project page, which links both to the actual OMW server and the documentation has been moved to a new location: https://omwn.org.
Francis Bond, Michael Wayne Goodman, Ewa Rudnicka, Luís Morgado da Costa, Alexandre Rademaker, John P. McCrae
GWC6
2023 Some Considerations in the Construction of a Historical Language WordNet
abstract
This article describes the manual construction of a part of the Old English WordNet (Old-EWN) covering the semantic field of emotion terms.This manually constructed part of the wordnet is to be eventually integrated with the automatically generated/manually checked part covering the whole of the rest of the Old English lexicon (currently under construction).We present the workflow for the definition of these emotion synsets on the basis of a dataset produced by a specialist in this area.We also look at the enrichment of the original Global WordNet Association Lexical Markup Framework (GWA LMF) schema to include the extra information which this part of the OldEWN requires.In the final part of the article we discuss how the wordnet style of lexicon organisation can be used to share and disseminate research findings/datasets in lexical semantics.
Anas Fahad Khan, John P. McCrae, Francisco J. Minaya Gómez, Rafael Cruz González, Javier E. Díaz-Vera
GWC2
2022 Semantic Aware Answer Sentence Selection Using Self-Learning Based Domain Adaptation
abstract
Selecting an appropriate and relevant context forms an essential component for the efficacy of several information retrieval applications like Question Answering (QA) systems. The problem of Answer Sentence Selection (AS2) refers to the task of selecting sentences, from a larger text, that are relevant and contain the answer to users' queries. While there has been a lot of success in building AS2 systems trained on open-domain data (e.g., SQuAD, NQ), they do not generalize well in closed-domain settings, since domain adaptation can be challenging due to poor availability and annotation expense of domain-specific data. This paper proposes SEDAN, an effective self-learning framework to adapt AS2 models for domain-specific applications. We leverage large pre-trained language models to automatically generate domain-specific QA pairs for domain adaptation. We further fine-tune a pre-trained Sentence-BERT architecture to capture semantic relatedness between questions and answer sentences for AS2. Extensive experiments demonstrate the effectiveness of our proposed approach (over existing state-of-the-art AS2 baselines) on different Question Answering benchmark datasets.
Rajdeep Sarkar, Sourav Dutta 0001, Haytham Assem, Mihael Arcan, John P. McCrae
KDD5
2021 Encoder-Attention-Based Automatic Term Recognition (EA-ATR)
abstract
Automated Term Recognition (ATR) is the task of finding terminology from raw text. It involves designing and developing techniques for the mining of possible terms from the text and filtering these identified terms based on their scores calculated using scoring methodologies like frequency of occurrence and then ranking the terms. Current approaches often rely on statistics and regular expressions over part-of-speech tags to identify terms, but this is error-prone. We propose a deep learning technique to improve the process of identifying a possible sequence of terms. We improve the term recognition by using Bidirectional Encoder Representations from Transformers (BERT) based embeddings to identify which sequence of words is a term. This model is trained on Wikipedia titles. We assume all Wikipedia titles to be the positive set, and random n-grams generated from the raw text as a weak negative set. The positive and negative set will be trained using the Embed, Encode, Attend and Predict (EEAP) formulation using BERT as embeddings. The model will then be evaluated against different domain-specific corpora like GENIA - annotated biological terms and Krapivin - scientific papers from the computer science domain.
Sampritha H. Manjunath, John P. McCrae
LDK2
2021 Automatic Construction of Knowledge Graphs from Text and Structured Data: A Preliminary Literature Review
abstract
Knowledge graphs have been shown to be an important data structure for many applications, including chatbot development, data integration, and semantic search. In the enterprise domain, such graphs need to be constructed based on both structured (e.g. databases) and unstructured (e.g. textual) internal data sources; preferentially using automatic approaches due to the costs associated with manual construction of knowledge graphs. However, despite the growing body of research that leverages both structured and textual data sources in the context of automatic knowledge graph construction, the research community has centered on either one type of source or the other. In this paper, we conduct a preliminary literature review to investigate approaches that can be used for the integration of textual and structured data sources in the process of automatic knowledge graph construction. We highlight the solutions currently available for use within enterprises and point areas that would benefit from further research.
Maraim Masoud, Bianca Pereira, John P. McCrae, Paul Buitelaar
LDK3
2021 Monolingual Word Sense Alignment as a Classification Problem
abstract
Words are defined based on their meanings in various ways in different resources.Aligning word senses across monolingual lexicographic resources increases domain coverage and enables integration and incorporation of data.In this paper, we explore the application of classification methods using manually-extracted features along with representation learning techniques in the task of word sense alignment and semantic relationship detection.We demonstrate that the performance of classification methods dramatically varies based on the type of semantic relationships due to the nature of the task but outperforms the previous experiments.
Sina Ahmadi, John P. McCrae
GWC2
2021 Towards a Linking between WordNet and Wikidata
abstract
WordNet is the most widely used lexical resource for English, while Wikidata is one of the largest knowledge graphs of entity and concepts available.While, there is a clear difference in the focus of these two resources, there is also a significant overlap and as such a complete linking of these resources would have many uses.We propose the development of such a linking, first by means of the hapax legomenon links and secondly by the use of natural language processing techniques.We show that these can be done with high accuracy but that human validation is still necessary.This has resulted in over 9,000 links being added between these two resources.
John P. McCrae, David Cillessen
GWC1
2021 The GlobalWordNet Formats: Updates for 2020
abstract
The Global Wordnet Formats have been introduced to enable wordnets to have a common representation that can be integrated through the Global WordNet Grid.As a result of their adoption, a number of shortcomings of the format were identified, and in this paper we describe the extensions to the formats that address these issues.These include: ordering of senses, dependencies between wordnets, pronunciation, syntactic modelling, relations, sense keys, metadata and RDF support.Furthermore, we provide some perspectives on how these changes help in the integration of wordnets.
John P. McCrae, Michael Wayne Goodman, Francis Bond, Alexandre Rademaker, Ewa Rudnicka, Luís Morgado da Costa
GWC1
2020 Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network
abstract
Exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices, but also enables people to express anti-social behavior like online harassment, cyberbul-lying, and hate speech. Numerous works have been proposed to utilize these data for social and anti-social behavior analysis, document characterization, and sentiment analysis by predicting the contexts mostly for highly resourced languages like English. However, some languages are under-resources, e.g., South Asian languages like Bengali, Tamil, Assamese, Malayalam, that lack of computational resources for natural language processing. In this paper1, we provide several classification benchmarks for Bengali, an under-resourced language. We prepared three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis. We built the largest Bengali word embedding models to date based on 250 million articles, which we call BengFastText. We perform three experiments, covering document classification, sentiment analysis, and hate speech detection. We incorporate word embeddings into a Multichannel Convolutional-LSTM (MC-LSTM) network for predicting different types of hate speech, document classification, and sentiment analysis. Experiments demonstrate that BengFastText can capture the semantics of words from respective contexts correctly. Evaluations against several baseline embedding models, e.g., Word2Vec and GloVe yield up to 92.30%, 82.25%, and 90.45% F1-scores in case of document classification, sentiment analysis, and hate speech detection, respectively during 5-fold cross-validation tests.
Md. Rezaul Karim 0001, Bharathi Raja Chakravarthi, John P. McCrae, Michael Cochez
DSAA3
2019 Comparison of Different Orthographies for Machine Translation of Under-Resourced Dravidian Languages
abstract
Under-resourced languages are a significant challenge for statistical approaches to machine translation, and recently it has been shown that the usage of training data from closely-related languages can improve machine translation quality of these languages. While languages within the same language family share many properties, many under-resourced languages are written in their own native script, which makes taking advantage of these language similarities difficult. In this paper, we propose to alleviate the problem of different scripts by transcribing the native script into common representation i.e. the Latin script or the International Phonetic Alphabet (IPA). In particular, we compare the difference between coarse-grained transliteration to the Latin script and fine-grained IPA transliteration. We performed experiments on the language pairs English-Tamil, English-Telugu, and English-Kannada translation task. Our results show improvements in terms of the BLEU, METEOR and chrF scores from transliteration and we find that the transliteration into the Latin script outperforms the fine-grained IPA transcription.
Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae
LDK3
2019 Crowd-Sourcing A High-Quality Dataset for Metaphor Identification in Tweets
abstract
Metaphor is one of the most important elements of human communication, especially in informal settings such as social media. There have been a number of datasets created for metaphor identification, however, this task has proven difficult due to the nebulous nature of metaphoricity. In this paper, we present a crowd-sourcing approach for the creation of a dataset for metaphor identification, that is able to rapidly achieve large coverage over the different usages of metaphor in a given corpus while maintaining high accuracy. We validate this methodology by creating a set of 2,500 manually annotated tweets in English, for which we achieve inter-annotator agreement scores over 0.8, which is higher than other reported results that did not limit the task. This methodology is based on the use of an existing classifier for metaphor in order to assist in the identification and the selection of the examples for annotation, in a way that reduces the cognitive load for annotators and enables quick and accurate annotation. We selected a corpus of both general language tweets and political tweets relating to Brexit and we compare the resulting corpus on these two domains. As a result of this work, we have published the first dataset of tweets annotated for metaphors, which we believe will be invaluable for the development, training and evaluation of approaches for metaphor identification in tweets.
Omnia Zayed, John P. McCrae, Paul Buitelaar
LDK2
2019 English WordNet 2019 - An Open-Source WordNet for English
abstract
We describe the release of a new wordnet for English based on the Princeton WordNet, but now developed under an open-source model.In particular, this version of WordNet, which we call English WordNet 2019, which has been developed by multiple people around the world through GitHub, fixes many errors in previous wordnets for English.We give some details of the changes that have been made in this version and give some perspectives about likely future changes that will be made as this project continues to evolve.
John P. McCrae, Alexandre Rademaker, Francis Bond, Ewa Rudnicka, Christiane Fellbaum
GWC1
2018 Improving Wordnets for Under-Resourced Languages Using Machine Translation
abstract
Wordnets are extensively used in natural language processing, but the current approaches for manually building a wordnet from scratch involves large research groups for a long period of time, which are typically not available for under-resourced languages.Even if wordnet-like resources are available for under-resourced languages, they are often not easily accessible, which can alter the results of applications using these resources.Our proposed method presents an expand approach for improving and generating wordnets with the help of machine translation.We apply our methods to improve and extend wordnets for the Dravidian languages, i.e., Tamil, Telugu, Kannada, which are severly under-resourced languages.We report evaluation results of the generated wordnet senses in term of precision for these languages.In addition to that, we carried out a manual evaluation of the translations for the Tamil language, where we demonstrate that our approach can aid in improving wordnet resources for under-resourced Dravidian languages.
Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae
GWC3
2018 Mapping WordNet Instances to Wikipedia
abstract
Lexical resource differ from encyclopaedic resources and represent two distinct types of resource covering general language and named entities respectively.However, many lexical resources, including Princeton WordNet, contain many proper nouns, referring to named entities in the world yet it is not possible or desirable for a lexical resource to cover all named entities that may reasonably occur in a text.In this paper, we propose that instead of including synsets for instance concepts PWN should instead provide links to Wikipedia articles describing the concept.In order to enable this we have created a gold-quality mapping between all of the 7,742 instances in PWN and Wikipedia (where such a mapping is possible).As such, this resource aims to provide a gold standard for link discovery, while also allowing PWN to distinguish itself from other resources such as DBpedia or BabelNet.Moreover, this linking connects PWN to the Linguistic Linked Open Data cloud, thus creating a richer, more usable resource for natural language processing.
John P. McCrae
GWC1
2018 Towards a Crowd-Sourced WordNet for Colloquial English
abstract
Princeton WordNet is one of the most widely-used resources for natural language processing, but is updated only infrequently and cannot keep up with the fast-changing usage of the English language on social media platforms such as Twitter.The Colloquial WordNet aims to provide an open platform whereby anyone can contribute, while still following the structure of WordNet.Many crowdsourced lexical resources often have significant quality issues, and as such care must be taken in the design of the interface to ensure quality.In this paper, we present the development of a platform that can be opened on the Web to any lexicographer who wishes to contribute to this resource and the lexicographic methodology applied by this interface.
John P. McCrae, Ian D. Wood, Amanda Hicks
GWC1
2018 ELEXIS - a European infrastructure fostering cooperation and information exchange among lexicographical research communities
abstract
The paper describes objectives, concept and methodology for ELEXIS, a European infrastructure fostering cooperation and information exchange among lexicographical research communities.The infrastructure is a newly granted project under the Horizon 2020 INFRAIA call, with the topic Integrating Activities for Starting Communities.The project is planned to start in January 2018.
Bolette S. Pedersen, John P. McCrae, Carole Tiberius, Simon Krek
GWC2
2017 An Evaluation Dataset for Linked Data Profiling
Andrejs Abele, John P. McCrae, Paul Buitelaar
LDK2
2017 OnLiT: An Ontology for Linguistic Terminology
abstract
Understanding the differences underlying the scope, usage and content of language data requires the provision of a clarifying terminological basis which is integrated in the metadata describing a particular language resource. While terminological resources such as the SIL Glossary of Linguistic Terms, ISOcat or the GOLD ontology provide a considerable amount of linguistic terms, their practical usage is limited to a look up of a defined term whose relation to other terms is unspecified or insufficient. Therefore, in this paper we propose an ontology for linguistic terminology, called OnLiT. It is a data model which can be used to represent linguistic terms and concepts in a semantically interrelated data structure and, thus, overcomes prevalent isolating definition-based term descriptions. OnLiT is based on the LiDo Glossary of Linguistic Terms and enables the creation of RDF datasets, that represent linguistic terms and their meanings within the whole or a subdomain of linguistics.
Bettina Klimek, John P. McCrae, Christian Lehmann, Christian Chiarcos, Sebastian Hellmann 0001
LDK2
2017 The Colloquial WordNet: Extending Princeton WordNet with Neologisms
John P. McCrae, Ian D. Wood, Amanda Hicks
LDK1
2016 CILI: the Collaborative Interlingual Index
abstract
This paper introduces the motivation for and design of the Collaborative InterLingual Index (CILI).It is designed to make possible coordination between multiple loosely coupled wordnet projects.The structure of the CILI is based on the Interlingual index first proposed in the Eu-roWordNet project with several pragmatic extensions: an explicit open license, definitions in English and links to wordnets in the Global Wordnet Grid.
Francis Bond, Piek Vossen, John P. McCrae, Christiane Fellbaum
GWC3
2016 Toward a truly multilingual GlobalWordnet Grid
abstract
In this paper, we describe a new and improved Global Wordnet Grid that takes advantage of the Collaborative InterLingual Index (CILI).Currently, the Open Multilingal Wordnet has made many wordnets accessible as a single linked wordnet, but as it used the Princeton Wordnet of English (PWN) as a pivot, it loses concepts that are not part of PWN.The technical solution to this, a central registry of concepts, as proposed in the EuroWordnet project through the InterLingual Index, has been known for many years.However, the practical issues of how to host this index and who decides what goes in remained unsolved.Inspired by current practice in the Semantic Web and the Linked Open Data community, we propose a way to solve this issue.In this paper we define the principles and protocols for contributing to the Grid.We tested them on two use cases, adding version 3.1 of the Princeton WordNet to a CILI based on 3.0 and adding the Open Dutch Wordnet, to validate the current set up.This paper aims to be a call for action that we hope will be further discussed and ultimately taken up by the whole wordnet community.
Piek Vossen, Francis Bond, John P. McCrae
GWC3
2016 Domain adaptation for ontology localization
John P. McCrae, Mihael Arcan, Kartik Asooja, Jorge Gracia, Paul Buitelaar, Philipp Cimiano
J. Web Semant.1
2015 LIME: The Metadata Module for OntoLex
Manuel Fiorelli, Armando Stellato, John P. McCrae, Philipp Cimiano, Maria Teresa Pazienza
ESWC3
2012 Challenges for the multilingual Web of Data
Jorge Gracia, Elena Montiel-Ponsoda, Philipp Cimiano, Asunción Gómez-Pérez, Paul Buitelaar, John P. McCrae
J. Web Semant.6
2011 Linking Lexical Resources and Ontologies on the Semantic Web with Lemon
John P. McCrae, Dennis Spohr, Philipp Cimiano
ESWC (1)1
2011 LexInfo: A declarative model for the lexicon-ontology interface
Philipp Cimiano, Paul Buitelaar, John P. McCrae, Michael Sintek
J. Web Semant.3