Paul Buitelaar

dblp:08/3075 · DBLP profile ↗
← Back
25ranked-venue papers in the field
4as first author
7since 2021 · last 2025
0000-0001-7238-9842ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 17 (3 first)Information Retrieval & Web Search · 5Database Systems & Data Management · 1Business Process & Enterprise Data · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 DiaSafety-CC: Annotating Dialogues with Safety Labels and Reasons for Cross-Cultural Analysis
abstract
A dialogue dataset developed in a language can have diverse safety annotations when presented to raters from different cultures. What is considered acceptable in one culture can be perceived as offensive in another culture. Cultural differences in dialogue safety annotation is yet to be fully explored. In this work, we use the geopolitical entity, Country, as our base for cultural study. We extend DiaSafety, an existing English dialogue safety dataset that was originally annotated by raters from Western culture, to create a new dataset, DiaSafety-CC. In our work, three raters each from Nigeria and India reannotate the DiaSafety dataset and provide reasons for their choice of labels. We perform pairwise comparisons of the annotations across the cultures studied. Furthermore, we compare the representative labels of each rater group to that of an existing large language model (LLM). Due to the subjectivity of the dialogue annotation task, 32.6% of the considered dialogues achieve unanimous annotation consensus across the labels of DiaSafety and the six raters. In our analyses, we observe that the Unauthorized Expertise and Biased Opinion categories have dialogues with the highest label disagreement ratio across the cultures studied. On manual inspection of the reasons provided for the choice of labels, we observe that raters across the cultures in DiaSafety-CC are sensitive to dialogues directed at target groups compared to dialogues directed at individuals. We also observe that GPT-4o annotation shows a more positive agreement with DiaSafety labels in terms of F1 score and phi coefficient.
Tunde Ajayi, Mihael Arcan, Paul Buitelaar
LDK3
2025 Towards Semantic Integration of Opinions: Unified Opinion Concepts Ontology and Extraction Task
abstract
This paper introduces the Unified Opinion Concepts (UOC) ontology to integrate opinions within their semantic context. The UOC ontology bridges the gap between the semantic representation of opinion across different formulations. It is a unified conceptualisation based on the facets of opinions studied extensively in NLP and semantic structures described through symbolic descriptions. We further propose the Unified Opinion Concept Extraction (UOCE) task of extracting opinions from the text with enhanced expressivity. Additionally, we provide a manually extended and re-annotated evaluation dataset for this task and tailored evaluation metrics to assess the adherence of extracted opinions to UOC semantics. Finally, we establish baseline performance for the UOCE task using state-of-the-art generative models.
Gaurav Negi, Dhairya Dalal, Omnia Zayed, Paul Buitelaar
LDK4
2025 Empowering Recommender Systems using Automatically Generated Knowledge Graphs and Reinforcement Learning
abstract
Personalized recommender systems play a crucial role in direct marketing, particularly in financial services, where delivering relevant content can enhance customer engagement and promote informed decision-making. This study explores interpretable knowledge graph (KG)-based recommender systems by proposing two distinct approaches for personalized article recommendations within a multinational financial services firm. The first approach leverages Reinforcement Learning (RL) to traverse a KG constructed from both structured (tabular) and unstructured (textual) data, enabling interpretability through Path Directed Reasoning (PDR). The second approach employs the XGBoost algorithm, with post-hoc explainability techniques such as SHAP and ELI5 to enhance transparency. By integrating machine learning with automatically generated KGs, our methods not only improve recommendation accuracy but also provide interpretable insights, facilitating more informed decision-making in customer relationship management.
Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, John P. McCrae, János A Perge, Shovon Sengupta, Paul Buitelaar
LDK8
2023 CURED4NLG: A Dataset for Table-to-Text Generation
Nivranshu Pasricha, Mihael Arcan, Paul Buitelaar
LDK3
2023 Multimodal Offensive Meme Classification with Natural Language Inference
Shardul Suryawanshi, Mihael Arcan, Suzanne Little, Paul Buitelaar
LDK4
2023 Identifying FrameNet Lexical Semantic Structures for Knowledge Graph Extraction from Financial Customer Interactions
abstract
We explore the use of the well established lexical resource and theory of the Berkeley FrameNet project to support the creation of a domain-specific knowledge graph in the financial domain, more precisely from financial customer interactions.We introduce a domain independent and unsupervised method that can be used across multiple applications, and test our experiments on the financial domain.We use an existing tool for term extraction and taxonomy generation in combination with information taken from FrameNet.By using principles from frame semantic theory, we show that we can connect domain-specific terms with their semantic concepts (semantic frames) and their properties (frame elements) to enrich knowledge about these terms, in order to improve the customer experience in customer-agent dialogue settings.
Cécile Robin, Atharva Kulkarni, Paul Buitelaar
GWC3
2021 Automatic Construction of Knowledge Graphs from Text and Structured Data: A Preliminary Literature Review
abstract
Knowledge graphs have been shown to be an important data structure for many applications, including chatbot development, data integration, and semantic search. In the enterprise domain, such graphs need to be constructed based on both structured (e.g. databases) and unstructured (e.g. textual) internal data sources; preferentially using automatic approaches due to the costs associated with manual construction of knowledge graphs. However, despite the growing body of research that leverages both structured and textual data sources in the context of automatic knowledge graph construction, the research community has centered on either one type of source or the other. In this paper, we conduct a preliminary literature review to investigate approaches that can be used for the integration of textual and structured data sources in the process of automatic knowledge graph construction. We highlight the solutions currently available for use within enterprises and point areas that would benefit from further research.
Maraim Masoud, Bianca Pereira, John P. McCrae, Paul Buitelaar
LDK4
2019 Utilizing Knowledge Graphs for Neural Machine Translation Augmentation
abstract
While neural networks have led to substantial progress in machine translation, their success depends heavily on large amounts of training data. However, parallel training corpora are not always readily available. Moreover, out-of-vocabulary words---mostly entities and terminological expressions---pose a difficult challenge to Neural Machine Translation systems. Recent efforts have tried to alleviate the data sparsity problem by augmenting the training data using different strategies, such as external knowledge injection. In this paper, we hypothesize that knowledge graphs enhance the semantic feature extraction of neural models, thus optimizing the translation of entities and terminological expressions in texts and consequently leading to better translation quality. We investigate two different strategies for incorporating knowledge graphs into neural models without modifying the neural network architectures. Additionally, we examine the effectiveness of our augmented models on domain-specific texts and ontologies. Our knowledge-graph-augmented neural translation model, dubbed KG-NMT, achieves significant and consistent improvements of +3 BLEU, METEOR and chrF3 on average on the newstest datasets between 2015 and 2018 for the WMT English-German translation task.
Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Paul Buitelaar, Mihael Arcan
K-CAP3
2019 Crowd-Sourcing A High-Quality Dataset for Metaphor Identification in Tweets
abstract
Metaphor is one of the most important elements of human communication, especially in informal settings such as social media. There have been a number of datasets created for metaphor identification, however, this task has proven difficult due to the nebulous nature of metaphoricity. In this paper, we present a crowd-sourcing approach for the creation of a dataset for metaphor identification, that is able to rapidly achieve large coverage over the different usages of metaphor in a given corpus while maintaining high accuracy. We validate this methodology by creating a set of 2,500 manually annotated tweets in English, for which we achieve inter-annotator agreement scores over 0.8, which is higher than other reported results that did not limit the task. This methodology is based on the use of an existing classifier for metaphor in order to assist in the identification and the selection of the examples for annotation, in a way that reduces the cognitive load for annotators and enables quick and accurate annotation. We selected a corpus of both general language tweets and political tweets relating to Brexit and we compare the resulting corpus on these two domains. As a result of this work, we have published the first dataset of tweets annotated for metaphors, which we believe will be invaluable for the development, training and evaluation of approaches for metaphor identification in tweets.
Omnia Zayed, John P. McCrae, Paul Buitelaar
LDK3
2017 An Evaluation Dataset for Linked Data Profiling
Andrejs Abele, John P. McCrae, Paul Buitelaar
LDK3
2016 ESSOT: An Expert Supporting System for Ontology Translation
Mihael Arcan, Mauro Dragoni, Paul Buitelaar
NLDB3
2016 Using Semantic Frames for Automatic Annotation of Regulatory Texts
Kartik Asooja, Georgeta Bordea, Paul Buitelaar
NLDB3
2016 Translating Ontologies in Real-World Settings
Mihael Arcan, Mauro Dragoni, Paul Buitelaar
ISWC (2)3
2016 Domain adaptation for ontology localization
John P. McCrae, Mihael Arcan, Kartik Asooja, Jorge Gracia, Paul Buitelaar, Philipp Cimiano
J. Web Semant.5
2015 Approximate and selective reasoning on knowledge graphs: A distributional semantics approach
André Freitas, João C. P. da Silva, Edward Curry, Paul Buitelaar
Data Knowl. Eng.4
2014 A Distributional Semantics Approach for Selective Reasoning on Commonsense Graph Knowledge Bases
André Freitas, João C. P. da Silva, Edward Curry, Paul Buitelaar
NLDB4
2013 Improving ESA with Document Similarity
Tamara Polajnar, Nitish Aggarwal, Kartik Asooja, Paul Buitelaar
ECIR4
2013 Cross-Lingual Natural Language Querying over the Web of Data
Nitish Aggarwal, Tamara Polajnar, Paul Buitelaar
NLDB3
2012 Challenges for the multilingual Web of Data
Jorge Gracia, Elena Montiel-Ponsoda, Philipp Cimiano, Asunción Gómez-Pérez, Paul Buitelaar, John P. McCrae
J. Web Semant.5
2011 LexInfo: A declarative model for the lexicon-ontology interface
Philipp Cimiano, Paul Buitelaar, John P. McCrae, Michael Sintek
J. Web Semant.2
2009 Towards Linguistically Grounded Ontologies
Paul Buitelaar, Philipp Cimiano, Peter Haase 0001, Michael Sintek
ESWC1
2009 Expertise mining from scientific literature
abstract
We describe an approach to pattern-based expertise topic extraction from publicly available scientific publications using Google Scholar. The approach is based on the observation that in the scientific text genre expertise topics will occur frequently in the context of particular phrasings that introduce them, such as 'method for', 'approach to', etc. The extracted knowledge can be used to analyze research structure in terms of expertise topics, researchers associated with these and relations between them.
Paul Buitelaar, Thomas Eigner
K-CAP1
2006 A Multilingual/Multimedia Lexicon Model for Ontologies
Paul Buitelaar, Michael Sintek, Malte Kiesel
ESWC1
2005 RelExt: A Tool for Relation Extraction from Text in Ontology Extension
Alexander Schutz, Paul Buitelaar
ISWC2
1992 The Use of a Lexicon to Interpret ER-Diagrams: A LIKE project
Paul Buitelaar, Reind P. van de Riet
ER1