EDBT 2026 Demo / reviewers in the wild / expert
Mihael Arcan
dblp:24/10894
· DBLP profile ↗
11ranked-venue papers in the field
2as first author
5since 2021 · last 2025
0000-0002-3116-621XORCID · reported
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 7 (1 first)Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiaSafety-CC: Annotating Dialogues with Safety Labels and Reasons for Cross-Cultural AnalysisabstractA dialogue dataset developed in a language can have diverse safety annotations when presented to raters from different cultures. What is considered acceptable in one culture can be perceived as offensive in another culture. Cultural differences in dialogue safety annotation is yet to be fully explored. In this work, we use the geopolitical entity, Country, as our base for cultural study. We extend DiaSafety, an existing English dialogue safety dataset that was originally annotated by raters from Western culture, to create a new dataset, DiaSafety-CC. In our work, three raters each from Nigeria and India reannotate the DiaSafety dataset and provide reasons for their choice of labels. We perform pairwise comparisons of the annotations across the cultures studied. Furthermore, we compare the representative labels of each rater group to that of an existing large language model (LLM). Due to the subjectivity of the dialogue annotation task, 32.6% of the considered dialogues achieve unanimous annotation consensus across the labels of DiaSafety and the six raters. In our analyses, we observe that the Unauthorized Expertise and Biased Opinion categories have dialogues with the highest label disagreement ratio across the cultures studied. On manual inspection of the reasons provided for the choice of labels, we observe that raters across the cultures in DiaSafety-CC are sensitive to dialogues directed at target groups compared to dialogues directed at individuals. We also observe that GPT-4o annotation shows a more positive agreement with DiaSafety labels in terms of F1 score and phi coefficient. Tunde Ajayi, Mihael Arcan, Paul Buitelaar |
LDK | 2 |
| 2023 | CURED4NLG: A Dataset for Table-to-Text Generation
Nivranshu Pasricha, Mihael Arcan, Paul Buitelaar |
LDK | 2 |
| 2023 | Multimodal Offensive Meme Classification with Natural Language Inference
Shardul Suryawanshi, Mihael Arcan, Suzanne Little, Paul Buitelaar |
LDK | 2 |
| 2023 | PICKD: In-Situ Prompt Tuning for Knowledge-Grounded Dialogue Generation
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae |
PAKDD (4) | 3 |
| 2022 | Semantic Aware Answer Sentence Selection Using Self-Learning Based Domain AdaptationabstractSelecting an appropriate and relevant context forms an essential component for the efficacy of several information retrieval applications like Question Answering (QA) systems. The problem of Answer Sentence Selection (AS2) refers to the task of selecting sentences, from a larger text, that are relevant and contain the answer to users' queries. While there has been a lot of success in building AS2 systems trained on open-domain data (e.g., SQuAD, NQ), they do not generalize well in closed-domain settings, since domain adaptation can be challenging due to poor availability and annotation expense of domain-specific data. This paper proposes SEDAN, an effective self-learning framework to adapt AS2 models for domain-specific applications. We leverage large pre-trained language models to automatically generate domain-specific QA pairs for domain adaptation. We further fine-tune a pre-trained Sentence-BERT architecture to capture semantic relatedness between questions and answer sentences for AS2. Extensive experiments demonstrate the effectiveness of our proposed approach (over existing state-of-the-art AS2 baselines) on different Question Answering benchmark datasets. Rajdeep Sarkar, Sourav Dutta 0001, Haytham Assem, Mihael Arcan, John P. McCrae |
KDD | 4 |
| 2019 | Utilizing Knowledge Graphs for Neural Machine Translation AugmentationabstractWhile neural networks have led to substantial progress in machine translation, their success depends heavily on large amounts of training data. However, parallel training corpora are not always readily available. Moreover, out-of-vocabulary words---mostly entities and terminological expressions---pose a difficult challenge to Neural Machine Translation systems. Recent efforts have tried to alleviate the data sparsity problem by augmenting the training data using different strategies, such as external knowledge injection. In this paper, we hypothesize that knowledge graphs enhance the semantic feature extraction of neural models, thus optimizing the translation of entities and terminological expressions in texts and consequently leading to better translation quality. We investigate two different strategies for incorporating knowledge graphs into neural models without modifying the neural network architectures. Additionally, we examine the effectiveness of our augmented models on domain-specific texts and ontologies. Our knowledge-graph-augmented neural translation model, dubbed KG-NMT, achieves significant and consistent improvements of +3 BLEU, METEOR and chrF3 on average on the newstest datasets between 2015 and 2018 for the WMT English-German translation task. Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Paul Buitelaar, Mihael Arcan |
K-CAP | 4 |
| 2019 | Comparison of Different Orthographies for Machine Translation of Under-Resourced Dravidian LanguagesabstractUnder-resourced languages are a significant challenge for statistical approaches to machine translation, and recently it has been shown that the usage of training data from closely-related languages can improve machine translation quality of these languages. While languages within the same language family share many properties, many under-resourced languages are written in their own native script, which makes taking advantage of these language similarities difficult. In this paper, we propose to alleviate the problem of different scripts by transcribing the native script into common representation i.e. the Latin script or the International Phonetic Alphabet (IPA). In particular, we compare the difference between coarse-grained transliteration to the Latin script and fine-grained IPA transliteration. We performed experiments on the language pairs English-Tamil, English-Telugu, and English-Kannada translation task. Our results show improvements in terms of the BLEU, METEOR and chrF scores from transliteration and we find that the transliteration into the Latin script outperforms the fine-grained IPA transcription. Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae |
LDK | 2 |
| 2018 | Improving Wordnets for Under-Resourced Languages Using Machine TranslationabstractWordnets are extensively used in natural language processing, but the current approaches for manually building a wordnet from scratch involves large research groups for a long period of time, which are typically not available for under-resourced languages.Even if wordnet-like resources are available for under-resourced languages, they are often not easily accessible, which can alter the results of applications using these resources.Our proposed method presents an expand approach for improving and generating wordnets with the help of machine translation.We apply our methods to improve and extend wordnets for the Dravidian languages, i.e., Tamil, Telugu, Kannada, which are severly under-resourced languages.We report evaluation results of the generated wordnet senses in term of precision for these languages.In addition to that, we carried out a manual evaluation of the translations for the Tamil language, where we demonstrate that our approach can aid in improving wordnet resources for under-resourced Dravidian languages. Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae |
GWC | 2 |
| 2016 | ESSOT: An Expert Supporting System for Ontology Translation
Mihael Arcan, Mauro Dragoni, Paul Buitelaar |
NLDB | 1 |
| 2016 | Translating Ontologies in Real-World Settings
Mihael Arcan, Mauro Dragoni, Paul Buitelaar |
ISWC (2) | 1 |
| 2016 | Domain adaptation for ontology localization
John P. McCrae, Mihael Arcan, Kartik Asooja, Jorge Gracia, Paul Buitelaar, Philipp Cimiano |
J. Web Semant. | 2 |