VLDB 2026 Research / reviewers in the wild / expert
Mihael Arcan
dblp:24/10894
· DBLP profile ↗
30ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0002-3116-621XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiaSafety-CC: Annotating Dialogues with Safety Labels and Reasons for Cross-Cultural AnalysisabstractA dialogue dataset developed in a language can have diverse safety annotations when presented to raters from different cultures. What is considered acceptable in one culture can be perceived as offensive in another culture. Cultural differences in dialogue safety annotation is yet to be fully explored. In this work, we use the geopolitical entity, Country, as our base for cultural study. We extend DiaSafety, an existing English dialogue safety dataset that was originally annotated by raters from Western culture, to create a new dataset, DiaSafety-CC. In our work, three raters each from Nigeria and India reannotate the DiaSafety dataset and provide reasons for their choice of labels. We perform pairwise comparisons of the annotations across the cultures studied. Furthermore, we compare the representative labels of each rater group to that of an existing large language model (LLM). Due to the subjectivity of the dialogue annotation task, 32.6% of the considered dialogues achieve unanimous annotation consensus across the labels of DiaSafety and the six raters. In our analyses, we observe that the Unauthorized Expertise and Biased Opinion categories have dialogues with the highest label disagreement ratio across the cultures studied. On manual inspection of the reasons provided for the choice of labels, we observe that raters across the cultures in DiaSafety-CC are sensitive to dialogues directed at target groups compared to dialogues directed at individuals. We also observe that GPT-4o annotation shows a more positive agreement with DiaSafety labels in terms of F1 score and phi coefficient. Tunde Ajayi, Mihael Arcan, Paul Buitelaar |
LDK | 2 |
| 2025 | Leveraging Visual Scene Graph to Enhance Translation Quality in Multimodal Machine TranslationabstractDespite significant advancements in Multimodal Machine Translation, understanding and effectively utilising visual scenes within multimodal models remains a complex challenge. Extracting comprehensive and relevant visual features requires extensive and detailed input data to ensure the model accurately captures objects, their attributes, and relationships within a scene. In this paper, we explore using visual scene graphs extracted from images to enhance the performance of translation models. We investigate this approach for integrating Visual Scene Graph information into translation models, focusing on representing this information in a semantic structure rather than relying on raw image data. The performance of our approach was evaluated on the Multi30K dataset for English into German, French, and Czech translations using BLEU, chrF2, TER and COMET metrics. Our results demonstrate that utilising visual scene graph information improves translation performance. Using information on semantic structure can improve the multimodal baseline model, leading to better contextual understanding and translation accuracy. Ali Hatami, Mihael Arcan, Paul Buitelaar |
MTSummit (1) | 2 |
| 2024 | Cross-lingual Transfer and Multilingual Learning for Detecting Harmful Behaviour in African Under-Resourced Language DialogueabstractMost harmful dialogue detection models are developed for high-resourced languages.Consequently, users who speak under-resourced languages cannot fully benefit from these models in terms of usage, development, detection and mitigation of harmful dialogue utterances.Our work aims at detecting harmful utterances in under-resourced African languages.We leverage transfer learning using pretrained models trained with multilingual embeddings to develop a cross-lingual model capable of detecting harmful content across various African languages.We first fine-tune a harmful dialogue detection model on a selected African dialogue dataset.Additionally, we fine-tune a model on a combined dataset in some African languages to develop a multilingual harmful dialogue detection model.We then evaluate the cross-lingual model's ability to generalise to an unseen African language by performing harmful dialogue detection in an under-resourced language not present during pretraining or finetuning.We evaluate our models on the test datasets.We show that our best performing models achieve impressive results in terms of F1 score.Finally, we discuss the results and limitations of our work. Tunde Ajayi, Mihael Arcan, Paul Buitelaar |
SIGDIAL | 2 |
| 2023 | CURED4NLG: A Dataset for Table-to-Text Generation
Nivranshu Pasricha, Mihael Arcan, Paul Buitelaar |
LDK | 2 |
| 2023 | Multimodal Offensive Meme Classification with Natural Language Inference
Shardul Suryawanshi, Mihael Arcan, Suzanne Little, Paul Buitelaar |
LDK | 2 |
| 2023 | A Filtering Approach to Object Region Detection in Multimodal Machine TranslationabstractRecent studies in Multimodal Machine Translation (MMT) have explored the use of visual information in a multimodal setting to analyze its redundancy with textual information. The aim of this work is to develop a more effective approach to incorporating relevant visual information into the translation process and improve the overall performance of MMT models. This paper proposes an object-level filtering approach in Multimodal Machine Translation, where the approach is applied to object regions extracted from an image to filter out irrelevant objects based on the image captions to be translated. Using the filtered image helps the model to consider only relevant objects and their relative locations to each other. Different matching methods, including string matching and word embeddings, are employed to identify relevant objects. Gaussian blurring is used to soften irrelevant objects from the image and to evaluate the effect of object filtering on translation quality. The performance of the filtering approaches was evaluated on the Multi30K dataset in English to German, French, and Czech translations, based on BLEU, ChrF2, and TER metrics. Ali Hatami, Paul Buitelaar, Mihael Arcan |
MTSummit (1) | 3 |
| 2023 | PICKD: In-Situ Prompt Tuning for Knowledge-Grounded Dialogue Generation
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae |
PAKDD (4) | 3 |
| 2023 | TrollsWithOpinion: A taxonomy and dataset for predicting domain-specific opinion manipulation in troll memes
Shardul Suryawanshi, Bharathi Raja Chakravarthi, Mihael Arcan, Paul Buitelaar |
Multim. Tools Appl. | 3 |
| 2022 | Semantic Aware Answer Sentence Selection Using Self-Learning Based Domain AdaptationabstractSelecting an appropriate and relevant context forms an essential component for the efficacy of several information retrieval applications like Question Answering (QA) systems. The problem of Answer Sentence Selection (AS2) refers to the task of selecting sentences, from a larger text, that are relevant and contain the answer to users' queries. While there has been a lot of success in building AS2 systems trained on open-domain data (e.g., SQuAD, NQ), they do not generalize well in closed-domain settings, since domain adaptation can be challenging due to poor availability and annotation expense of domain-specific data. This paper proposes SEDAN, an effective self-learning framework to adapt AS2 models for domain-specific applications. We leverage large pre-trained language models to automatically generate domain-specific QA pairs for domain adaptation. We further fine-tune a pre-trained Sentence-BERT architecture to capture semantic relatedness between questions and answer sentences for AS2. Extensive experiments demonstrate the effectiveness of our proposed approach (over existing state-of-the-art AS2 baselines) on different Question Answering benchmark datasets. Rajdeep Sarkar, Sourav Dutta 0001, Haytham Assem, Mihael Arcan, John P. McCrae |
KDD | 4 |
| 2020 | Suggest me a movie for tonight: Leveraging Knowledge Graphs for Conversational RecommendationabstractConversational recommender systems focus on the task of suggesting products to users based on the conversation flow.Recently, the use of external knowledge in the form of knowledge graphs has shown to improve the performance in recommendation and dialogue systems.Information from knowledge graphs aids in enriching those systems by providing additional information such as closely related products and textual descriptions of the items.However, knowledge graphs are incomplete since they do not contain all factual information present on the web.Furthermore, when working on a specific domain, knowledge graphs in its entirety contribute towards extraneous information and noise.In this work, we study several subgraph construction methods and compare their performance across the recommendation task.We incorporate pre-trained embeddings from the subgraphs along with positional embeddings in our models.Extensive experiments show that our method has a relative improvement of at least 5.62% compared to the state-of-the-art on multiple metrics on the recommendation task. Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae |
COLING | 3 |
| 2019 | Utilizing Knowledge Graphs for Neural Machine Translation AugmentationabstractWhile neural networks have led to substantial progress in machine translation, their success depends heavily on large amounts of training data. However, parallel training corpora are not always readily available. Moreover, out-of-vocabulary words---mostly entities and terminological expressions---pose a difficult challenge to Neural Machine Translation systems. Recent efforts have tried to alleviate the data sparsity problem by augmenting the training data using different strategies, such as external knowledge injection. In this paper, we hypothesize that knowledge graphs enhance the semantic feature extraction of neural models, thus optimizing the translation of entities and terminological expressions in texts and consequently leading to better translation quality. We investigate two different strategies for incorporating knowledge graphs into neural models without modifying the neural network architectures. Additionally, we examine the effectiveness of our augmented models on domain-specific texts and ontologies. Our knowledge-graph-augmented neural translation model, dubbed KG-NMT, achieves significant and consistent improvements of +3 BLEU, METEOR and chrF3 on average on the newstest datasets between 2015 and 2018 for the WMT English-German translation task. Diego Moussallem, Axel-Cyrille Ngonga Ngomo, Paul Buitelaar, Mihael Arcan |
K-CAP | 4 |
| 2019 | Comparison of Different Orthographies for Machine Translation of Under-Resourced Dravidian LanguagesabstractUnder-resourced languages are a significant challenge for statistical approaches to machine translation, and recently it has been shown that the usage of training data from closely-related languages can improve machine translation quality of these languages. While languages within the same language family share many properties, many under-resourced languages are written in their own native script, which makes taking advantage of these language similarities difficult. In this paper, we propose to alleviate the problem of different scripts by transcribing the native script into common representation i.e. the Latin script or the International Phonetic Alphabet (IPA). In particular, we compare the difference between coarse-grained transliteration to the Latin script and fine-grained IPA transliteration. We performed experiments on the language pairs English-Tamil, English-Telugu, and English-Kannada translation task. Our results show improvements in terms of the BLEU, METEOR and chrF scores from transliteration and we find that the transliteration into the Latin script outperforms the fine-grained IPA transcription. Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae |
LDK | 2 |
| 2019 | Leveraging Rule-Based Machine Translation Knowledge for Under-Resourced Neural Machine Translation Models
Daniel Torregrosa, Nivranshu Pasricha, Maraim Masoud, Bharathi Raja Chakravarthi, Juan A. Alonso, Noe Casas, Mihael Arcan |
MTSummit (2) | 7 |
| 2018 | Automatic Enrichment of Terminological Resources: the IATE RDF Example
Mihael Arcan, Elena Montiel-Ponsoda, John P. McCrae, Paul Buitelaar |
LREC | 1 |
| 2018 | Improving Wordnets for Under-Resourced Languages Using Machine TranslationabstractWordnets are extensively used in natural language processing, but the current approaches for manually building a wordnet from scratch involves large research groups for a long period of time, which are typically not available for under-resourced languages.Even if wordnet-like resources are available for under-resourced languages, they are often not easily accessible, which can alter the results of applications using these resources.Our proposed method presents an expand approach for improving and generating wordnets with the help of machine translation.We apply our methods to improve and extend wordnets for the Dravidian languages, i.e., Tamil, Telugu, Kannada, which are severly under-resourced languages.We report evaluation results of the generated wordnet senses in term of precision for these languages.In addition to that, we carried out a manual evaluation of the translations for the Tamil language, where we demonstrate that our approach can aid in improving wordnet resources for under-resourced Dravidian languages. Bharathi Raja Chakravarthi, Mihael Arcan, John P. McCrae |
GWC | 2 |
| 2018 | MixedEmotions: An Open-Source Toolbox for Multimodal Emotion AnalysisabstractRecently, there is an increasing tendency to embed functionalities for recognizing emotions from user-generated media content in automated systems such as call-centre operations, recommendations, and assistive technologies, providing richer and more informative user and content profiles. However, to date, adding these functionalities was a tedious, costly, and time-consuming effort, requiring identification and integration of diverse tools with diverse interfaces as required by the use case at hand. The MixedEmotions Toolbox leverages the need for such functionalities by providing tools for text, audio, video, and linked data processing within an easily integrable plug-and-play platform. These functionalities include: 1) for text processing: emotion and sentiment recognition; 2) for audio processing: emotion, age, and gender recognition; 3) for video processing: face detection and tracking, emotion recognition, facial landmark localization, head pose estimation, face alignment, and body pose estimation; and 4) for linked data: knowledge graph integration. Moreover, the MixedEmotions Toolbox is open-source and free. In this paper, we present this toolbox in the context of the existing landscape, and provide a range of detailed benchmarks on standard test-beds showing its state-of-the-art performance. Furthermore, three real-world use cases show its effectiveness, namely, emotion-driven smart TV, call center monitoring, and brand reputation analysis. Paul Buitelaar, Ian D. Wood, Sapna Negi, Mihael Arcan, John P. McCrae, Andrejs Abele, Cécile Robin, Vladimir Andryushechkin, Housam Ziad, Hesam Sagha, Maximilian Schmitt, Björn W. Schuller, J. Fernando Sánchez-Rada, Carlos Angel Iglesias, Carlos Navarro, Andreas Giefer, Nicolaus Heise, Vincenzo Masucci, Francesco A. Danza, Ciro Caterino, Pavel Smrz, Michal Hradis, Filip Povolný, Marek Klimes, Pavel Matejka, Giovanni Tummarello |
IEEE Trans. Multim. | 4 |
| 2017 | Leveraging bilingual terminology to improve machine translation in a CAT environmentabstractAbstract This work focuses on the extraction and integration of automatically aligned bilingual terminology into a Statistical Machine Translation (SMT) system in a Computer Aided Translation scenario. We evaluate the proposed framework that, taking as input a small set of parallel documents, gathers domain-specific bilingual terms and injects them into an SMT system to enhance translation quality. Therefore, we investigate several strategies to extract and align terminology across languages and to integrate it in an SMT system. We compare two terminology injection methods that can be easily used at run-time without altering the normal activity of an SMT system: XML markup and cache-based model. We test the cache-based model on two different domains (information technology and medical) in English, Italian and German, showing significant improvements ranging from 2.23 to 6.78 BLEU points over a baseline SMT system and from 0.05 to 3.03 compared to the widely-used XML markup approach. Mihael Arcan, Marco Turchi, Sara Tonelli, Paul Buitelaar |
Nat. Lang. Eng. | 1 |
| 2016 | Expanding wordnets to new languages with multilingual sense disambiguationabstractPrinceton WordNet is one of the most important resources for natural language processing, but is only available for English. While it has been translated using the expand approach to many other languages, this is an expensive manual process. Therefore it would be beneficial to have a high-quality automatic translation approach that would support NLP techniques, which rely on WordNet in new languages. The translation of wordnets is fundamentally complex because of the need to translate all senses of a word including low frequency senses, which is very challenging for current machine translation approaches. For this reason we leverage existing translations of WordNet in other languages to identify contextual information for wordnet senses from a large set of generic parallel corpora. We evaluate our approach using 10 translated wordnets for European languages. Our experiment shows a significant improvement over translation without any contextual information. Furthermore, we evaluate how the choice of pivot languages affects performance of multilingual word sense disambiguation. Mihael Arcan, John P. McCrae, Paul Buitelaar |
COLING | 1 |
| 2016 | Potential and Limits of Using Post-edits as Reference Translations for MT Evaluation
Maja Popovic, Mihael Arcan, Arle Lommel |
EAMT | 2 |
| 2016 | IRIS: English-Irish Machine Translation System
Mihael Arcan, Caoilfhionn Lane, Eoin Ó Droighneáin, Paul Buitelaar |
LREC | 1 |
| 2016 | PE2rr Corpus: Manual Error Annotation of Automatically Pre-annotated MT Post-edits
Maja Popovic, Mihael Arcan |
LREC | 2 |
| 2016 | ESSOT: An Expert Supporting System for Ontology Translation
Mihael Arcan, Mauro Dragoni, Paul Buitelaar |
NLDB | 1 |
| 2016 | Translating Ontologies in Real-World Settings
Mihael Arcan, Mauro Dragoni, Paul Buitelaar |
ISWC (2) | 1 |
| 2016 | Domain adaptation for ontology localization
John P. McCrae, Mihael Arcan, Kartik Asooja, Jorge Gracia, Paul Buitelaar, Philipp Cimiano |
J. Web Semant. | 2 |
| 2015 | Knowledge Portability with Semantic Expansion of Ontology LabelsabstractMihael Arcan, Marco Turchi, Paul Buitelaar. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mihael Arcan, Marco Turchi, Paul Buitelaar |
ACL (1) | 1 |
| 2015 | MixedEmotions: Social Semantic Emotion Analysis for Innovative Multilingual Big Data Analytics Markets
Mihael Arcan, Paul Buitelaar |
EAMT | 1 |
| 2015 | Identifying main obstacles for statistical machine translation of morphologically rich South Slavic languages
Maja Popovic, Mihael Arcan |
EAMT | 2 |
| 2015 | Poor man's lemmatisation for automatic error classification
Maja Popovic, Mihael Arcan, Eleftherios Avramidis, Aljoscha Burchardt, Arle Lommel |
EAMT | 2 |
| 2013 | Ontology Label Translation
Mihael Arcan, Paul Buitelaar |
HLT-NAACL | 1 |
| 2012 | Experiments with Term Translation
Mihael Arcan, Christian Federmann, Paul Buitelaar |
COLING | 1 |