VLDB 2026 Research / reviewers in the wild / expert
Tharindu Ranasinghe
dblp:242/4755
· DBLP profile ↗
26ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-3207-3821ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUNIChus: MUltilingual News Image Captioning Benchmark
Yuji Chen, Alistair Plum, Hansi Hettiarachchi, Diptesh Kanojia, Saroj Basnet, Marcos Zampieri, Tharindu Ranasinghe |
LREC | 7 |
| 2026 | Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
Alistair Plum, Laura Bernardy, Tharindu Ranasinghe |
LREC | 3 |
| 2026 | UKSC-JP: A Legal Judgement Prediction Benchmark for the United Kingdom Supreme Court
Damith Premasiri, Tharindu Ranasinghe, Sandani Abeywardena, Ruslan Mitkov |
WorldCIST (2) | 2 |
| 2026 | On the performance of large language models on introductory programming assignmentsabstractAbstract Recent advances in artificial intelligence (AI), machine learning (ML), and natural language processing (NLP) have led to the development of a new generation of Large Language Models (LLMs) trained on massive amounts of data. Commercial applications (e.g., ChatGPT) have made this available to the general public, enabling the use of LLMs to produce high-quality texts for academic and professional purposes. Educational institutions are increasingly aware of students’ use of AI-generated content and are researching its impact and potential misuse. Computer Science (CS) and related fields are particularly affected, as LLMs can also generate programming code in various languages. To understand the potential impact of publicly available LLMs in CS education, we extend our previously introduced (Raihan et al. 2024), a framework comprising hundreds of programming exercise prompts and multiple-choice questions from introductory CS and programming courses. We provide experimental results on , evaluating the performance of several LLMs in generating Python code and answering basic computer science and programming questions, offering insights into the implications of this technology for CS education. Dhiman Goswami, Sadiya Sayara Chowdhury Puspo, Mohammed Latif Siddiq, Christian D. Newman, Tharindu Ranasinghe, Joanna C. S. Santos, Marcos Zampieri |
J. Intell. Inf. Syst. | 6 |
| 2025 | Sinhala Encoder-only Language Models and EvaluationabstractRecently, language models (LMs) have produced excellent results in many natural language processing (NLP) tasks. However, their effectiveness is highly dependent on available pre-training resources, which is particularly challenging for low-resource languages such as Sinhala. Furthermore, the scarcity of benchmarks to evaluate LMs is also a major concern for low-resource languages. In this paper, we address these two challenges for Sinhala by (i) collecting the largest monolingual corpus for Sinhala, (ii) training multiple LMs on this corpus and (iii) compiling the first Sinhala NLP benchmark (Sinhala-GLUE) and evaluating LMs on it. We show the Sinhala LMs trained in this paper outperform the popular multilingual LMs, such as XLM-R and existing Sinhala LMs in downstream NLP tasks. All the trained LMs are publicly available. We also make Sinhala-GLUE publicly available as a public leaderboard, and we hope that it will enable further advancements in developing and evaluating LMs for Sinhala. Tharindu Ranasinghe, Hansi Hettiarachchi, Nadeesha Pathirana, Damith Premasiri, Lasitha Uyangodage, Isuri Anuradha, Alistair Plum, Paul Rayson, Ruslan Mitkov |
ACL (1) | 1 |
| 2025 | Deep learning approaches to lexical simplification: A surveyabstractAbstract Lexical Simplification (LS) is the task of substituting complex words within a sentence for simpler alternatives while maintaining the sentence’s original meaning. LS is the lexical component of Text Simplification (TS) systems with the aim of improving accessibility to various target populations such as individuals with low literacy or reading disabilities. Prior surveys have been published several years before the introduction of transformers, transformer-based large language models (LLMs), and prompt learning that have drastically changed the field of NLP. The high performance of these models has sparked renewed interest in LS. To reflect these recent advances, we present a comprehensive survey of papers published since 2017 on LS and its sub-tasks focusing on deep learning. Finally, we describe available benchmark datasets for the future development of LS systems. Kai North, Tharindu Ranasinghe, Matthew Shardlow, Marcos Zampieri |
J. Intell. Inf. Syst. | 2 |
| 2025 | Survey on legal information extraction: current status and open challengesabstractAbstract The goal of information extraction is to extract structural knowledge (such as entities, relations and events) from plain and unstructured texts. Information extraction in legal documents has recently gained a lot of attention in the natural language processing (NLP) community due to the high demand for efficient information extraction for legal practitioners and companies. Given that the legal documents are unique and their processing is challenging, there is a pressing need for applications of NLP techniques to tackle these challenges. In this research, we present a survey on the recent advancements in legal information extraction focusing on three tasks: named entity recognition, relationship extraction and event detection. We report language resources and systems in multiple jurisdictions and languages for each task. Based on the thorough review conducted, we identify insights into the techniques employed and promising research directions that merit further exploration in future studies. We maintain a public repository and consistently update related resources at https://github.com/DamithDR/legalinformationextraction . Damith Premasiri, Tharindu Ranasinghe, Ruslan Mitkov, Mo El-Haj, Ingo Frommholz |
Knowl. Inf. Syst. | 2 |
| 2025 | Complex Concept-Based Readability Estimation from Arabic CurriculumabstractThis article presents an approach to readability estimation that focuses on conceptual rather than linguistic complexity, using the extensive SaudiTextBooks textbooks. We introduce DARES 2.0 , an enhanced concept-based readability training dataset designed to estimate the readability of Saudi educational texts. Building on DARES 1.0, DARES 2.0 extends the scope of conceptual complexity by replacing repetitive concepts and manually revising the input features with unique terms and their surrounding contexts from the SaudiTextBooks, spanning grades 1 to 12. The refined DARES 2.0 is employed to fine-tune pre-trained transformer models, including XLM-R Base, mBERT, AraELECTRA, AraBERTv2, and CAMeLBERTmix. The findings suggest that both the dataset and experimental setup require further development to ensure a larger, higher-quality dataset and to support more extensive fine-tuning experiments, in addition to exploring transfer learning from other languages and enhancing the diversity and richness of Arabic concepts. These developments pave the way for further advancements in concept-based readability estimation in educational contexts in future work. Sultan Almujaiwel, Damith Premasiri, Tharindu Ranasinghe, Mo El-Haj, Ruslan Mitkov |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | Towards Generalized Offensive Language Identification
Alphaeus Dmonte, Tejas Arya, Tharindu Ranasinghe, Marcos Zampieri |
ASONAM (1) | 3 |
| 2024 | DORE: A Dataset for Portuguese Definition GenerationabstractDefinition modelling (DM) is the task of automatically generating a dictionary definition of a specific word. Computational systems that are capable of DM can have numerous applications benefiting a wide range of audiences. As DM is considered a supervised natural language generation problem, these systems require large annotated datasets to train the machine learning (ML) models. Several DM datasets have been released for English and other high-resource languages. While Portuguese is considered a mid/high-resource language in most natural language processing tasks and is spoken by more than 200 million native speakers, there is no DM dataset available for Portuguese. In this research, we fill this gap by introducing DORE; the first dataset for Definition MOdelling for PoRtuguEse containing more than 100,000 definitions. We also evaluate several deep learning based DM models on DORE and report the results. The dataset and the findings of this paper will facilitate research and study of Portuguese in wider contexts. Anna Beatriz Dimas Furtado, Tharindu Ranasinghe, Frédéric Blain, Ruslan Mitkov |
LREC/COLING | 2 |
| 2024 | NSina: A News Corpus for SinhalaabstractThe introduction of large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. This is especially evident in low-resource languages, such as Sinhala, which face two primary challenges: the lack of substantial training data and limited benchmarking datasets. In response, this study introduces NSina, a comprehensive news corpus of over 500,000 articles from popular Sinhala news websites, along with three NLP tasks: news media identification, news category prediction, and news headline generation. The release of NSina aims to provide a solution to challenges in adapting LLMs to Sinhala, offering valuable resources and benchmarks for improving NLP in the Sinhala language. NSina is the largest news corpus for Sinhala, available up to date. Hansi Hettiarachchi, Damith Premasiri, Lasitha Uyangodage, Tharindu Ranasinghe |
LREC/COLING | 4 |
| 2024 | Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New LanguageabstractRelation extraction is essential for extracting and understanding biographical information in the context of digital humanities and related subjects. There is a growing interest in the community to build datasets capable of training machine learning models to extract relationships. However, annotating such datasets can be expensive and time-consuming, in addition to being limited to English. This paper applies guided distant supervision to create a large biographical relationship extraction dataset for German. Our dataset, composed of more than 80,000 instances for nine relationship types, is the largest biographical German relationship extraction dataset. We also create a manually annotated dataset with 2000 instances to evaluate the models and release it together with the dataset compiled using guided distant supervision. We train several state-of-the-art machine learning models on the automatically created dataset and release them as well. Furthermore, we experiment with multilingual and cross-lingual zero-shot experiments that could benefit many low-resource languages. Alistair Plum, Tharindu Ranasinghe, Christoph Purschke |
LREC/COLING | 2 |
| 2024 | MentalHelp: A Multi-Task Dataset for Mental Health in Social MediaabstractEarly detection of mental health disorders is an essential step in treating and preventing mental health conditions. Computational approaches have been applied to users’ social media profiles in an attempt to identify various mental health conditions such as depression, PTSD, schizophrenia, and eating disorders. The interest in this topic has motivated the creation of various depression detection datasets. However, annotating such datasets is expensive and time-consuming, limiting their size and scope. To overcome this limitation, we present MentalHelp, a large-scale semi-supervised mental disorder detection dataset containing 14 million instances. The corpus was collected from Reddit and labeled in a semi-supervised way using an ensemble of three separate models - flan-T5, Disor-BERT, and Mental-BERT. Sadiya Sayara Chowdhury Puspo, Shafkat Farabi, Ana-Maria Bucur, Tharindu Ranasinghe, Marcos Zampieri |
LREC/COLING | 5 |
| 2024 | What do Large Language Models Need for Machine Translation Evaluation?abstractShenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Fred Blain. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Frédéric Blain |
EMNLP | 6 |
| 2024 | A Survey of Multimodal Sarcasm Detection
Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong 0001, Marcos Zampieri |
IJCAI | 2 |
| 2024 | CSEPrompts: A Benchmark of Introductory Computer Science Prompts
Dhiman Goswami, Sadiya Sayara Chowdhury Puspo, Christian D. Newman, Tharindu Ranasinghe, Marcos Zampieri |
ISMIS | 5 |
| 2023 | Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveabstractTharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh |
EMNLP | 3 |
| 2023 | Offensive language identification with multi-task learning
Marcos Zampieri, Tharindu Ranasinghe, Diptanu Sarkar, Alexander Ororbia |
J. Intell. Inf. Syst. | 2 |
| 2023 | OffensEval 2023: Offensive language identification in the age of Large Language ModelsabstractAbstract The OffensEval shared tasks organized as part of SemEval-2019–2020 were very popular, attracting over 1300 participating teams. The two editions of the shared task helped advance the state of the art in offensive language identification by providing the community with benchmark datasets in Arabic, Danish, English, Greek, and Turkish. The datasets were annotated using the OLID hierarchical taxonomy, which since then has become the de facto standard in general offensive language identification research and was widely used beyond OffensEval. We present a survey of OffensEval and related competitions, and we discuss the main lessons learned. We further evaluate the performance of Large Language Models (LLMs), which have recently revolutionalized the field of Natural Language Processing. We use zero-shot prompting with six popular LLMs and zero-shot learning with two task-specific fine-tuned BERT models, and we compare the results against those of the top-performing teams at the OffensEval competitions. Our results show that while some LMMs such as Flan-T5 achieve competitive performance, in general LLMs lag behind the best OffensEval systems. Marcos Zampieri, Sara Rosenthal, Preslav Nakov, Alphaeus Dmonte, Tharindu Ranasinghe |
Nat. Lang. Eng. | 5 |
| 2022 | ALEXSIS-PT: A New Resource for Portuguese Lexical SimplificationabstractLexical simplification (LS) is the task of automatically replacing complex words for easier ones making texts more accessible to various target populations (e.g. individuals with low literacy, individuals with learning disabilities, second language learners). To train and test models, LS systems usually require corpora that feature complex words in context along with their potential substitutions. To continue improving the performance of LS systems we introduce ALEXSIS-PT, a novel multi-candidate dataset for Brazilian Portuguese LS containing 9,605 candidate substitutions for 387 complex words. ALEXSIS-PT has been compiled following the ALEXSIS-ES protocol for Spanish opening exciting new avenues for cross-lingual models. ALEXSIS-PT is the first LS multi-candidate dataset that contains Brazilian newspaper articles. We evaluated three models for substitute generation on this dataset, namely mBERT, XLM-R, and BERTimbau. The latter achieved the highest performance across all evaluation metrics. Kai North, Marcos Zampieri, Tharindu Ranasinghe |
COLING | 3 |
| 2022 | Biographical Semi-Supervised Relation Extraction DatasetabstractExtracting biographical information from online documents is a popular research topic among the information extraction (IE) community. Various natural language processing (NLP) techniques such as text classification, text summarisation and relation extraction are commonly used to achieve this. Among these techniques, RE is the most common since it can be directly used to build biographical knowledge graphs. RE is usually framed as a supervised machine learning (ML) problem, where ML models are trained on annotated datasets. However, there are few annotated datasets for RE since the annotation process can be costly and time-consuming. To address this, we developedBiographical, the first semi-supervised dataset for RE. The dataset, which is aimed towards digital humanities (DH) and historical research, is automatically compiled by aligning sentences from Wikipedia articles with matching structured data from sources including Pantheon and Wikidata. By exploiting the structure of Wikipedia articles and robust named entity recognition (NER), we match information with relatively high precision in order to compile annotated relation pairs for ten different relations that are important in the DH domain. Furthermore, we demonstrate the effectiveness of the dataset by training a state-of-the-art neural model to classify relation pairs, and evaluate it on a manually annotated gold standard set.Biographical is primarily aimed at training neural models for RE within the domain of digital humanities and history, but as we discuss at the end of this paper, it can be useful for other purposes as well. Alistair Plum, Tharindu Ranasinghe, Spencer Jones, Constantin Orasan, Ruslan Mitkov |
SIGIR | 2 |
| 2022 | Multilingual Offensive Language Identification for Low-resource LanguagesabstractOffensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such content (e.g., hate speech, cyberbullying, and cyberaggression). The clear majority of these studies deal with English partially because most annotated datasets available contain English data. In this article, we take advantage of available English datasets by applying cross-lingual contextual word embeddings and transfer learning to make predictions in low-resource languages. We project predictions on comparable data in Arabic, Bengali, Danish, Greek, Hindi, Spanish, and Turkish. We report results of 0.8415 F1 macro for Bengali in TRAC-2 shared task [23], 0.8532 F1 macro for Danish and 0.8701 F1 macro for Greek in OffensEval 2020 [58], 0.8568 F1 macro for Hindi in HASOC 2019 shared task [27], and 0.7513 F1 macro for Spanish in in SemEval-2019 Task 5 (HatEval) [7], showing that our approach compares favorably to the best systems submitted to recent shared tasks on these three languages. Additionally, we report competitive performance on Arabic and Turkish using the training and development sets of OffensEval 2020 shared task. The results for all languages confirm the robustness of cross-lingual contextual embeddings and transfer learning for this task. Tharindu Ranasinghe, Marcos Zampieri |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2020 | TransQuest: Translation Quality Estimation with Cross-lingual TransformersabstractRecent years have seen big advances in the field of sentence-level quality estimation (QE), largely as a result of using neural-based architectures.However, the majority of these methods work only on the language pair they are trained on and need retraining for new language pairs.This process can prove difficult from a technical point of view and is usually computationally expensive.In this paper we propose a simple QE framework based on cross-lingual transformers, and we use it to implement and evaluate two different neural architectures.Our evaluation shows that the proposed methods achieve state-of-the-art results outperforming current open-source quality estimation frameworks when trained on datasets from WMT.In addition, the framework proves very useful in transfer learning settings, especially when dealing with low-resourced languages, allowing us to obtain very competitive results. Tharindu Ranasinghe, Constantin Orasan, Ruslan Mitkov |
COLING | 1 |
| 2020 | Intelligent Translation Memory Matching and Retrieval with Sentence EncodersabstractMatching and retrieving previously translated segments from the Translation Memory is a key functionality in Translation Memories systems. However this matching and retrieving process is still limited to algorithms based on edit distance which we have identified as a major drawback in Translation Memories systems. In this paper, we introduce sentence encoders to improve matching and retrieving process in Translation Memories systems - an effective and efficient solution to replace edit distance-based algorithms. Tharindu Ranasinghe, Constantin Orasan, Ruslan Mitkov |
EAMT | 1 |
| 2020 | Multilingual Offensive Language Identification with Cross-lingual EmbeddingsabstractOffensive content is pervasive in social media and a reason for concern to companies and government organizations.Several studies have been recently published investigating methods to detect the various forms of such content (e.g.hate speech, cyberbulling, and cyberaggression).The clear majority of these studies deal with English partially because most annotated datasets available contain English data.In this paper, we take advantage of English data available by applying cross-lingual contextual word embeddings and transfer learning to make predictions in languages with less resources.We project predictions on comparable data in Bengali, Hindi, and Spanish and we report results of 0.8415 F1 macro for Bengali, 0.8568 F1 macro for Hindi, and 0.7513 F1 macro for Spanish.Finally, we show that our approach compares favorably to the best systems submitted to recent shared tasks on these three languages, confirming the robustness of cross-lingual contextual embeddings and transfer learning for this task. Tharindu Ranasinghe, Marcos Zampieri |
EMNLP (1) | 1 |
| 2020 | Offensive Language Identification in GreekabstractAs offensive language has become a rising issue for online communities and social media platforms, researchers have been investigating ways of coping with abusive content and developing systems to detect its different types: cyberbullying, hate speech, aggression, etc. With a few notable exceptions, most research on this topic so far has dealt with English. This is mostly due to the availability of language resources for English. To address this shortcoming, this paper presents the first Greek annotated dataset for offensive language identification: the Offensive Greek Tweet Dataset (OGTD). OGTD is a manually annotated dataset containing 4,779 posts from Twitter annotated as offensive and not offensive. Along with a detailed description of the dataset, we evaluate several computational models trained and tested on this data. Zeses Pitenis, Marcos Zampieri, Tharindu Ranasinghe |
LREC | 3 |