Surangika Ranathunga

dblp:46/10481 · DBLP profile ↗
← Back
9ranked-venue papers in the field
0as first author
5since 2021 · last 2026
0000-0003-0701-0204ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4Information Retrieval & Web Search · 3Database Systems & Data Management · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Georeferencing complex relative locality descriptions with large language models
Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones
Int. J. Geogr. Inf. Sci.2
2025 Linguistic entity masking to improve cross-lingual representation of multilingual language models for low-resource languages
abstract
Abstract Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM) objective are commonly being used for cross-lingual tasks such as bitext mining. However, the performance of these models is still suboptimal for low-resource languages (LRLs). To improve the language representation of a given multiPLM, it is possible to further pre-train it. This is known as continual pre-training. Previous research has shown that continual pre-training with MLM and subsequently with Translation Language Modelling (TLM) improves the cross-lingual representation of multiPLMs. However, during masking, both MLM and TLM give equal weight to all tokens in the input sequence, irrespective of the linguistic properties of the tokens. In this paper, we introduce a novel masking strategy, Linguistic Entity Masking (LEM) to be used in the continual pre-training step to further improve the cross-lingual representations of existing multiPLMs. In contrast to MLM and TLM, LEM limits masking to the linguistic entity types nouns, verbs and named entities, which hold a higher prominence in a sentence. Secondly, we limit masking to a single token within the linguistic entity span thus keeping more context, whereas, in MLM and TLM, tokens are masked randomly. We evaluate the effectiveness of LEM using three downstream tasks, namely bitext mining, parallel data curation and code-mixed sentiment analysis using three low-resource language pairs English-Sinhala, English-Tamil, and Sinhala-Tamil. Experiment results show that continually pre-training a multiPLM with LEM outperforms a multiPLM continually pre-trained with MLM+TLM for all three tasks.
Aloka Fernando, Surangika Ranathunga
Knowl. Inf. Syst.2
2025 Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning
abstract
This research pioneers the use of fine-tuned Large Language Models (LLMs) to automate Systematic Literature Reviews (SLRs), presenting a significant and novel contribution in integrating AI to enhance academic research methodologies. Our study employed advanced fine-tuning methodologies on open sourced LLMs, applying textual data mining techniques to automate the knowledge discovery and synthesis phases of an SLR process, thus demonstrating a practical and efficient approach for extracting and analyzing high-quality information from large academic datasets. The results maintained high fidelity in factual accuracy in LLM responses, and were validated through the replication of an existing PRISMA-conforming SLR. Our research proposed solutions for mitigating LLM hallucination and proposed mechanisms for tracking LLM responses to their sources of information, thus demonstrating how this approach can meet the rigorous demands of scholarly research. The findings ultimately confirmed the potential of fine-tuned LLMs in streamlining various labor-intensive processes of conducting literature reviews. As a scalable proof-of-concept, this study highlights the broad applicability of our approach across multiple research domains. The potential demonstrated here advocates for updates to PRISMA reporting guidelines, incorporating AI-driven processes to ensure methodological transparency and reliability in future SLRs. This study broadens the appeal of AI-enhanced tools across various academic and research fields, demonstrating how to conduct comprehensive and accurate literature reviews with more efficiency in the face of ever-increasing volumes of academic studies while maintaining high standards.
Teo Susnjak, Peter Hwang, Napoleon H. Reyes, Andre L. C. Barczak, Timothy R. McIntosh, Surangika Ranathunga
ACM Trans. Knowl. Discov. Data6
2023 Exploiting bilingual lexicons to improve multilingual embedding-based document and sentence alignment for low-resource languages
Aloka Fernando, Surangika Ranathunga, Dilan Sachintha, Lakmali Piyarathna, Charith Rajitha
Knowl. Inf. Syst.2
2022 Adapter-based fine-tuning of pre-trained multilingual language models for code-mixed and code-switched text classification
Himashi Rathnayake, Janani Sumanapala, Raveesha Rukshani, Surangika Ranathunga
Knowl. Inf. Syst.4
2019 Mathematical Expression Extraction from Unstructured Plain Text
Kulakshi Fernando, Surangika Ranathunga, Gihan Dias
NLDB2
2019 Model Answer Generation for Word-Type Questions in Elementary Mathematics
Sakthithasan Rajpirathap, Surangika Ranathunga
NLDB2
2016 Tamil Morphological Analyzer Using Support Vector Machines
Mokanarangan Thayaparan, Pranavan Theivendiram, Megala Uthayakumar, Nilusija Nadarasamoorthy, Gihan Dias, Sanath Jayasena, Surangika Ranathunga
NLDB7
2016 Domain-Specific Term Extraction for Concept Identification in Ontology Construction
abstract
An ontology is a formal and explicit specification of a shared conceptualization. Manual construction of domain ontology does not adequately satisfy requirements of new applications, because they need a more dynamic ontology and the possibility to manage a considerable quantity of concepts that humans cannot achieve alone. Researchers have discussed ontology learning as a solution to overcome issues related to the manual construction of ontology. Ontology learning is either an automatic or semi-automatic process to apply methods for building ontology from scratch, or enriching or adapting an existing ontology. This research focuses on improving the process of term extraction for identifying concepts in ontology learning. Available approaches for term extraction process are limited in various ways. These limitations include: (1) obtaining domain-specific terms from a domain expert as seed words without automatically discovering them from the corpus, and (2) unsuitable usage of corpora in discovering domain-specific terms for multiple domains. Our study uses linguistic analysis and statistical calculations to extract domain-specific simple and complex terms to overcome this first limitation. To eliminate the second limitation, we use multiple contrastive corpora that reduce the biasness in using a single contrastive corpus. Evaluations show that our system is better at extracting terms when compared with the previous research that used the same corpora.
Kiruparan Balachandran, Surangika Ranathunga
WI2