Surangika Ranathunga

dblp:46/10481 · DBLP profile ↗
← Back
42ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-0701-0204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 5Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
Hannah Liu, Murphy Tian, Iqra Ali, Haonan Gao, Qiaoyiwen Wu, Blair Yang, Uthayasanker Thayasivam, Annie En-Shiun Lee, Pakawat Nakwijit, Surangika Ranathunga, Ravi Shekhar
LREC10
2026 Georeferencing complex relative locality descriptions with large language models
Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones
Int. J. Geogr. Inf. Sci.2
2026 Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-Resource Language Translation
abstract
Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language’s representation in the model are limited. This restricts the capabilities of domain-specific NMT systems for low-resource languages (LRLs). As a solution, parallel data from auxiliary domains can be used either to fine-tune or to further pre-train the msLM. We present an evaluation of the effectiveness of these two techniques in the context of domain-specific LRL-NMT. We also explore the impact of domain divergence on NMT model performance. We recommend several strategies for utilizing auxiliary parallel data in building domain-specific NMT models for LRLs.
Surangika Ranathunga, Shravan Nayak, Annie En-Shiun Lee, Shih-Ting Cindy Huang, Yuchen Zeng 0001, Yanke Mao, Yun-Hsiang Ray Chan, Songchen Yuan, Anthony Rinaldi
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2025 Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
Kavindu Warnakulasuriya, Prabhash Dissanayake, Navindu De Silva, Stephen Cranefield, Bastin Tony Roy Savarimuthu, Surangika Ranathunga, Nisansa de Silva
COINE6
2025 Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics
abstract
Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from webmined corpora.Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique.However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs.This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view.We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics.The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs.We publicly release the source code and the curated datasets 1 .
Aloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake, Surangika Ranathunga
EMNLP5
2025 Linguistic entity masking to improve cross-lingual representation of multilingual language models for low-resource languages
abstract
Abstract Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM) objective are commonly being used for cross-lingual tasks such as bitext mining. However, the performance of these models is still suboptimal for low-resource languages (LRLs). To improve the language representation of a given multiPLM, it is possible to further pre-train it. This is known as continual pre-training. Previous research has shown that continual pre-training with MLM and subsequently with Translation Language Modelling (TLM) improves the cross-lingual representation of multiPLMs. However, during masking, both MLM and TLM give equal weight to all tokens in the input sequence, irrespective of the linguistic properties of the tokens. In this paper, we introduce a novel masking strategy, Linguistic Entity Masking (LEM) to be used in the continual pre-training step to further improve the cross-lingual representations of existing multiPLMs. In contrast to MLM and TLM, LEM limits masking to the linguistic entity types nouns, verbs and named entities, which hold a higher prominence in a sentence. Secondly, we limit masking to a single token within the linguistic entity span thus keeping more context, whereas, in MLM and TLM, tokens are masked randomly. We evaluate the effectiveness of LEM using three downstream tasks, namely bitext mining, parallel data curation and code-mixed sentiment analysis using three low-resource language pairs English-Sinhala, English-Tamil, and Sinhala-Tamil. Experiment results show that continually pre-training a multiPLM with LEM outperforms a multiPLM continually pre-trained with MLM+TLM for all three tasks.
Aloka Fernando, Surangika Ranathunga
Knowl. Inf. Syst.2
2025 SiTSE: Sinhala Text Simplification Dataset and Evaluation
abstract
Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this article, we present a human-curated sentence-level text simplification dataset for the Sinhala language. Our evaluation dataset contains 1,000 complex sentences and 3,000 corresponding simplified sentences produced by three different human annotators. We model the text simplification task as a zero-shot and zero-resource sequence-to-sequence (seq-seq) task on the multilingual language models mT5 and mBART. We exploit auxiliary data from related seq-seq tasks and explore the possibility of using intermediate task transfer learning (ITTL). Our analysis shows that ITTL outperforms the previously proposed zero-resource methods for text simplification. Our findings also highlight the challenges in evaluating text simplification systems and support the calls for improved metrics for measuring the quality of automated text simplification systems that would suit low-resource languages as well. Our code and data are publicly available: https://github.com/brainsharks-fyp17/Sinhala-Text-Simplification-Dataset-and-Evaluation .
Surangika Ranathunga, Rumesh Sirithunga, Himashi Rathnayake, Lahiru de Silva, Thamindu Aluthwala, Saman Peramuna, Ravi Shekhar
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2025 Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning
abstract
This research pioneers the use of fine-tuned Large Language Models (LLMs) to automate Systematic Literature Reviews (SLRs), presenting a significant and novel contribution in integrating AI to enhance academic research methodologies. Our study employed advanced fine-tuning methodologies on open sourced LLMs, applying textual data mining techniques to automate the knowledge discovery and synthesis phases of an SLR process, thus demonstrating a practical and efficient approach for extracting and analyzing high-quality information from large academic datasets. The results maintained high fidelity in factual accuracy in LLM responses, and were validated through the replication of an existing PRISMA-conforming SLR. Our research proposed solutions for mitigating LLM hallucination and proposed mechanisms for tracking LLM responses to their sources of information, thus demonstrating how this approach can meet the rigorous demands of scholarly research. The findings ultimately confirmed the potential of fine-tuned LLMs in streamlining various labor-intensive processes of conducting literature reviews. As a scalable proof-of-concept, this study highlights the broad applicability of our approach across multiple research domains. The potential demonstrated here advocates for updates to PRISMA reporting guidelines, incorporating AI-driven processes to ensure methodological transparency and reliability in future SLRs. This study broadens the appeal of AI-enhanced tools across various academic and research fields, demonstrating how to conduct comprehensive and accurate literature reviews with more efficiency in the face of ever-increasing volumes of academic studies while maintaining high standards.
Teo Susnjak, Peter Hwang, Napoleon H. Reyes, Andre L. C. Barczak, Timothy R. McIntosh, Surangika Ranathunga
ACM Trans. Knowl. Discov. Data6
2024 Norm Violation Detection in Multi-Agent Systems Using Large Language Models - A Pilot Study
Shawn He, Surangika Ranathunga, Stephen Cranefield, Bastin Tony Roy Savarimuthu
COINE2
2024 Harnessing the Power of LLMs for Normative Reasoning in MASs
Bastin Tony Roy Savarimuthu, Surangika Ranathunga, Stephen Cranefield
COINE2
2024 Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
abstract
Surangika Ranathunga, Nisansa De Silva, Velayuthan Menan, Aloka Fernando, Charitha Rathnayake. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Surangika Ranathunga, Nisansa de Silva, Menan Velayuthan, Aloka Fernando, Charitha Rathnayake
EACL (1)1
2024 LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
abstract
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001
EMNLP31
2024 AdapterFusion-based multi-task learning for code-mixed and code-switched text classification
Himashi Rathnayake, Janani Sumanapala, Raveesha Rukshani, Surangika Ranathunga
Eng. Appl. Artif. Intell.4
2024 Use of prompt-based learning for code-mixed and code-switched text classification
abstract
Abstract Code-mixing and code-switching (CMCS) are prevalent phenomena observed in social media conversations and various other modes of communication. When developing applications such as sentiment analysers and hate-speech detectors that operate on this social media data, CMCS text poses challenges. Recent studies have demonstrated that prompt-based learning of pre-trained language models outperforms full fine-tuning across various tasks. Despite the growing interest in classifying CMCS text, the effectiveness of prompt-based learning for the task remains unexplored. This paper presents an extensive exploration of prompt-based learning for CMCS text classification and the first comprehensive analysis of the impact of the script on classifying CMCS text. Our study reveals that the performance in classifying CMCS text is significantly influenced by the inclusion of multiple scripts and the intensity of code-mixing. In response, we introduce a novel method, Dynamic+AdapterPrompt , which employs distinct models for each script, integrated with adapters. While DynamicPrompt captures the script-specific representation of the text, AdapterPrompt emphasizes capturing the task-oriented functionality. Our experiments on Sinhala-English, Kannada-English, and Hindi-English datasets for sentiment classification, hate-speech detection, and humour detection tasks show that our method outperforms strong fine-tuning baselines and basic prompting strategies.
Pasindu Udawatta, Indunil Udayangana, Chathulanka Gamage, Ravi Shekhar, Surangika Ranathunga
World Wide Web (WWW)5
2023 Exploiting bilingual lexicons to improve multilingual embedding-based document and sentence alignment for low-resource languages
Aloka Fernando, Surangika Ranathunga, Dilan Sachintha, Lakmali Piyarathna, Charith Rajitha
Knowl. Inf. Syst.2
2022 BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification
abstract
This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis shows that out of the pre-trained multilingual models that include Sinhala (XLM-R, LaBSE, and LASER), XLM-R is the best model by far for Sinhala text classification. We also pre-train two RoBERTa-based monolingual Sinhala models, which are far superior to the existing pre-trained language models for Sinhala. We show that when fine-tuned, these pre-trained language models set a very strong baseline for Sinhala text classification and are robust in situations where labeled data is insufficient for fine-tuning. We further provide a set of recommendations for using pre-trained models for Sinhala text classification. We also introduce new annotated datasets useful for future research in Sinhala text classification and publicly release our pre-trained models.
Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, Sanath Jayasena
LREC3
2022 Dataset and Baseline for Automatic Student Feedback Analysis
abstract
In this paper, we present a student feedback corpus, which contains 3000 instances of feedback written by university students. This dataset has been annotated for aspect terms, opinion terms, polarities of the opinion terms towards targeted aspects, document-level opinion polarities and sentence separations. We develop a hierarchical taxonomy for aspect categorization, which covers all the areas of the teaching-learning process. We annotated both implicit and explicit aspects using this taxonomy. Annotation methodology, difficulties faced during the annotation, and the details about the aspect term categorization have been discussed in detail. This annotated corpus can be used for Aspect Extraction, Aspect Level Sentiment Analysis, and Document Level Sentiment Analysis. Also the baseline results for all three tasks are given in the paper.
Missaka Herath, Kushan Chamindu, Hashan Maduwantha, Surangika Ranathunga
LREC4
2022 Adapter-based fine-tuning of pre-trained multilingual language models for code-mixed and code-switched text classification
Himashi Rathnayake, Janani Sumanapala, Raveesha Rukshani, Surangika Ranathunga
Knowl. Inf. Syst.4
2021 Data Augmentation to Address Out of VocabularyProblem in Low Resource Sinhala English Neural Machine Translation
Aloka Fernando, Surangika Ranathunga
PACLIC2
2021 Sentiment Analysis of Sinhala News Comments
abstract
Sinhala is a low-resource language, for which basic language and linguistic tools have not been properly defined. This affects the development of NLP-based end-user applications for Sinhala. Thus, when implementing NLP tools such as sentiment analyzers, we have to rely only on language-independent techniques. This article presents the use of such language-independent techniques in implementing a sentiment analysis system for Sinhala news comments. We demonstrate that for low-resource languages such as Sinhala, the use of recently introduced word embedding models as semantic features can compensate for the lack of well-developed language-specific linguistic or language resources, and text classification with acceptable accuracy is indeed possible using both traditional statistical classifiers and Deep Learning models. The developed classification models, a corpus of 8.9 million tokens extracted from Sinhala news articles and user comments, and Sinhala Word2Vec and fastText word embedding models are now available for public use; 9,048 news comments annotated with POSITIVE/NEGATIVE/NEUTRAL polarities have also been released.
Surangika Ranathunga, Isuru Udara Liyanage
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2020 Word Embedding Evaluation for Sinhala
abstract
This paper presents the first ever comprehensive evaluation of different types of word embeddings for Sinhala language. Three standard word embedding models, namely, Word2Vec (both Skipgram and CBOW), FastText, and Glove are evaluated under two types of evaluation methods: intrinsic evaluation and extrinsic evaluation. Word analogy and word relatedness evaluations were performed in terms of intrinsic evaluation, while sentiment analysis and part-of-speech (POS) tagging were conducted as the extrinsic evaluation tasks. Benchmark datasets used for intrinsic evaluations were carefully crafted considering specific linguistic features of Sinhala. In general, FastText word embeddings with 300 dimensions reported the finest accuracies across all the evaluation tasks, while Glove reported the lowest results.
Dimuthu Lakmal, Surangika Ranathunga, Saman Peramuna, Indu Herath
LREC2
2020 Multi-lingual Mathematical Word Problem Generation using Long Short Term Memory Networks with Enhanced Input Features
abstract
A Mathematical Word Problem (MWP) differs from a general textual representation due to the fact that it is comprised of numerical quantities and units, in addition to text. Therefore, MWP generation should be carefully handled. When it comes to multi-lingual MWP generation, language specific morphological and syntactic features become additional constraints. Standard template-based MWP generation techniques are incapable of identifying these language specific constraints, particularly in morphologically rich yet low resource languages such as Sinhala and Tamil. This paper presents the use of a Long Short Term Memory (LSTM) network that is capable of generating elementary level MWPs, while satisfying the aforementioned constraints. Our approach feeds a combination of character embeddings, word embeddings, and Part of Speech (POS) tag embeddings to the LSTM, in which attention is provided for numerical values and units. We trained our model for three languages, English, Sinhala and Tamil using separate MWP datasets. Irrespective of the language and the type of the MWP, our model could generate accurate single sentenced and multi sentenced problems. Accuracy reported in terms of average BLEU score for English, Sinhala and Tamil languages were 22.97%, 24.49% and 20.74%, respectively.
Vijini Liyanage, Surangika Ranathunga
LREC2
2020 Two-Step Memory Networks for Deep Semantic Parsing of Geometry Word Problems
Ishadi Jayasinghe, Surangika Ranathunga
SOFSEM2
2019 Mathematical Expression Extraction from Unstructured Plain Text
Kulakshi Fernando, Surangika Ranathunga, Gihan Dias
NLDB2
2019 Model Answer Generation for Word-Type Questions in Elementary Mathematics
Sakthithasan Rajpirathap, Surangika Ranathunga
NLDB2
2018 Anomaly Detection in Industrial Software Systems - Using Variational Autoencoders
Tharindu Kumarage, Nadun De Silva, Malsha Ranawaka, Chamal Kuruppu, Surangika Ranathunga
ICPRAM5
2018 Annotating Opinions and Opinion Targets in Student Course Feedback
Janaka Chathuranga, Shanika Ediriweera, Ravindu Hasantha, Pranidhith Munasinghe, Surangika Ranathunga
LREC5
2018 Improving domain-specific SMT for low-resourced languages using data from different domains
Fathima Farhath, Pranavan Theivendiram, Surangika Ranathunga, Sanath Jayasena, Gihan Dias
LREC3
2018 Handling Rare Word Problem using Synthetic Training Data for Sinhala and Tamil Neural Machine Translation
Pasindu Tennage, Prabath Sandaruwan, Malith Thilakarathne, Achini Herath, Surangika Ranathunga
LREC5
2018 Graph Based Semi-Supervised Learning Approach for Tamil POS tagging
Mokanarangan Thayaparan, Surangika Ranathunga, Uthayasanker Thayasivam
LREC2
2017 Monitoring Health of Large Scale Software Systems Using Drift Detection Techniques
L. H. C. Prabodha, W. R. R. Vithanage, L. T. Ranaweera, D. M. M. A. I. B. Dissanayake, Surangika Ranathunga
CISIS5
2017 Automatic Identification of Errors in Multi-step Answers to Algebra Questions
abstract
This paper presents a system that automatically identifies errors made by students in answering algebra questions that require multiple steps. The types of algebra questions we consider include linear equations with fractions and quadratic equations. We have already developed a system that is capable of grading multi-step answers to the aforementioned two types of questions and awarding full/partial credit according to a marking scheme. The error identification module works on top of this previous system. It was evaluated using data from two sources: government schools and a tuition class in Sri Lanka. The mistakes identified by the system were compared against feedback by two independent teachers. The results showed that the system identified the student mistakes with more than 85% accuracy for both types of questions.
Buddhiprabha Erabadda, Surangika Ranathunga, Gihan Dias
ICALT2
2017 Assessment and Error Identification of Answers to Mathematical Word Problems
abstract
Mathematical word problems can be broadly divided into two categories as numerical word problems and algebraic word problems. These can be further categorised according to the domain, such as interest calculation and mensuration. Although most of the popular Mathematics examinations contain word problems, there are slight differences in the syllabi. The existing research has produced solutions for some categories of the word type problems. However, these solutions cannot be used for other types of word problems nor can these systems be used in the context of other examinations where there are differences in the grading schemes. We introduce a system that can be easily used to assess answers to both numerical and algebraic type word problems and automatically identifies the exact errors (if any) made by students by using a (teacher provided) marking rubric. The system is modularized and can be extended to support different types of word problems. If the answer contains a short textual phrase along with the numerical or algebraic expression, it is also evaluated in order to check whether the student has actually understood the question.
J. C. S. Kadupitiya, Surangika Ranathunga, Gihan Dias
ICALT2
2017 Automatic Assessment of Student Answers Consisting of Venn and Euler Diagrams
abstract
Venn and Euler diagrams are well-defined mathematical diagram types, which are the major representation methods of Set Theory. Venn and Euler diagrams are part of major Mathematics examinations in secondary education such as London Ordinary Level and SAT. Although computer assessment of different diagram types has been addressed, no such research has been done for Venn and Euler diagrams. In this research, we present a system capable of automatically assessing student answers consisting of Venn and Euler diagrams. The student answer is compared against a model answer, and marks are allocated according to a marking rubric.
Diunuge Buddhika Wijesinghe, J. C. S. Kadupitiya, Surangika Ranathunga, Gihan Dias
ICALT3
2017 Automatic Assessment of Student Answers for Geometric Construction Questions
abstract
In this paper, we present a system that evaluates answers given by students for geometric construction questions and awards full/partial marks according to a marking rubric. We focus on geometric construction questions found in high school mathematics, where students only use a straightedge and compass. The assessment system is automatic, and requires teacher involvement only at question set up time.
Buddhima S. Wijeweera, Gihan Dias, Surangika Ranathunga
ICALT3
2016 Named-Entity-Recognition (NER) for Tamil Language Using Margin-Infused Relaxed Algorithm (MIRA)
Pranavan Theivendiram, Megala Uthayakumar, Nilusija Nadarasamoorthy, Mokanarangan Thayaparan, Sanath Jayasena, Gihan Dias, Surangika Ranathunga
CICLing (1)7
2016 Computer Aided Evaluation of Multi-Step Answers to Algebra Questions
abstract
This paper presents a system that automatically assesses multi-step answers to algebra questions. The system requires teacher involvement only during the question set-up stage. Two types of algebra questions are currently supported: questions with linear equations containing fractions, and questions with quadratic equations. The system evaluates each step of a student's answer and awards full/partial marks according to a marking scheme. The system was evaluated for its performance using a set of student answer scripts from a government school in Sri Lanka and also by undergraduate students. The system accuracy was over 95.4%, and over 97.5%, respectively for the aforementioned data sets.
Buddhiprabha Erabadda, Surangika Ranathunga, Gihan Dias
ICALT2
2016 An Episode-based Approach to Identify Website User Access Patterns
abstract
Mining web access log data is a popular technique to identify frequent access patterns of website users. There are many mining techniques such as clustering, sequential pattern mining and association rule mining to identify these frequent access patterns. Each can find interesting access patterns and group the users, but they cannot identify the slight differences between accesses patterns included in individual clusters. But in reality these could refer to important information about attacks. This paper introduces a methodology to identify these access patterns at a much lower level than what is provided by traditional clustering techniques, such as nearest neighbour based techniques and classification techniques. This technique makes use of the concept of episodes to represent web sessions. These episodes are expressed in the form of regular expressions. To the best of our knowledge, this is the first time to apply the concept of regular expressions to identify user access patterns in web server log data. In addition to identifying frequent patterns, we demonstrate that this technique is able to identify access patterns that occur rarely, which would have been simply treated as noise in traditional clustering mechanisms.
Madhuka Udantha, Surangika Ranathunga, Gihan Dias
ICPRAM2
2016 Tamil Morphological Analyzer Using Support Vector Machines
Mokanarangan Thayaparan, Pranavan Theivendiram, Megala Uthayakumar, Nilusija Nadarasamoorthy, Gihan Dias, Sanath Jayasena, Surangika Ranathunga
NLDB7
2016 Domain-Specific Term Extraction for Concept Identification in Ontology Construction
abstract
An ontology is a formal and explicit specification of a shared conceptualization. Manual construction of domain ontology does not adequately satisfy requirements of new applications, because they need a more dynamic ontology and the possibility to manage a considerable quantity of concepts that humans cannot achieve alone. Researchers have discussed ontology learning as a solution to overcome issues related to the manual construction of ontology. Ontology learning is either an automatic or semi-automatic process to apply methods for building ontology from scratch, or enriching or adapting an existing ontology. This research focuses on improving the process of term extraction for identifying concepts in ontology learning. Available approaches for term extraction process are limited in various ways. These limitations include: (1) obtaining domain-specific terms from a domain expert as seed words without automatically discovering them from the corpus, and (2) unsuitable usage of corpora in discovering domain-specific terms for multiple domains. Our study uses linguistic analysis and statistical calculations to extract domain-specific simple and complex terms to overcome this first limitation. To eliminate the second limitation, we use multiple contrastive corpora that reduce the biasness in using a single contrastive corpus. Evaluations show that our system is better at extracting terms when compared with the previous research that used the same corpora.
Kiruparan Balachandran, Surangika Ranathunga
WI2
2015 Handling Agent Perception in Heterogeneous Distributed Systems: A Policy-Based Approach
Stephen Cranefield, Surangika Ranathunga
COORDINATION2
2015 VISIRI - Distributed Complex Event Processing System for Handling Large Number of Queries
Malinda Kumarasinghe, Geeth Tharanga, Lasitha Weerasinghe, Ujitha Wickramarathna, Surangika Ranathunga
COORDINATION5