Nisansa de Silva

dblp:180/9385 · also Nisansa Dilushan de Silva · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-5361-4810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 2 first-author · 22 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adapter-Based Multi-Document Summarisation: Opinion Summarisation Use Case
Kushan Hewapathirana, Nisansa de Silva, C. D. Athuralya
ICAART (4)2
2026 SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas
LREC2
2026 JavaBackports: A Dataset for Benchmarking Automated Backporting in Java
abstract
Manually backporting critical patches to long-term support versions is both error-prone and often overlooked, resulting in substantial security risks. Progress in this area is constrained by the absence of datasets that capture the semantic complexities across versions, inherent to backporting in large Java ecosystems. To address this gap, we present JavaBackports, a curated dataset of 491 real-world backport instances, systematically selected and manually validated from more than 11,000 candidate patches in fifteen widely used open-source Java projects: Druid, Elasticsearch, Hadoop, Kafka, among others and four major JDK versions (jdk11, jdk17, jdk21, jdk25). To assess the utility of JavaBackports, we conduct preliminary experiments to evaluate the effectiveness of the state-of-the-art Large Language Models (LLMs) in zero-shot automatic patch backporting. The results indicate that current LLMs struggle with backporting tasks, particularly when the required changes involve non-trivial logical or structural modifications. These findings demonstrate both the difficulty of the problem and the potential of JavaBackports to stimulate new research directions in automated software maintenance and repair.
Kaushal Kahapola, Sharada Galappaththi, Dinith Ranasinghe, Ridwan Salihin Shariffdeen, Nisansa de Silva, Srinath Perera, Sandareka Wickramanayake
MSR5
2026 From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks
Mineth Weerasinghe, Himindu Kularathne, Methmini Madhushika, Danuka Lakshan, Nisansa de Silva, Adeesha Wijayasiri, Srinath Perera
WorldCIST (4)5
2026 An emotion aware context adaptive machine learning approach for detecting hate speech in social media
abstract
Abstract Detecting hate speech is challenging because language constantly evolves, often carrying subtle emotional undertones that can be difficult to interpret. While many existing approaches rely on semantic analysis, they frequently miss these emotional cues and struggle to adapt to different contexts. This paper introduces an enhanced Dual Contrastive Learning (DCL) framework, integrating emotion profiling with fine-tuned language models (BERTweet, RoBERTa, TimeLMs) to address these gaps. By using SHAP (SHapley Additive exPlanations) for interpretability, our approach enhances both detection accuracy and explainability. Experiments on SemEval-2019, Davidson, and a Unified Twitter corpus show that our EmotionDCL models outperform baseline methods, with TimeLMs achieving 81.20% accuracy on SemEval and 94.04% on Davidson. Ablation studies confirm the synergy between contrastive learning and emotion integration, highlighting anger, fear, and anticipation as dominant emotional drivers of hate speech. These findings show the importance of emotion-aware, context-adaptive detection systems and the need for strategies to address class imbalance and linguistic evolution.
Krishan Chavinda, Pasan Kalansooriya, Thushalya Weerasuriya, Nisansa de Silva, Uthayasanker Thayasivam, Kogul Srikandabala, Kirishnni Prabagar, Damminda Alahakoon
Neural Comput. Appl.4
2025 Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
Kavindu Warnakulasuriya, Prabhash Dissanayake, Navindu De Silva, Stephen Cranefield, Bastin Tony Roy Savarimuthu, Surangika Ranathunga, Nisansa de Silva
COINE7
2025 Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics
abstract
Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from webmined corpora.Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique.However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs.This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view.We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics.The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs.We publicly release the source code and the curated datasets 1 .
Aloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake, Surangika Ranathunga
EMNLP2
2025 Automatic Analysis of App Reviews Using LLMs
Sadeep Gunathilaka, Nisansa de Silva
ICAART (2)2
2025 DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction
Vishal Thenuwara, Nisansa de Silva
ICCCI (1)2
2025 Domain Adaptation for Multi-document Summarisation: A Case Study in the Medical Research Domain
Kushan Hewapathirana, Nisansa de Silva, C. D. Athuralya, Piumi Kandanaarachchi
PACLIC2
2025 SiDiaC: Sinhala Diachronic Corpus
Nevidu Jayatilleke, Nisansa de Silva
PACLIC2
2025 Fine-tuning an LLM to Generate Lore Coherent Encounters for Dungeons and Dragons
Aravinth Sivaganeshan, Nisansa de Silva, Akila Peiris
PACLIC2
2025 ScheduleMe: Multi-Agent Calendar Assistant
Oshadha Wijerathne, Amandi Nimasha, Dushan Fernando, Nisansa de Silva, Srinath Perera
PACLIC4
2024 Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
abstract
Surangika Ranathunga, Nisansa De Silva, Velayuthan Menan, Aloka Fernando, Charitha Rathnayake. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Surangika Ranathunga, Nisansa de Silva, Menan Velayuthan, Aloka Fernando, Charitha Rathnayake
EACL (1)2
2024 A Multi-Stage Approach to Image Consistency in Zero-Shot Character Art Generation for the D&D Domain
Gayashan Weerasundara, Nisansa de Silva
ICAART (3)2
2024 Abstract Generation with Hybrid Model Supported by Relevance Matrix
Dushan Kumarasinghe, Nisansa de Silva
ICCCI (1)2
2024 Enhanced Aspect-Based Sentiment Analysis with Integrated Category Extraction for Instruct-DeBERTa
Dineth Jayakody, Koshila Isuranda, A V. A Malkith, Nisansa de Silva, Sachintha Rajith Ponnamperuma, Gammana Guruge Nadeesha Sandamali, Kushan Sudheera Kalupahana Liyanage, Kashnika Gimhani Sarathchandra
PACLIC4
2023 Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
Kasun Wickramasinghe, Nisansa de Silva
PACLIC2
2022 Learning Sentence Embeddings In The Legal Domain with Low Resource Settings
Sahan Jayasinghe, Lakith Rambukkanage, Ashan Silva, Nisansa de Silva, Shehan Perera, Madhavi Perera
PACLIC4
2022 Synthesis and Evaluation of a Domain-specific Large Data Set for Dungeons & Dragons
Akila Peiris, Nisansa de Silva
PACLIC2
2022 Sinhala Sentence Embedding: A Two-Tiered Structure for Low-Resource Languages
Gihan Weeraprameshwara, Vihanga Jayawickrama, Nisansa de Silva, Yudhanjaya Wijeratne
PACLIC3
2022 Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
abstract
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi
Trans. Assoc. Comput. Linguistics39
2021 Sigmalaw PBSA - A Deep Learning Model for Aspect-Based Sentiment Analysis for the Legal Domain
Isanka Rajapaksha, Chanika Ruchini Mudalige, Dilini Karunarathna, Nisansa de Silva, Amal Perera, Gathika Ratnayaka
DEXA (1)4
2021 Semantic Oppositeness Assisted Deep Contextual Modeling for Automatic Rumor Detection in Social Networks
abstract
Social networks face a major challenge in the form of rumors and fake news, due to their intrinsic nature of connecting users to millions of others, and of giving any individual the power to post anything.Given the rapid, widespread dissemination of information in social networks, manually detecting suspicious news is sub-optimal.Thus, research on automatic rumor detection has become a necessity.Previous works in the domain have utilized the reply relations between posts, as well as the semantic similarity between the main post and its context, consisting of replies, in order to obtain state-of-the-art performance.In this work, we demonstrate that semantic oppositeness can improve the performance on the task of rumor detection.We show that semantic oppositeness captures elements of discord, which are not properly covered by previous efforts, which only utilize semantic similarity or reply structure.Our proposed model learns both explicit and implicit relations between the main tweet and its replies, by utilizing both semantic similarity and semantic oppositeness.Both of these employ the self-attention mechanism in neural text modeling, with semantic oppositeness utilizing word-level self-attention, and with semantic similarity utilizing post-level self-attention.We show, with extensive experiments on recent data sets for this problem, that our proposed model achieves state-of-theart performance.Further, we show that our model is more resistant to the variances in performance introduced by randomness.
Nisansa de Silva, Dejing Dou
EACL1
2021 Identifying Legal Party Members from Legal Opinion Documents using Natural Language Processing
abstract
Law and order is a field that can highly benefit from the contribution of Natural Language Processing (NLP) to its betterment. An area in which NLP can be of immense help is, information retrieval from legal documents which function as legal databases. The extraction of legal parties from the aforementioned legal documents can be identified as a task of high importance since it has a significant impact on the proceedings of contemporary legal cases. This study proposes a novel deep learning methodology which can be effectively used to find a solution to the problem of identifying legal party members in legal documents. In addition to that, in this paper, we introduce a novel data set which is annotated with legal party information by an expert in the legal domain. The deep learning model proposed in this study provides a benchmark for the legal party identification task on this data set. Evaluations for the solution presented in the paper show that our system has 90.89% precision and 91.69% recall for an unseen paragraph from a legal document, thus conforming the success of our attempt.
Melonie de Almeida, Chamodi Samarawickrama, Nisansa de Silva, Gathika Ratnayaka, Amal Perera
iiWAS3
2020 Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial Regularization
abstract
Network representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods.
Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001
COLING2
2020 Effective Approach to Develop a Sentiment Annotator For Legal Domain in a Low Resource Setting
Gathika Ratnayaka, Nisansa de Silva, Amal Perera, Ramesh Pathirana
PACLIC2
2019 Semantic Oppositeness Embedding Using an Autoencoder-Based Learning Model
Nisansa de Silva, Dejing Dou
DEXA (1)1
2018 Concept and Attention-Based CNN for Question Retrieval in Multi-View Learning
abstract
Question retrieval, which aims to find similar versions of a given question, is playing a pivotal role in various question answering (QA) systems. This task is quite challenging, mainly in regard to five aspects: synonymy, polysemy, word order, question length, and data sparsity. In this article, we propose a unified framework to simultaneously handle these five problems. We use the word combined with corresponding concept information to handle the synonymy problem and the polysemous problem. Concept embedding and word embedding are learned at the same time from both the context-dependent and context-independent views. To handle the word-order problem, we propose a high-level feature-embedded convolutional semantic model to learn question embedding by inputting concept embedding and word embedding. Due to the fact that the lengths of some questions are long, we propose a value-based convolutional attentional method to enhance the proposed high-level feature-embedded convolutional semantic model in learning the key parts of the question and the answer. The proposed high-level feature-embedded convolutional semantic model nicely represents the hierarchical structures of word information and concept information in sentences with their layer-by-layer convolution and pooling. Finally, to resolve data sparsity, we propose using the multi-view learning method to train the attention-based convolutional semantic model on question–answer pairs. To the best of our knowledge, we are the first to propose simultaneously handling the above five problems in question retrieval using one framework. Experiments on three real question-answering datasets show that the proposed framework significantly outperforms the state-of-the-art solutions.
Pengwei Wang 0004, Lei Ji 0001, Jun Yan 0001, Dejing Dou, Nisansa de Silva
ACM Trans. Intell. Syst. Technol.5
2014 Building a WordNet for Sinhala
abstract
Indeewari Wijesiri, Malaka Gallage, Buddhika Gunathilaka, Madhuranga Lakjeewa, Daya Wimalasuriya, Gihan Dias, Rohini Paranavithana, Nisansa de Silva. Proceedings of the Seventh Global Wordnet Conference. 2014.
Indeewari Wijesiri, Malaka Gallage, Buddhika Gunathilaka, Madhuranga Lakjeewa, Daya C. Wimalasuriya, Gihan Dias, Rohini Paranavithana, Nisansa de Silva
GWC8