VLDB 2026 Research / reviewers in the wild / expert
Nisansa de Silva
dblp:180/9385 · also Nisansa Dilushan de Silva
· DBLP profile ↗
30ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-5361-4810ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 22 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adapter-Based Multi-Document Summarisation: Opinion Summarisation Use Case
Kushan Hewapathirana, Nisansa de Silva, C. D. Athuralya |
ICAART (4) | 2 |
| 2026 | SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas |
LREC | 2 |
| 2026 | JavaBackports: A Dataset for Benchmarking Automated Backporting in JavaabstractManually backporting critical patches to long-term support versions is both error-prone and often overlooked, resulting in substantial security risks. Progress in this area is constrained by the absence of datasets that capture the semantic complexities across versions, inherent to backporting in large Java ecosystems. To address this gap, we present JavaBackports, a curated dataset of 491 real-world backport instances, systematically selected and manually validated from more than 11,000 candidate patches in fifteen widely used open-source Java projects: Druid, Elasticsearch, Hadoop, Kafka, among others and four major JDK versions (jdk11, jdk17, jdk21, jdk25). To assess the utility of JavaBackports, we conduct preliminary experiments to evaluate the effectiveness of the state-of-the-art Large Language Models (LLMs) in zero-shot automatic patch backporting. The results indicate that current LLMs struggle with backporting tasks, particularly when the required changes involve non-trivial logical or structural modifications. These findings demonstrate both the difficulty of the problem and the potential of JavaBackports to stimulate new research directions in automated software maintenance and repair. Kaushal Kahapola, Sharada Galappaththi, Dinith Ranasinghe, Ridwan Salihin Shariffdeen, Nisansa de Silva, Srinath Perera, Sandareka Wickramanayake |
MSR | 5 |
| 2026 | From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks
Mineth Weerasinghe, Himindu Kularathne, Methmini Madhushika, Danuka Lakshan, Nisansa de Silva, Adeesha Wijayasiri, Srinath Perera |
WorldCIST (4) | 5 |
| 2026 | An emotion aware context adaptive machine learning approach for detecting hate speech in social mediaabstractAbstract Detecting hate speech is challenging because language constantly evolves, often carrying subtle emotional undertones that can be difficult to interpret. While many existing approaches rely on semantic analysis, they frequently miss these emotional cues and struggle to adapt to different contexts. This paper introduces an enhanced Dual Contrastive Learning (DCL) framework, integrating emotion profiling with fine-tuned language models (BERTweet, RoBERTa, TimeLMs) to address these gaps. By using SHAP (SHapley Additive exPlanations) for interpretability, our approach enhances both detection accuracy and explainability. Experiments on SemEval-2019, Davidson, and a Unified Twitter corpus show that our EmotionDCL models outperform baseline methods, with TimeLMs achieving 81.20% accuracy on SemEval and 94.04% on Davidson. Ablation studies confirm the synergy between contrastive learning and emotion integration, highlighting anger, fear, and anticipation as dominant emotional drivers of hate speech. These findings show the importance of emotion-aware, context-adaptive detection systems and the need for strategies to address class imbalance and linguistic evolution. Krishan Chavinda, Pasan Kalansooriya, Thushalya Weerasuriya, Nisansa de Silva, Uthayasanker Thayasivam, Kogul Srikandabala, Kirishnni Prabagar, Damminda Alahakoon |
Neural Comput. Appl. | 4 |
| 2025 | Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
Kavindu Warnakulasuriya, Prabhash Dissanayake, Navindu De Silva, Stephen Cranefield, Bastin Tony Roy Savarimuthu, Surangika Ranathunga, Nisansa de Silva |
COINE | 7 |
| 2025 | Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing HeuristicsabstractParallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from webmined corpora.Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique.However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs.This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view.We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics.The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs.We publicly release the source code and the curated datasets 1 . Aloka Fernando, Nisansa de Silva, Menan Velayuthan, Charitha Rathnayake, Surangika Ranathunga |
EMNLP | 2 |
| 2025 | Automatic Analysis of App Reviews Using LLMs
Sadeep Gunathilaka, Nisansa de Silva |
ICAART (2) | 2 |
| 2025 | DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction
Vishal Thenuwara, Nisansa de Silva |
ICCCI (1) | 2 |
| 2025 | Domain Adaptation for Multi-document Summarisation: A Case Study in the Medical Research Domain
Kushan Hewapathirana, Nisansa de Silva, C. D. Athuralya, Piumi Kandanaarachchi |
PACLIC | 2 |
| 2025 | SiDiaC: Sinhala Diachronic Corpus
Nevidu Jayatilleke, Nisansa de Silva |
PACLIC | 2 |
| 2025 | Fine-tuning an LLM to Generate Lore Coherent Encounters for Dungeons and Dragons
Aravinth Sivaganeshan, Nisansa de Silva, Akila Peiris |
PACLIC | 2 |
| 2025 | ScheduleMe: Multi-Agent Calendar Assistant
Oshadha Wijerathne, Amandi Nimasha, Dushan Fernando, Nisansa de Silva, Srinath Perera |
PACLIC | 4 |
| 2024 | Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel CorporaabstractSurangika Ranathunga, Nisansa De Silva, Velayuthan Menan, Aloka Fernando, Charitha Rathnayake. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Surangika Ranathunga, Nisansa de Silva, Menan Velayuthan, Aloka Fernando, Charitha Rathnayake |
EACL (1) | 2 |
| 2024 | A Multi-Stage Approach to Image Consistency in Zero-Shot Character Art Generation for the D&D Domain
Gayashan Weerasundara, Nisansa de Silva |
ICAART (3) | 2 |
| 2024 | Abstract Generation with Hybrid Model Supported by Relevance Matrix
Dushan Kumarasinghe, Nisansa de Silva |
ICCCI (1) | 2 |
| 2024 | Enhanced Aspect-Based Sentiment Analysis with Integrated Category Extraction for Instruct-DeBERTa
Dineth Jayakody, Koshila Isuranda, A V. A Malkith, Nisansa de Silva, Sachintha Rajith Ponnamperuma, Gammana Guruge Nadeesha Sandamali, Kushan Sudheera Kalupahana Liyanage, Kashnika Gimhani Sarathchandra |
PACLIC | 4 |
| 2023 | Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
Kasun Wickramasinghe, Nisansa de Silva |
PACLIC | 2 |
| 2022 | Learning Sentence Embeddings In The Legal Domain with Low Resource Settings
Sahan Jayasinghe, Lakith Rambukkanage, Ashan Silva, Nisansa de Silva, Shehan Perera, Madhavi Perera |
PACLIC | 4 |
| 2022 | Synthesis and Evaluation of a Domain-specific Large Data Set for Dungeons & Dragons
Akila Peiris, Nisansa de Silva |
PACLIC | 2 |
| 2022 | Sinhala Sentence Embedding: A Two-Tiered Structure for Low-Resource Languages
Gihan Weeraprameshwara, Vihanga Jayawickrama, Nisansa de Silva, Yudhanjaya Wijeratne |
PACLIC | 3 |
| 2022 | Quality at a Glance: An Audit of Web-Crawled Multilingual DatasetsabstractAbstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases. Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi |
Trans. Assoc. Comput. Linguistics | 39 |
| 2021 | Sigmalaw PBSA - A Deep Learning Model for Aspect-Based Sentiment Analysis for the Legal Domain
Isanka Rajapaksha, Chanika Ruchini Mudalige, Dilini Karunarathna, Nisansa de Silva, Amal Perera, Gathika Ratnayaka |
DEXA (1) | 4 |
| 2021 | Semantic Oppositeness Assisted Deep Contextual Modeling for Automatic Rumor Detection in Social NetworksabstractSocial networks face a major challenge in the form of rumors and fake news, due to their intrinsic nature of connecting users to millions of others, and of giving any individual the power to post anything.Given the rapid, widespread dissemination of information in social networks, manually detecting suspicious news is sub-optimal.Thus, research on automatic rumor detection has become a necessity.Previous works in the domain have utilized the reply relations between posts, as well as the semantic similarity between the main post and its context, consisting of replies, in order to obtain state-of-the-art performance.In this work, we demonstrate that semantic oppositeness can improve the performance on the task of rumor detection.We show that semantic oppositeness captures elements of discord, which are not properly covered by previous efforts, which only utilize semantic similarity or reply structure.Our proposed model learns both explicit and implicit relations between the main tweet and its replies, by utilizing both semantic similarity and semantic oppositeness.Both of these employ the self-attention mechanism in neural text modeling, with semantic oppositeness utilizing word-level self-attention, and with semantic similarity utilizing post-level self-attention.We show, with extensive experiments on recent data sets for this problem, that our proposed model achieves state-of-theart performance.Further, we show that our model is more resistant to the variances in performance introduced by randomness. Nisansa de Silva, Dejing Dou |
EACL | 1 |
| 2021 | Identifying Legal Party Members from Legal Opinion Documents using Natural Language ProcessingabstractLaw and order is a field that can highly benefit from the contribution of Natural Language Processing (NLP) to its betterment. An area in which NLP can be of immense help is, information retrieval from legal documents which function as legal databases. The extraction of legal parties from the aforementioned legal documents can be identified as a task of high importance since it has a significant impact on the proceedings of contemporary legal cases. This study proposes a novel deep learning methodology which can be effectively used to find a solution to the problem of identifying legal party members in legal documents. In addition to that, in this paper, we introduce a novel data set which is annotated with legal party information by an expert in the legal domain. The deep learning model proposed in this study provides a benchmark for the legal party identification task on this data set. Evaluations for the solution presented in the paper show that our system has 90.89% precision and 91.69% recall for an unseen paragraph from a legal document, thus conforming the success of our attempt. Melonie de Almeida, Chamodi Samarawickrama, Nisansa de Silva, Gathika Ratnayaka, Amal Perera |
iiWAS | 3 |
| 2020 | Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial RegularizationabstractNetwork representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods. Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001 |
COLING | 2 |
| 2020 | Effective Approach to Develop a Sentiment Annotator For Legal Domain in a Low Resource Setting
Gathika Ratnayaka, Nisansa de Silva, Amal Perera, Ramesh Pathirana |
PACLIC | 2 |
| 2019 | Semantic Oppositeness Embedding Using an Autoencoder-Based Learning Model
Nisansa de Silva, Dejing Dou |
DEXA (1) | 1 |
| 2018 | Concept and Attention-Based CNN for Question Retrieval in Multi-View LearningabstractQuestion retrieval, which aims to find similar versions of a given question, is playing a pivotal role in various question answering (QA) systems. This task is quite challenging, mainly in regard to five aspects: synonymy, polysemy, word order, question length, and data sparsity. In this article, we propose a unified framework to simultaneously handle these five problems. We use the word combined with corresponding concept information to handle the synonymy problem and the polysemous problem. Concept embedding and word embedding are learned at the same time from both the context-dependent and context-independent views. To handle the word-order problem, we propose a high-level feature-embedded convolutional semantic model to learn question embedding by inputting concept embedding and word embedding. Due to the fact that the lengths of some questions are long, we propose a value-based convolutional attentional method to enhance the proposed high-level feature-embedded convolutional semantic model in learning the key parts of the question and the answer. The proposed high-level feature-embedded convolutional semantic model nicely represents the hierarchical structures of word information and concept information in sentences with their layer-by-layer convolution and pooling. Finally, to resolve data sparsity, we propose using the multi-view learning method to train the attention-based convolutional semantic model on question–answer pairs. To the best of our knowledge, we are the first to propose simultaneously handling the above five problems in question retrieval using one framework. Experiments on three real question-answering datasets show that the proposed framework significantly outperforms the state-of-the-art solutions. Pengwei Wang 0004, Lei Ji 0001, Jun Yan 0001, Dejing Dou, Nisansa de Silva |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | Building a WordNet for SinhalaabstractIndeewari Wijesiri, Malaka Gallage, Buddhika Gunathilaka, Madhuranga Lakjeewa, Daya Wimalasuriya, Gihan Dias, Rohini Paranavithana, Nisansa de Silva. Proceedings of the Seventh Global Wordnet Conference. 2014. Indeewari Wijesiri, Malaka Gallage, Buddhika Gunathilaka, Madhuranga Lakjeewa, Daya C. Wimalasuriya, Gihan Dias, Rohini Paranavithana, Nisansa de Silva |
GWC | 8 |