VLDB 2026 Research / reviewers in the wild / expert
Malsha V. Perera
dblp:270/4731
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-1976-4647ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating Social Biases in Multimodal LLMsabstractWith the rapid advancement of Multimodal Large Language Models (MLLMs) and their ability to integrate multimodal inputs, these models are increasingly being applied to real-world tasks. However, alongside their impressive capabilities, MLLMs often exhibit undesirable characteristics, such as social biases. In this study, we conduct a comprehensive evaluation of bias in MLLMs concerning gender, race, and age attributes. To achieve this, we design a set of visual-question-answering (VQA)-based queries that prompt the models to perform attribute estimation given a face image. We assess these models using class-wise accuracies and bias-related metrics, revealing that while gender biases are relatively minimal, significant biases persist in race and age estimations. Our findings highlight the need for further research to mitigate these biases before deploying MLLMs in real-world applications. Malsha V. Perera, Kartik Narayan, Vishal M. Patel |
FG | 1 |
| 2025 | Debias-DPO: Debiasing Diffusion-based Face Image Generation with Direct Preference OptimizationabstractThe exceptional ability of diffusion-based generative models to produce high-quality images has led to their widespread adoption many in real-world applications. However, despite their impressive capabilities, these models often exhibit undesirable characteristics, such as social biases related to attributes like gender and race, which can have significant negative societal implications. In this study, we propose a Direct Preference Optimization (DPO)-based debiasing algorithm to effectively mitigate these biases in diffusion models used for face generation. We achieve this by giving greater preference to generating underrepresented classes, thereby achieving a balanced distribution in terms of the considered social attribute. Extensive experiments demonstrate that our method successfully reduces social biases in diffusion-based face generation models across multiple datasets while outperforming state-of-the-art debiasing techniques. Malsha V. Perera, Vishal M. Patel |
IJCB | 1 |
| 2025 | Frame by Familiar Frame: Understanding Replication in Video Diffusion ModelsabstractBuilding on the momentum of image generation diffusion models, there is an increasing interest in video-based diffusion models. However, video generation poses greater challenges due to its higher-dimensional nature, the scarcity of training data, and the complex spatiotemporal relationships involved. Image generation models, due to their extensive data requirements, have already strained computational resources to their limits. There have been instances of these models reproducing elements from the training samples, leading to concerns and even legal disputes over sample replication. Video diffusion models, which operate with even more constrained datasets and are tasked with generating both spatial and temporal content, may be more prone to replicating samples from their training sets. Compounding the issue, these models are often evaluated using metrics that inadvertently reward replication. In our paper, we present a systematic investigation into the phenomenon of sample replication in video diffusion models. We scrutinize various recent diffusion models for video synthesis, assessing their tendency to replicate spatial and temporal content in both unconditional and conditional generation scenarios. Our study identifies strategies that are less likely to lead to replication. Furthermore, we propose new evaluation strategies that take replication into account, offering a more accurate measure of a model's ability to generate the original content. Aimon Rahman, Malsha V. Perera, Vishal M. Patel |
WACV | 2 |
| 2023 | Analyzing Bias in Diffusion-based Face Generation ModelsabstractDiffusion models are becoming increasingly popular in synthetic data generation and image editing applications. However, these models can amplify existing biases and propagate them to downstream applications. Therefore, it is crucial to understand the sources of bias in their outputs. In this paper, we investigate the presence of bias in diffusion-based face generation models with respect to attributes such as gender, race, and age. Moreover, we examine how dataset size affects the attribute composition and perceptual quality of both diffusion and Generative Adversarial Network (GAN) based face generation models across various attribute classes. Our findings suggest that diffusion models tend to worsen distribution bias in the training data for various attributes, which is heavily influenced by the size of the dataset. Conversely, GAN models trained on balanced datasets with a larger number of samples show less bias across different attributes. Malsha V. Perera, Vishal M. Patel |
IJCB | 1 |
| 2023 | SAR Despeckling Using a Denoising Diffusion Probabilistic ModelabstractSpeckle is a type of multiplicative noise that affects all coherent imaging modalities including Synthetic Aperture Radar (SAR) images. The presence of speckle degrades the image quality and can adversely affect the performance of SAR image applications such as automatic target recognition and change detection. Thus, SAR despeckling is an important problem in remote sensing. In this paper, we introduce SAR-DDPM, a denoising diffusion probabilistic model for SAR despeckling. The proposed method employs a Markov chain that transforms clean images to white Gaussian noise by successively adding random noise. The despeckled image is obtained through a reverse process that predicts the added noise iteratively, using a noise predictor conditioned on the speckled image. Additionally, we propose a new inference strategy based on cycle spinning to improve the despeckling performance. Our experiments on both synthetic and real SAR images demonstrate that the proposed method leads to significant improvements in both quantitative and qualitative results over the state-of-the-art despeckling methods. The code is available at: https://github.com/malshaV/SAR_DDPM. Malsha V. Perera, Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M. Patel |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | SAR Despeckling Using Overcomplete Convolutional NetworksabstractSynthetic Aperture Radar (SAR) despeckling is an important problem in remote sensing as speckle degrades SAR images, affecting downstream tasks like detection and segmentation. Recent studies show that convolutional neural networks (CNNs) outperform classical despeckling methods. Traditional CNNs try to increase the receptive field size as the network goes deeper, thus extracting global features. However, speckle is relatively small, and increasing receptive field does not help in extracting speckle features. This study employs an overcomplete CNN architecture to focus on learning low-level features by restricting the receptive field. The proposed network consists of an overcomplete branch to focus on the local structures and an undercomplete branch that focuses on the global structures. We show that the proposed network improves despeckling performance compared to recent despeckling methods on synthetic and real SAR images. Our code is available at: https://github.com/malshaV/sar_overcomplete Malsha V. Perera, Wele Gedara Chaminda Bandara, Jeya Maria Jose Valanarasu, Vishal M. Patel |
IGARSS | 1 |
| 2022 | Transformer-Based SAR Image DespecklingabstractSynthetic Aperture Radar (SAR) images are usually degraded by a multiplicative noise known as speckle which makes processing and interpretation of SAR images difficult. In this paper, we introduce a transformer-based network for SAR image despeckling. The proposed despeckling network comprises of a transformer-based encoder which allows the network to learn global dependencies between different image regions - aiding in better despeckling. The network is trained end-to-end with synthetically generated speckled images using a composite loss function. Experiments show that the proposed method achieves significant improvements over traditional and convolutional neural network-based despeckling methods on both synthetic and real SAR images. Our code is available at: https://github.com/malshaV/sar_transformer Malsha V. Perera, Wele Gedara Chaminda Bandara, Jeya Maria Jose Valanarasu, Vishal M. Patel |
IGARSS | 1 |
| 2021 | A Joint Convolutional and Spatial Quad-Directional LSTM Network for Phase UnwrappingabstractPhase unwrapping is a classical ill-posed problem which aims to recover the true phase from wrapped phase. In this paper, we introduce a novel Convolutional Neural Network (CNN) that incorporates a Spatial Quad-Directional Long Short Term Memory (SQD-LSTM) for phase unwrapping, by formulating it as a regression problem. Incorporating SQD-LSTM can circumvent the typical CNNs’ inherent difficulty of learning global spatial dependencies which are vital when recovering the true phase. Furthermore, we employ a problem specific composite loss function to train this network. The proposed network is found to be performing better than the existing methods under severe noise conditions (Normalized Root Mean Square Error of 1.3% at SNR = 0 dB) while spending a significantly less computational time (0.054s). The network also does not require a large scale dataset during training, thus making it ideal for applications with limited data that require fast and accurate phase unwrapping. Malsha V. Perera, Ashwin De Silva |
ICASSP | 1 |
| 2020 | Real-Time Hand Gesture Recognition Using Temporal Muscle Activation Maps of Multi-Channel Semg SignalsabstractAccurate and real-time hand gesture recognition is essential for controlling advanced hand prostheses. Surface Electromyography (sEMG) signals obtained from the forearm are widely used for this purpose. Here, we introduce a novel hand gesture representation called Temporal Muscle Activation (TMA) maps which captures information about the activation patterns of muscles in the forearm. Based on these maps, we propose an algorithm that can recognize hand gestures in real-time using a Convolution Neural Network. The algorithm was tested on 8 healthy subjects with sEMG signals acquired from 8 electrodes placed along the circumference of the forearm. The average classification accuracy of the proposed method was 94%, which is comparable to state-of-the-art methods. The average computation time of a prediction was 5.5ms, making the algorithm ideal for the real-time gesture recognition applications. Ashwin De Silva, Malsha V. Perera, Kithmin Wickremasinghe, Asma M. Naim, Thilina Dulantha Lalitharatne, Simon Lind Kappel |
ICASSP | 2 |
| 2020 | Low-cost Active Dry-Contact Surface EMG Sensor for Bionic ArmsabstractSurface electromyography (sEMG) is a popular bio-signal used for controlling prostheses and finger gesture recognition mechanisms. Myoelectric prostheses are costly, and most commercially available sEMG acquisition systems are not suitable for real-time gesture recognition. In this paper, a method of acquiring sEMG signals using novel low-cost, active, dry-contact, flexible sensors has been proposed. Since the active sEMG sensor was developed to be used along with a bionic arm, the sensor was tested for its ability to acquire sEMG signals that could be used for real-time classification of five selected gestures. In a study of 4 subjects, the average classification accuracy for real-time gesture classification using the active sEMG sensor system was 85%. The common-mode rejection ratio of the sensor was measured to 59 dB, and thus the sensor's performance was not substantially limited by its active circuitry. The proposed sensors can be interfaced with a variety of amplifiers to perform fully wearable sEMG acquisition. This satisfies the need for a low-cost sEMG acquisition system for prostheses. Asma M. Naim, Kithmin Wickremasinghe, Ashwin De Silva, Malsha V. Perera, Thilina Dulantha Lalitharatne, Simon Lind Kappel |
SMC | 4 |