EDBT 2026 Demo / reviewers in the wild / expert
Suranga Seneviratne
dblp:126/2555
· DBLP profile ↗
55ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0002-5485-5595ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 1 first-author · 20 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrafficLLM: LLMs for improved open-set encrypted traffic analysisabstractEncrypted traffic has been known to be vulnerable to traffic analysis attacks that exploit the statistical features of encrypted traffic flows, such as packet sizes, timing, and direction, to infer information about the underlying content, which undermines the privacy guarantees of end-to-end encryption. Existing methods such as CNNs lack generalizability, requiring model changes based on the dataset. Furthermore, while state-of-the-art attacks leverage deep learning models to achieve high accuracy, most attacks work under the less realistic closed-set assumption, failing in the open-set setting. Deploying such attacks in practice requires addressing the open-set scenario, which allows the models to filter out target content from other background traffic. The effectiveness of open-set traffic classification largely relies on the model’s ability to generalize and accurately extract features from traffic traces, which are essentially sequential data. Concurrently, Large Language Models (LLM) are increasingly becoming popular in modeling sequential data beyond their typical applications in natural language processing. Inspired by this, our work introduces TrafficLLM, a novel traffic analysis attack method that leverages pre-trained LLMs, such as GPT-2 and LLaMA-2-7B, to extract features from network traffic traces with minimal fine-tuning. Using seven existing encrypted traffic datasets, we show that LLMs improve the open-set performance of traffic classification; for instance, our method, TrafficLLM outperforms ET-BERT and CNN-based approaches by 12.7 % and 13.7 % with GPT-2 feature extractor and 17.6 % and 21.5 % with LLaMA-2-7B feature extractor, respectively. Yasod Ginige, Bhanuka Silva, Thilini Dahanayaka, Suranga Seneviratne |
Comput. Networks | 4 |
| 2026 | Personalizing Federated Learning for Hierarchical Edge Networks With Non-IID DataabstractHierarchical Federated Learning (HFL) frameworks place edge servers between IoT devices and the cloud server to reduce communication costs and preserve privacy. In practice, however, HFL must handle hierarchical non-IID data across both device and edge levels. At the edge-level, heterogeneity arises because devices connected to the same edge server often share geographic or contextual similarities, giving each server its own optimization goal aligned with its region-specific data distribution rather than with a shared global objective. Existing HFL methods largely ignore this distinction, focusing on training a single global model that can obscure severe underperformance at the edge-level with underrepresented data. Since edge servers often act as operational units, poor performance at an edge implies degraded service quality, undermining system reliability and user trust. We propose Personalized Hierarchical Edge-enabled Federated Learning (PHE-FL), a novel method that produces personalized edge models by adaptively integrating edge- and cloud-level knowledge based on the data distribution of each edge, without incurring additional computational overhead or compromising client privacy. We deploy edge-specific test sets at each edge to ensure its unique data distribution is accurately reflected during evaluation. To the best of our knowledge, this is the first work to explicitly address hierarchical data heterogeneity in a 3-level HFL framework, both in terms of personalization and evaluation. Extensive experiments show that PHE-FL achieves up to 83% higher accuracy than existing edge-accommodated FL methods and maintains robust performance across edge-level non-IIDness, with reduced accuracy fluctuations compared to the state-of-the-art FedAvg with two levels (edge and cloud) aggregation. Omid Tavallaie, Shuaijun Chen, Kanchana Thilakarathna, Suranga Seneviratne, Adel Nadjaran Toosi, Albert Y. Zomaya |
IEEE Internet Things J. | 5 |
| 2026 | Device Type Classification Using WiFi Probe Requests: From Signals to InsightsabstractWiFi devices are ubiquitous in modern environments, from smartphones and laptops to IoT sensors and AR/VR headsets. Identifying device types/models within these populations enables crowd analysis, network optimization, and detection of unusual devices. Current identification methods struggle with MAC address randomization, require large training datasets, and perform poorly in real-world deployments. This paper introduces a device identification method based on Information Element (IE) attributes extracted from WiFi probe requests. We evaluate the approach using probe requests captured in the 2.4 GHz band. Evaluation across 70+ device types yields 99% precision, 98% recall, and 99% F1 score, exceeding deep learning approaches (92% F1 score) under similar training conditions. Our approach maintains accuracy despite MAC randomization and requires minimal training data. We demonstrate practical applicability through an operational dashboard tested in real-world scenarios for urban planning and network management. Case studies across diverse environments confirm the effectiveness of the method for operational use. Niruth Bogahawatta, Yasiru Senarath Karunanayaka, Suranga Seneviratne, Kanchana Thilakarathna, Rahat Masood, Salil S. Kanhere, Aruna Seneviratne, Albert Y. Zomaya |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | FeUDA-BERT: Federated URL Domain-Aware BERT for Phishing URL DetectionabstractPhishing URL classification is a critical component of cybersecurity. Although numerous phishing detection models based on machine learning and deep learning have been proposed, many face challenges related to generalisability and domain adaptation. This is largely due to the absence of representative training data and restricted threat intelligence sharing caused by data privacy and commercial concerns. To address these challenges, we introduce FeUDA-BERT, a federated, domain-aware BERT-based framework designed for privacy-preserving phishing URL detection. Our approach includes a novel URL domain-aware self-attention mechanism and a self-attention-based client selection method that tackles client data heterogeneity and enhances detection accuracy. We assess our model in a federated learning setup and show that each contribution — the domain-aware attention and the client selection strategy — independently boosts performance over baseline methods. Combined, they yield even greater improvements, achieving an average F1 score of 0.93. Using statistical analyses, we further validate that these performance gains are statistically significant, demonstrating the robustness and effectiveness of our proposed framework. Fariza Rashid, Ben Doyle, Suranga Seneviratne |
LCN | 3 |
| 2025 | Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly DetectionabstractThe proliferation of edge devices has dramatically increased the generation of multivariate time-series (MVTS) data, essential for applications from healthcare to smart cities. Such data streams, however, are vulnerable to anomalies that signal crucial problems like system failures or security incidents. Traditional MVTS anomaly detection methods, encompassing statistical and centralized machine learning approaches, struggle with the heterogeneity, variability, and privacy concerns of large-scale, distributed environments. In response, we introduce FedKO, a novel unsupervised Federated Learning framework that leverages the linear predictive capabilities of Koopman operator theory along with the dynamic adaptability of Reservoir Computing. This enables effective spatiotemporal processing and privacy-preserving for MVTS data. FedKO is formulated as a bi-level optimization problem, utilizing a specific federated algorithm to explore a shared Reservoir-Koopman model across diverse datasets. Such a model is then deployable on edge devices for efficient detection of anomalies in local MVTS streams. Experimental results across various datasets showcase FedKO’s superior performance against state-of-the-art methods in MVTS anomaly detection. Moreover, FedKO reduces up to 8x communication size and 2x memory usage, making it highly suitable for large-scale systems. Tung-Anh Nguyen, Han Shu, Suranga Seneviratne, Choong Seon Hong, Nguyen H. Tran |
SDM | 4 |
| 2025 | Detecting Content Rating Violations in Android Applications: A Vision-Language ApproachabstractDespite regulatory efforts to establish reliable content-rating guidelines for mobile apps, the process of assigning content ratings in the Google Play Store remains self-regulated by the app developers. There is no straightforward method of verifying developer-assigned content ratings manually due to the overwhelming scale or automatically due to the challenging problem of interpreting textual and visual data and correlating them with content ratings. We propose and evaluate a vision-language approach to predict the content ratings of mobile game applications and detect content rating violations, using a dataset of metadata of popular Android games.Our method achieves ∼6% better relative accuracy compared to the state-of-the-art CLIP-fine-tuned model in a multi-modal setting. Applying our classifier in the wild, we detected more than 70 possible cases of content rating violations, including nine instances with the ‘Teacher Approved’ badge. Additionally, our findings indicate that 34.5% of the apps identified by our classifier as violating content ratings were later removed from the Play Store. In contrast, the removal rate for correctly classified apps was only 27%. This discrepancy highlights the practical effectiveness of our classifier in identifying apps likely to be removed based on user complaints. Dishanika Denipitiyage, Bhanuka Silva, Suranga Seneviratne, Aruna Seneviratne, Sanjay Chawla |
TrustCom | 3 |
| 2025 | AutoPentester: An LLM Agent-based Framework for Automated PentestingabstractPenetration testing and vulnerability assessment are essential industry practices for safeguarding computer systems. As cyber threats grow in scale and complexity, the demand for pentesting has surged, surpassing the capacity of human professionals to meet it effectively. With advances in AI, particularly Large Language Models (LLMs), there have been attempts to automate the pentesting process. However, existing tools such as PentestGPT are still semi-manual, requiring significant professional human interaction to conduct pentests. To this end, we propose a novel LLM agent-based framework, AutoPentester, which automates the pentesting process. Given a target IP, AutoPentester automatically conducts pentesting steps using common security tools in an iterative process. It can dynamically generate attack strategies based on the tool outputs from the previous iteration, mimicking the human pentester approach. We evaluate AutoPentester using Hack The Box and custom-made VMs, comparing the results with the state-of-the-art PentestGPT. Results show that AutoPentester achieves a 27.0% better subtask completion rate and 39.5% more vulnerability coverage with fewer steps. Most importantly, it requires significantly fewer human interactions and interventions compared to PentestGPT. Furthermore, we recruit a group of security industry professional volunteers for a user survey and perform a qualitative analysis to evaluate AutoPentester against industry practices and compare it with PentestGPT. On average, AutoPentester received a score of 3.93 out of 5 based on user reviews, which was 19.8% higher than PentestGPT. Yasod Ginige, Akila Niroshan, Sajal Jain, Suranga Seneviratne |
TrustCom | 4 |
| 2025 | CRAFT: Class Ranking Aware Fine-Tuning for Enhanced Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection remains a key challenge preventing the rollout of key AI technologies like autonomous vehicles into the mainstream as classifiers trained on in-distribution (ID) data are unable to gracefully handle OOD data. While OOD detection remains an active area of research, current post-hoc methods often suffer from limited separability between ID and OOD, and outlier exposure-based methods lack generalisation to unseen outlier types. We present CRAFT, a fine-tuning approach for arming pre-trained classifiers against OOD inputs without requiring access to outliers. The key insight that underpins our approach is that during pre-training, classifiers implicitly learn a ranking across the ID classes that is not respected by OOD data. Therefore, a form of fine-tuning without outliers of a pre-trained classifier can sharpen the rank order of the classes, making them sensitive to the presence of OOD data. Furthermore, the fine-tuned model does not impact the ability of the classifier to correctly classify ID inputs to their respective classes. Experiments on CIFAR-10, CIFAR-100, and ImageNet-200 demonstrate that CRAFT outperforms 33 existing methods, particularly in the more challenging near-OOD detection, as well as in overall OOD detection consistency and ID classification accuracy. Naveen Karunanayake, Suranga Seneviratne, Sanjay Chawla |
WACV | 2 |
| 2025 | Demo: P4 Based In-network ML with Federated Learning to Secure and Slice IoT NetworksabstractRecent cyberattacks have increasingly targeted distributed networking environments like IoT networks. To detect these attacks, hidden under network traffic encryption, many centralized Machine Learning (ML) based solutions have been introduced, which are not well suited for IoT networks. This work proposes PIFL a practical approach to secure IoT networks by combining federated learning, in-network ML using P4-enabled devices, software-defined networks, and binarized neural networks. PIFL detects compromised edge devices and isolates them into separate network slices based on trust parameters derived from their behavior. We demonstrate the feasibility of PIFL using an experimental testbed with three intelligent network devices and seven IoT devices implemented on Raspberry Pi devices. Chamara Manoj Madarasingha Kattadige, Thilini Dahanayaka, Kanchana Thilakarathna, Suranga Seneviratne, Young Choon Lee, Salil S. Kanhere, Albert Y. Zomaya, Aruna Seneviratne, Phil Ridley |
WoWMoM | 4 |
| 2025 | LLMs are one-shot URL classifiers and explainers
Fariza Rashid, Nishavi Ranaweera, Ben Doyle, Suranga Seneviratne |
Comput. Networks | 4 |
| 2025 | Long-tail learning with rebalanced contrastive lossabstractIntegrating supervised contrastive loss to cross entropy-based classification has recently been proposed as a solution to address the long-tail learning problem. However, when the class imbalance ratio is high, it requires adjusting the supervised contrastive loss to support the tail classes, as the conventional contrastive learning is biased towards head classes by default. To this end, we present Rebalanced Contrastive Learning (RCL), an efficient means to increase the long-tail classification accuracy by addressing three main aspects: 1. Feature space balancedness – Equal division of the feature space among all the classes 2. Intra-Class compactness – Reducing the distance between same-class embeddings 3. Regularization – Enforcing larger margins for tail classes to reduce overfitting. RCL adopts class frequency-based SoftMax loss balancing to supervised contrastive learning loss and exploits scalar multiplied features fed to the contrastive learning loss to enforce compactness. We implement RCL on the Balanced Contrastive Learning (BCL) Framework, which has the SOTA performance. Our experiments on three benchmark datasets CIFAR10-LT,CIFAR100-LT and ImageNet-LT demonstrate the richness of the learnt embeddings and increased top-1 balanced accuracy RCL provides to the BCL framework. We further demonstrate that the performance of RCL as a standalone loss also achieves state-of-the-art level accuracy. • Rebalances supervised contrastive learning to support long-tail classification. • Optimizes the learnt feature distribution to support rare class classification. • Improved performance over datasets: CIFAR10 Lt, CIFAR100 Lt, and ImageNet Lt. • Enhanced class separability, feature space balancedness and intra-class compactness. • Can apply complementary to existing long-tail classification frameworks. Charika De Alvis, Dishanika Denipitiyage, Suranga Seneviratne |
Neurocomputing | 3 |
| 2025 | Quantifying and Exploiting Adversarial Vulnerability: Gradient-Based Input Pre-Filtering for Enhanced Performance in Black-Box AttacksabstractWe investigate the vulnerability of inputs in an adversarial setting and demonstrate that certain samples are more susceptible to adversarial perturbations compared to others. Specifically, we employ a simple yet effective approach to quantify the adversarial vulnerability of inputs, which relies on the clipped gradients of the loss with respect to the input. Our observations indicate that inputs with a low percentage of zero gradient components tend to be more vulnerable to attacks. These findings are supported by a theoretical explanation on a linear model and empirical evidence on deep neural networks. Across all datasets we tested, we find that inputs with the lowest zero gradient percentage, on average, exhibit 34.5% more susceptibility to adversarial attacks than randomly selected inputs. Additionally, we demonstrate that the zero gradient percentage, as a metric, transfers across different model architectures. Finally, we propose a novel black-box attack pipeline that enhances the efficiency of conventional query-based black-box attacks and show that input pre-filtering based on Zero Gradient Percentage can boost the attack success rates, particularly under low perturbation levels. On average, across all datasets we test, our approach outperforms the conventional shadow model-based and query-based black-box attack pipelines by 44.9% and 30.4%, respectively. Naveen Karunanayake, Bhanuka Silva, Yasod Ginige, Suranga Seneviratne, Sanjay Chawla |
ACM Trans. Priv. Secur. | 4 |
| 2025 | Detecting and Characterising Mobile App Metamorphosis in Google Play StoreabstractApp markets have evolved into highly competitive and dynamic environments for developers. While the traditional app life cycle involves incremental updates for feature enhancements and issue resolution, some apps deviate from this norm by undergoing significant transformations in their use cases or market positioning. We define this previously unstudied phenomenon as ‘app metamorphosis'. In this paper, we propose a novel and efficient multi-modal search methodology to identify apps undergoing metamorphosis and apply it to analyse two snapshots of the Google Play Store taken five years apart. Our methodology uncovers various metamorphosis scenarios, including re-births, re-branding, re-purposing, and others, enabling comprehensive characterisation. Although these transformations may register as successful for app developers based on our defined success score metric (e.g., re-branded apps performing approximately 11.3% better than an average top app), we shed light on the concealed security and privacy risks that lurk within, potentially impacting even tech-savvy end-users. Dishanika Denipitiyage, Bhanuka Silva, Kavishka Gunathilaka, Suranga Seneviratne, Anirban Mahanti, Aruna Seneviratne, Sanjay Chawla |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Federated Deep Equilibrium Learning: Harnessing Compact Global Representations to Enhance Personalization
Tuan Dung Nguyen, Tung-Anh Nguyen, Choong Seon Hong, Suranga Seneviratne, Wei Bao 0001, Nguyen Hoang Tran |
CIKM | 5 |
| 2024 | Passive Identification of WiFi Devices At-Scale: A Data-Driven ApproachabstractWiFi has emerged as the standard method for local connectivity across various devices, including smart assistants, IoT devices, smart TVs, and AR/VR devices. Identifying WiFi devices in neighborhoods has implications for law enforcement, urban planning, and socio-economic analysis. This paper introduces a novel approach to constructing WiFi device-type signatures using Information Element attributes from wildcard WiFi probe requests. Our method accurately identifies device types even when dealing with randomized MAC addresses and requires minimal training data, thus addressing limitations of existing machine learning and deep learning approaches. We evaluate our approach using a dataset of 51,726 probe requests across 50 device types, achieving an average F1 score of 99%, precision of 99%, and recall of 98% in device-type identification. Importantly, our method outperforms deep learning methods with significantly less training data, achieving a 92% F1 score with only one training sample per device type. Niruth Bogahawatta, Yasiru Senarath Karunanayaka, Suranga Seneviratne, Kanchana Thilakarathna, Rahat Masood, Salil S. Kanhere, Aruna Seneviratne |
LCN | 3 |
| 2024 | Evaluating Web-Based Privacy Controls: A User Study on Expectations and PreferencesabstractIn response to growing privacy concerns, many websites have implemented privacy controls that aim to enhance user autonomy and compliance with data protection regulations. While existing literature has evaluated individual privacy controls such as cookie consent interfaces, less attention has been given to how combining multiple privacy controls in realistic online environments affects the overall user experience. We conducted an online user study with 75 participants to explore the usability of privacy controls offered by websites across four widely used categories. Participants were asked to interact with website prototypes that differed in terms of where and how the privacy control variants were presented, then answer a survey about their experience. Our findings revealed the usability impact of design parameters on privacy controls and highlighted how user expectations vary across different demographics and website categories. We provide design recommendations that combine informative elements (privacy notices and policies) with actionable elements (privacy nudges and settings) to enhance the usability of website privacy controls for users. Yuemeng Yin, Rahat Masood, Suranga Seneviratne, Aruna Seneviratne |
TrustCom | 3 |
| 2024 | Phishing URL detection generalisation using Unsupervised Domain AdaptationabstractPhishing attacks are a prevailing problem in cybersecurity. In many data breaches, the initial entry can be traced back to phishing. URL-based phishing detection is one of the many ways of phishing attempt detection where solely the properties of the URLs are used to decide whether a given URL is phishing or not. While there are multiple existing works that use machine learning and deep learning to detect phishing URLs, in this paper, we show that such methods lack generalisation (i.e., they work effectively only when the test sets are split from the same training dataset). This is a significant issue since the vast majority of phishing attempts are short-lived and use freshly created domain names. Also, many network vantage points and middleboxes record URLs in slightly different formats and as such, URL data collected at various companies may be different. To address this, we propose an Unsupervised Domain Adaptation-based framework to increase the model transferability between datasets. We evaluate our approach using three datasets and show that the increase in cross-dataset F1 score performance is 0.06 on average and in some cases approximately as high as 0.2. Fariza Rashid, Ben Doyle, Soyeon Caren Han, Suranga Seneviratne |
Comput. Networks | 4 |
| 2024 | Single-Sensor Sparse Adversarial Perturbation Attacks Against Behavioral BiometricsabstractIn Internet of Things (IoT) deployments, sensing applications have emerged as critical tools. They combine data streams from heterogeneous, untrusted sensors to provide valuable insights or make automated decisions. This paper shows that such systems can be easily manipulated by only compromising a single sensor and perturbing the data from specific time slots rather than entire data streams in grey-box and black-box settings -attack scenarios not considered in traditional machine learning literature. Drawing from two datasets related to behavioural biometrics of smart headsets, we demonstrate that by altering just 6.2% of the data, an attacker can significantly reduce the system’s accuracy—achieving drops of 85% in grey-box scenarios and 74.5% in black-box settings. Next, we show that while adversarial training can mitigate such attacks, an attacker can overcome such defences by increasing the perturbation only in specific time steps. To this end, we propose a two-step defence where we detect more significant perturbations in IoT sensor readings using anomaly detection and mitigate more minor perturbations through adversarial training. Overall, our proposed method can limit the accuracy drop to a maximum of 9.59% across all magnitudes of perturbations, thus protecting against adversarial attacks on multi-sensor systems. Ravin Gunawardena, Sandani Jayawardena, Suranga Seneviratne, Rahat Masood, Salil S. Kanhere |
IEEE Internet Things J. | 3 |
| 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly DetectionabstractWith the proliferation of the Internet of Things (IoT) and the rising interconnectedness of devices, network security faces significant challenges, especially from anomalous activities. While traditional machine learning-based intrusion detection systems (ML-IDS) effectively employ supervised learning methods, they possess limitations such as the requirement for labeled data and challenges with high dimensionality. Recent unsupervised ML-IDS approaches such as AutoEncoders and Generative Adversarial Networks (GAN) offer alternative solutions but pose challenges in deployment onto resource-constrained IoT devices and in interpretability. To address these concerns, this paper proposes a novel federated unsupervised anomaly detection framework – FedPCA – that leverages Principal Component Analysis (PCA) and the Alternating Directions Method Multipliers (ADMM) to learn common representations of distributed non-i.i.d. datasets. Building on the FedPCA framework, we propose two algorithms, FedPE in Euclidean space and FedPG on Grassmann manifolds. Our approach enables real-time threat detection and mitigation at the device level, enhancing network resilience while ensuring privacy. Moreover, the proposed algorithms are accompanied by theoretical convergence rates even under a sub-sampling scheme, a novel result. Experimental results on the UNSW-NB15 and TON-IoT datasets show that our proposed methods offer performance in anomaly detection comparable to non-linear baselines, while providing significant improvements in communication and memory efficiency, underscoring their potential for securing IoT networks. Tung-Anh Nguyen, Tuan Dung Nguyen, Wei Bao 0001, Suranga Seneviratne, Choong Seon Hong, Nguyen Hoang Tran |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | POSTER: Performance Characterization of Binarized Neural Networks in Traffic FingerprintingabstractTraffic fingerprinting allows making inferences about encrypted traffic flows through passive observation. They have been used for tasks such as network performance management and analytics and in attacker settings such as censorship and surveillance. A key challenge when implementing traffic fingerprinting in real-time settings is how the state-of-the-art traffic fingerprint models can be ported into programmable in-network computing devices with limited computing resources. Towards this, in this work, we characterize the performance of binarized traffic fingerprinting neural networks that are efficient and well-suited for in-network computing devices and propose a new data encoding method that is better suited for network traffic. Overall, we show that the proposed binary neural network with first-layer binarization and last-layer quantization reduces the performance requirement of hardware equipment while retaining the accuracies of those models of binary datasets over 70%. Furthermore, when combined with our proposed encoding algorithm, accuracies of binarized models of numeric datasets show further improvements to achieve over 65% accuracy. Yiyan Wang, Thilini Dahanayaka, Guillaume Jourjon, Suranga Seneviratne |
AsiaCCS | 4 |
| 2023 | Robust open-set classification for encrypted traffic fingerprintingabstractEncrypted network traffic has been known to leak information about their underlying content through side-channel information leaks. Traffic fingerprinting attacks exploit this by using machine learning techniques to threaten user privacy by identifying user activities such as website visits, videos streamed, and messenger app activities. Although state-of-the-art traffic fingerprinting attacks have high performances, even undermining the latest defenses, most of them are developed under the closed-set assumption. To deploy them in practical situations, it is important to adapt them to the open-set scenario, which allows the attacker to identify its target content while rejecting other background traffic. At the same time, in practice, these models need to be deployed on in-networking devices such as programmable switches, which have limited memory and computation power. Model weight quantization can reduce the memory footprint of deep learning models while at the same time, allowing inference to be done as integer operations as opposed to floating point operations. Open-set classification in the domain of traffic fingerprinting has not been explored well in prior work and none of them explored the effect of quantization on the open-set performance of such models. In this work, we propose a framework for robust open-set classification of encrypted traffic based on three key ideas. First, we show that a well-regularized deep learning model improves the open-set classification and then we propose a novel open-set classification method with three variants that perform consistently over multiple datasets. Next, we show that traffic fingerprinting models can be quantized without a significant drop in both closed-set and open-set accuracy and therefore, they can be readily deployed on in-network computing devices. Finally, we show that when the above three components are combined, the resulting open-set classifier outperforms all other open-set classification methods evaluated across five datasets with a minimum and maximum increase in F1_Score of 8.9% and 77.3% respectively. Thilini Dahanayaka, Yasod Ginige, Yi Huang 0023, Guillaume Jourjon, Suranga Seneviratne |
Comput. Networks | 5 |
| 2023 | Privacy-preserving spam filtering using homomorphic and functional encryption
Tham Nguyen, Naveen Karunanayake, Suranga Seneviratne, Peizhao Hu |
Comput. Commun. | 4 |
| 2023 | Calibrated reconstruction based adversarial autoencoder model for novelty detection
Yi Huang 0023, Ying Li 0039, Guillaume Jourjon, Suranga Seneviratne, Kanchana Thilakarathna, Adriel Cheng, Darren Webb |
Pattern Recognit. Lett. | 4 |
| 2022 | Inline Traffic Analysis Attacks on DNS over HTTPSabstractEven though end-to-end encryption was introduced to Domain Name System (DNS) communications to ensure user privacy and there is an increase in adoption of DNS over HTTPS (DoH), prior research has demonstrated that encrypted DNS traffic is vulnerable to traffic analysis attacks. However, these attacks were demonstrated under strong assumptions such as handling only closed-set classification or doing only post-event analysis. In this work we demonstrate traffic analysis attacks on DoH without such strong assumptions. We first show the feasibility of website fingerprinting over DoH traffic and present an inline traffic analysis attack that achieve over 90% accuracy using DoH traces of length as short as ten packets. Next, we propose a novel open-set classification method and achieve over 75% accuracy on both closed-set and open-set samples for the open-set scenario. Finally, we demonstrate that the same attack can be performed without any knowledge on the start of the activity. Thilini Dahanayaka, Guillaume Jourjon, Suranga Seneviratne |
LCN | 4 |
| 2022 | Dissecting traffic fingerprinting CNNs with filter activations
Thilini Dahanayaka, Guillaume Jourjon, Suranga Seneviratne |
Comput. Networks | 3 |
| 2022 | From traffic classes to content: A hierarchical approach for encrypted traffic classification
Ying Li 0039, Yi Huang 0023, Suranga Seneviratne, Kanchana Thilakarathna, Adriel Cheng, Guillaume Jourjon, Darren Webb, David B. Smith 0001 |
Comput. Networks | 3 |
| 2022 | Task adaptive siamese neural networks for open-set recognition of encrypted network traffic with bidirectional dropout
Yi Huang 0023, Ying Li 0039, Timothy Heyes, Guillaume Jourjon, Adriel Cheng, Suranga Seneviratne, Kanchana Thilakarathna, Darren Webb |
Pattern Recognit. Lett. | 6 |
| 2022 | A Multi-Modal Neural Embeddings Approach for Detecting Mobile Counterfeit Apps: A Case Study on Google Play StoreabstractCounterfeit apps impersonate existing popular apps in attempts to misguide users to install them for various reasons such as collecting personal information, spreading malware, or simply to increase their advertisement revenue. Many counterfeits can be identified once installed, however even a tech-savvy user may struggle to detect them before installation as app icons and descriptions can be quite similar to the original app. To this end, this paper proposes to leverage the recent advances in deep learning methods to create image and text embeddings so that counterfeit apps can be efficiently identified when they are submitted to be published in app markets. We show that for the problem of counterfeit detection, a novel approach of combiningcontent embeddingsandstyle embeddings(given by the Gram matrix of CNN feature maps) outperforms the baseline methods for image similarity such as SIFT, SURF, LATCH, and various image hashing methods. We first evaluate the performance of the proposed method on two well-known datasets for evaluating image similarity methods and show that, content, style, and combined embeddings increaseprecision@kandrecall@kby 10-15 percent and 12-25 percent, respectively when retrieving five nearest neighbours. Second specifically for the app counterfeit detection problem, combined content and style embeddings achieve 12 and 14 percent increase inprecision@kandrecall@k, respectively compared to the baseline methods. We also show that adding text embeddings further increases the performance by 5 and 6 percent in terms ofprecision@kandrecall@k, respectively when$k$is five. Third, we present an analysis of approximately 1.2 million apps from Google Play Store and identify a set of potential counterfeits for top-10,000 popular apps. Under a conservative assumption, we were able to find 2,040 potential counterfeits that contain malware in a set of 49,608 apps that showed high similarity to one of the top-10,000 popular apps in Google Play Store. We also find 1,565 potential counterfeits asking for at least five additional dangerous permissions than the original app and 1,407 potential counterfeits having at least five extra third party advertisement libraries. Naveen Karunanayake, Jathushan Rajasegaran, Ashanie Gunathillake, Suranga Seneviratne, Guillaume Jourjon |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | SMAUG: Streaming Media Augmentation Using CGANs as a Defence Against Video FingerprintingabstractTraffic fingerprinting and developing defenses against it has always been an arms race between the attackers and the defenders. The rapid evolution of deep learning methods makes developing stronger traffic fingerprinting models much easier, while overhead, latency, and deployment constraints restrict the abilities of the defenses. As such, there is always the need of coming up with novel defenses against traffic fingerprinting. In this paper, we propose SMAUG, a novel CGAN-based (Conditional Generative Adversarial Network) defense to protect video streaming traffic against fingerprinting. We first assess the performance of various GANs in video streaming traffic synthesis using multiple GAN quality metrics and show that CGAN outperforms other types of GANs such as basic GANs and WGANs (Wasserstein GAN). Our proposed defense, SMAUG, uses CGANs to synthesize video traffic flows and use those synthesized flows to camouflage the original traffic that needs protection. We compare SMAUG with other state-of-the-art defenses - FPA and d*-private methods, as well as a kernel density estimation-based baseline and show that SMAUG provides better privacy with lower overhead and delay. Alexander Vaskevich, Thilini Dahanayaka, Guillaume Jourjon, Suranga Seneviratne |
NCA | 4 |
| 2021 | MusicID: A Brainwave-Based User Authentication System for Internet of Things
Jinani Sooriyaarachchi, Suranga Seneviratne, Kanchana Thilakarathna, Albert Y. Zomaya |
IEEE Internet Things J. | 2 |
| 2021 | Power Control for Body Area Networks: Accurate Channel Prediction by Lightweight Deep LearningabstractRecent advances in the Internet of Things (IoT) are reforming the health care industry by providing higher communication efficiency, lower costs, and higher mobility. Among the many IoT applications, wireless body area networks (BANs) are a remarkable solution caring for a rapidly growing aged population. Predictive transmit power control schemes improve BAN communications' reliability and energy efficiency through long-term optimal radio resources allocation that supports consistent pervasive healthcare services. Here, we propose LSTM-based neural network (NN) prediction methods that provide long-term accurate channel gain prediction of up to 2 s over nonstationary BAN on-body channels. An incremental learning scheme, which enables the LSTM predictor to operate online, is also developed for dynamic scenarios. Our main contribution is a lightweight NN predictor, “LiteLSTM,” that has a compact structure and higher computational efficiency than other variants. We show that LiteLSTM remains functional under an incremental learning scheme, with only marginal performance degradation when implemented on hand-held devices. For optimal power allocation, we develop an interquartile range (IQR)-based power control for our channel prediction. When extensively tested using empirical channel measurements at different sampling rates, our proposed methods outperform the existing state-of-the-art methods in terms of prediction accuracy, power consumption, level crossing rate (LCR), and outage probability and duration. Yizhou Yang, David B. Smith 0001, Jathushan Rajasegaran, Suranga Seneviratne |
IEEE Internet Things J. | 4 |
| 2021 | SETA++: Real-Time Scalable Encrypted Traffic Analytics in Multi-Gbps NetworksabstractThe security and privacy of the end-users are a few of the most important components of a communication network. Though end-to-end encryption (e.g., TLS/SSL) fulfils this requirement, it makes inspecting network traffic with legacy solutions such as Deep Packet Inspection difficult. Recent Machine Learning techniques have shown outstanding performance in encrypted traffic classification. Nevertheless, such approaches require efficient flow sampling at real enterprise-scale networks due to the sheer volume of transferred data. Through this paper, we propose a holistic architecture to extract flow information of encrypted data at multi Gbps line rate using sampling and sketching mechanisms, enabling network operators to estimate flow size distribution accurately and understand the behavior of VPN-obfuscated traffic. Using over 6000 video traffic traces, under three main evaluation scenarios based on trace duration and starting time point, we show that it is possible to achieve 99% accuracy for service provider classification and over 90% accuracy for content classification for a given service provider in the best case. We also deploy our solution at an operational enterprise-scale network leveraging kernel bypassing to demonstrate its capability to efficiently sample live traffic for analytics. Chamara Manoj Madarasingha Kattadige, Kwon Nung Choi, Achintha Wijesinghe, Arpit Nama, Kanchana Thilakarathna, Suranga Seneviratne, Guillaume Jourjon |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2020 | Poster Abstract: Passive Activity Classification of Smart Homes through Wireless Packet SniffingabstractNetwork communications, despite being encrypted, leak crucial information via side channels. WiFi networks are more prone to such side-channel attacks since any attacker within the network’s range can passively eavesdrop the channel. With the increasing number of smart home devices and sensors connecting to private WiFi networks, it is essential to understand the inadvertent information leakage through WiFi side-channels. Our work demonstrates how fine-granular information on the activities happening inside a house can be inferred by passively monitoring WiFi network traffic. In particular, we were able to correctly classify various user interactions with simple IoT devices such as smart bulbs or power sockets as well as advanced voice-based intelligent assistants. Kwon Nung Choi, Thilini Dahanayaka, David Kennedy, Kanchana Thilakarathna, Suranga Seneviratne, Salil S. Kanhere, Prasant Mohapatra |
IPSN | 5 |
| 2020 | Triplet Mining-based Phishing Webpage DetectionabstractPhishing web pages impersonate legitimate websites to trick users into entering sensitive information such as their credentials. In many high profile data breaches, the initial entry points have been traced back to phishing attacks. Attackers are using increasingly sophisticated methods such as code obfuscation to bypass existing phishing detection systems. Since phishing websites show very high visual similarity to the respective target pages, recent advances in Convolutional Neural Networks (CNN) can be leveraged to build better phishing detection systems. In this work, we propose a novel CNN architecture consisting of two paths to capture the content similarity and structural similarity between web pages. Leveraging the fact that web pages of the same web site are visually similar, we use triplet learning to train our model without any labelled phishing examples. Kalana Abeywardena, Lexi Brent, Suranga Seneviratne, Ralph Holz |
LCN | 4 |
| 2020 | SETA: Scalable Encrypted Traffic Analytics in Multi-Gbps NetworksabstractWhile end-to-end encryption brings security and privacy to the end-users, it makes legacy solutions such as Deep Packet Inspection ineffective. Despite the recent work in machine learning-based encrypted traffic classification, these new techniques would require, if they were to be deployed in real enterprise-scale networks, an enhanced flow sampling due to sheer volume of data being traversed. In this paper, we propose a holistic architecture that can cope with encryption and multi-Gbps line rate with sampling and sketching flow statistics, which allows network operators to both accurately estimate the flow size distribution and identify the nature of VPN-obfuscated traffic. With over 6000 video traffic traces, we show that it is possible to achieve 99% accuracy for service provider classification even with sampled possibly inaccurate data. Kwon Nung Choi, Achintha Wijesinghe, Chamara Manoj Madarasingha Kattadige, Kanchana Thilakarathna, Suranga Seneviratne, Guillaume Jourjon |
LCN | 5 |
| 2020 | Understanding Traffic Fingerprinting CNNsabstractHTTPS encrypted traffic can leak information about underlying contents through various statistical properties of traffic flows like packet lengths and timing, opening doors to traffic fingerprinting attacks. Recently proposed traffic fingerprinting attacks leveraged Convolutional Neural Networks (CNNs) and recorded very high accuracies undermining the state-of-the-art mitigation techniques. In this paper, we methodically dissect such CNNs with the objectives of building further accurate and scalable traffic classifiers and understanding the inner workings of such CNNs to develop effective mitigation techniques. By conducting experiments with three datasets, we show that website fingerprinting CNNs focus majorly on the initial parts of traces instead of longer windows of continuous uploads or downloads. Next, we show that traffic fingerprinting CNNs exhibit transfer-learning capabilities allowing identification of new websites with fewer data. Finally, we show that traffic fingerprinting CNNs outperform RNNs because of their resilience to random shifts in data happening due to varying network conditions. Thilini Dahanayaka, Guillaume Jourjon, Suranga Seneviratne |
LCN | 3 |
| 2020 | Security Apps under the Looking Glass: An Empirical Analysis of Android Security AppsabstractThird-party security apps are an integral part of the Android app ecosystem. Many users install them as an extra layer of protection for their devices. By installing security apps, the smartphone users place a significant amount of trust on them allowing access to many smartphone resources that contain personal information such as the storage, text messages, email, and browser history. As such, it is essential to understand the mobile security apps ecosystem. In this paper, we present the first empirical study of Android security apps. We analyse 100 Android security apps from multiple aspects and offer insights to their operations and behaviours. Our results show that 20% of the security apps resell the data they collect to third parties; in some cases, even without the user consent. Also, we show that around 50% of the security apps fail to identify known malware. Weixian Yao, Yexuan Li, Weiye Lin, Tianhui Hu, Imran Chowdhury, Rahat Masood, Suranga Seneviratne |
LCN | 7 |
| 2020 | Side-channel information leaks of Z-wave smart home IoT devices: demo abstractabstractZ-Wave is one of the key access protocols of the Internet of Things (IoT). It is highly popular in home automation and security system applications due to its minimum power consumption, reliability, and cost effectiveness. With an estimate of over 100 million deployed Z-Wave devices around the globe, it is essential to understand their security landscape. For instance, Z-Wave devices can leak personal information about the home dwellers as well as their possessions and buglers can use compromised Z-Wave devices to disable security systems or even to feed incorrect information. In this paper, we present an experiment setup and early results of side-channel information leaks of Z-Wave. We show that Z-Wave traffic despite being encrypted, leaks information through side-channels and an attacker who can passively capture Z-Wave frames by simply being in the vicinity of a house can identify Z-Wave devices inside the house. Jung-Chang Liou, Sajal Jain, Sooraj Randhir Singh, Dhit Taksinwarajan, Suranga Seneviratne |
SenSys | 5 |
| 2020 | Making Sense of Occluded Scenes using Light Field Pre-processing and Deep-learningabstractA combined approach of low-complexity light field depth filtering and deep learning is proposed for object classification in the presence of partial occlusions. The proposed approach exploits depth information embedded in multi-perspective four-dimensional (4-D) light fields via low-complexity 4-D sparse depth filtering and deep-learning. The proposed 4-D depth filter, designed using numerical optimization techniques by formulating as an ℓ1- ℓ∞minimization problem, is shown to outperform typical light field refocusing based on 4-D shift-sum averaging filters. Experiments conducted using a light field dataset acquired by a Lytro camera verify 45% and 27% better performance in terms of object classification accuracy compared to the cases when no depth filtering is employed and standard shift-sum refocusing is employed, respectively. Namalka Liyanage, Kalana Abeywardena, Sakila S. Jayaweera, Chamith Wijenayake, Chamira U. S. Edussooriya, Suranga Seneviratne |
TENCON | 6 |
| 2019 | DeepCaps: Going Deeper With Capsule NetworksabstractCapsule Network is a promising concept in deep learning, yet its true potential is not fully realized thus far, providing sub-par performance on several key benchmark datasets with complex data. Drawing intuition from the success achieved by Convolutional Neural Networks (CNNs) by going deeper, we introduce DeepCaps, a deep capsule network architecture which uses a novel 3D convolution based dynamic routing algorithm. With DeepCaps, we surpass the state-of-the-art capsule domain networks results on CIFAR10, SVHN and Fashion MNIST, while achieving a 68% reduction in the number of parameters. Further, we propose a class independent decoder network, which strengthens the use of reconstruction loss as a regularization term. This leads to an interesting property of the decoder, which allows us to identify and control the physical attributes of the images represented by the instantiation parameters. Jathushan Rajasegaran, Vinoj Jayasundara 0001, Sandaru Jayasekara, Hirunima Jayasekara, Suranga Seneviratne, Ranga Rodrigo |
CVPR | 5 |
| 2019 | Deep Learning Channel Prediction for Transmit Power Control in Wireless Body Area NetworksabstractThe general non-stationarity of the wireless body area network (WBAN) narrowband radio channel makes long-term prediction very challenging. However, long short-term memory (LSTM) is a deep learning recurrent neural network (RNN) architecture that is proposed here to learn these atypical radio channel dynamics and make channel predictions. Thus, here we propose an LSTM-based RNN channel prediction framework providing long-term channel prediction up to 2s with low error. To address practical scenarios where information packets are transmitted continuously, we outline a timing scheme, which enables the LSTM predictor to operate online. We employ the proposed method in transmit power control for everyday on-body, measured, WBAN channels. When compared with existing approaches, the proposed channel prediction reduces circuit power consumption significantly while improving communications reliability. Yizhou Yang, David B. Smith 0001, Suranga Seneviratne |
ICC | 3 |
| 2019 | TextCaps: Handwritten Character Recognition With Very Small DatasetsabstractMany localized languages struggle to reap the benefits of recent advancements in character recognition systems due to the lack of substantial amount of labeled training data. This is due to the difficulty in generating large amounts of labeled data for such languages and inability of deep learning techniques to properly learn from small number of training samples. We solve this problem by introducing a technique of generating new training samples from the existing samples, with realistic augmentations which reflect actual variations that are present in human hand writing, by adding random controlled noise to their corresponding instantiation parameters. Our results with a mere 200 training samples per class surpass existing character recognition results in the EMNIST-letter dataset while achieving the existing results in the three datasets: EMNIST-balanced, EMNIST-digits, and MNIST. We also develop a strategy to effectively use a combination of loss functions to improve reconstructions. Our system is useful in character recognition for localized languages that lack much labeled training data and even in other related more general contexts such as object recognition. Vinoj Jayasundara 0001, Sandaru Jayasekara, Hirunima Jayasekara, Jathushan Rajasegaran, Suranga Seneviratne, Ranga Rodrigo |
WACV | 5 |
| 2019 | A Multi-modal Neural Embeddings Approach for Detecting Mobile Counterfeit AppsabstractCounterfeit apps impersonate existing popular apps in attempts to misguide users. Many counterfeits can be identified once installed, however even a tech-savvy user may struggle to detect them before installation. In this paper, we propose a novel approach of combining content embeddings and style embeddings generated from pre-trained convolutional neural networks to detect counterfeit apps. We present an analysis of approximately 1.2 million apps from Google Play Store and identify a set of potential counterfeits for top-10,000 apps. Under conservative assumptions, we were able to find 2,040 potential counterfeits that contain malware in a set of 49,608 apps that showed high similarity to one of the top-10,000 popular apps in Google Play Store. We also find 1,565 potential counterfeits asking for at least five additional dangerous permissions than the original app and 1,407 potential counterfeits having at least five extra third party advertisement libraries. Jathushan Rajasegaran, Naveen Karunanayake, Ashanie Gunathillake, Suranga Seneviratne, Guillaume Jourjon |
WWW | 4 |
| 2019 | Light weight and fine-grained access mechanism for secure access to outsourced dataabstractSummary In this paper, we explore the problem of providing selective read/write access to the outsourced data for clients using mobile devices in an environment that supports users from multiple domains and where attributes are generated by multiple authorities. We consider Ciphertext‐Policy Attribute‐based Encryption (CP‐ABE) scheme as it can provide access control on encrypted outsourced data. One limitation of CP‐ABE is that the users can modify the access policy specified by the data owner if write operations are introduced in the scheme. We propose a protocol for providing different levels of access to outsourced data that permits the authorized users to perform write operation without altering the access policy specified by the data owner. Our scheme provides fine‐grained read/write access to the users, accompanied with a light weight signature scheme and computationally inexpensive user revocation mechanism suitable for resource‐constrained mobile devices. We provide a theoretical analysis of the security of the proposed protocol and the experimental results measured from a real‐world testbed. Mosarrat Jahan, Suranga Seneviratne, Partha Sarathi Roy 0001, Kouichi Sakurai, Aruna Seneviratne, Sanjay K. Jha |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | A First Look at SIM-Enabled Wearables in the Wild
Harini Kolamunna, Ilias Leontiadis, Diego Perino, Suranga Seneviratne, Kanchana Thilakarathna, Aruna Seneviratne |
Internet Measurement Conference | 4 |
| 2018 | Deep Content: Unveiling Video Streaming Content from Encrypted WiFi Trafficabstract© 2018 IEEE. The proliferation of smart devices has led to an exponential growth in digital media consumption, especially mobile video for content marketing. The vast majority of the associated Internet traffic is now end-to-end encrypted, and while encryption provides better user privacy and security, it has made network surveillance an impossible task. The result is an unchecked environment for exploiters and attackers to distribute content such as fake, radical and propaganda videos. Recent advances in machine learning techniques have shown great promise in characterising encrypted traffic captured at the end points. However, video fingerprinting from passively listening to encrypted traffic, especially wireless traffic, has been reported as a challenging task due to the difficulty in distinguishing retransmissions and multiple flows on the same link. We show the potential of fingerprinting videos by passively sniffing WiFi frames in air, even without connecting to the WiFi network. We have developed Multi-Layer Perceptron (MLP) and Recurrent Neural Networks (RNNs) that are able to identify streamed YouTube videos from a closed set, by sniffing WiFi traffic encrypted at both Media Access Control (MAC) and Network layers. We compare these models to the state-of-the-art wired traffic classifier based on Convolutional Neural Networks (CNNs), and show that our models obtain similar results while requiring significantly less computational power and time (approximately a threefold reduction). Ying Li 0039, Yi Huang 0023, Suranga Seneviratne, Kanchana Thilakarathna, Adriel Cheng, Darren Webb, Guillaume Jourjon |
NCA | 4 |
| 2017 | BreathPrint: Breathing Acoustics-based User AuthenticationabstractWe propose BreathPrint, a new behavioural biometric signature based on audio features derived from an individual's commonplace breathing gestures. Specifically, BreathPrint uses the audio signatures associated with the three individual gestures: sniff, normal, and deep breathing, which are sufficiently different across individuals. Using these three breathing gestures, we develop the processing pipeline that identifies users via the microphone sensor on smartphones and wearable devices. In BreathPrint, a user performs breathing gestures while holding the device very close to their nose. Using off-the-shelf hardware, we experimentally evaluate the BreathPrint prototype with 10 users, observed over seven days. We show that users can be authenticated reliably with an accuracy of over 94% for all the three breathing gestures in intra-sessions and deep breathing gesture provides the best overall balance between true positives (successful authentication) and false positives (resiliency to directed impersonation and replay attacks). Moreover, we show that this breathing sound based biometric is also robust to some typical changes in both physiological and environmental context, and that it can be applied on multiple smartphone platforms. Early results suggest that breathing based biometrics show promise as either to be used as a secondary authentication modality in a multimodal biometric authentication system or as a user disambiguation technique for some daily lifestyle scenarios. Jagmohan Chauhan, Yining Hu 0001, Suranga Seneviratne, Archan Misra, Aruna Seneviratne, Youngki Lee 0001 |
MobiSys | 3 |
| 2017 | Privacy preserving data access scheme for IoT devicesabstractAttribute-based encryption schemes provide read access to data based on users' attributes. In these schemes, user privacy is compromised as the access policies are visible. This privacy issue has been addressed in literature by enabling the data owner to obfuscate the policy in a setting where a single authority generates decryption keys. However, a single authority can figure out the hidden access policy which violates user privacy. We present PPDAS, a scheme which overcomes these limitations and makes two contributions. Firstly, we present a mechanism which supports fine-grained read and write operations in a setting where decryption keys are generated by multiple attribute authorities, and the access policy is hidden from all unauthorized entities including the attribute authorities. Our scheme is also accompanied with a user revocation mechanism. Secondly, we show that it is possible to adapt the scheme for accessing data through resource-constrained devices such as smart watches and IoT devices through extensive experimental evaluations. Mosarrat Jahan, Suranga Seneviratne, Ben Chu, Aruna Seneviratne, Sanjay K. Jha |
NCA | 2 |
| 2017 | A deep dive into location-based communities in social discovery networks
Kanchana Thilakarathna, Suranga Seneviratne, Mohamed Ali Kâafar, Aruna Seneviratne |
Comput. Commun. | 2 |
| 2017 | App Miscategorization Detection: A Case Study on Google PlayabstractAn ongoing challenge in the rapidly evolving app market ecosystem is to maintain the integrity of app categories. At the time of registration, app developers have to select, what they believe, is the most appropriate category for their apps. Besides the inherent ambiguity of selecting the right category, the approach leaves open the possibility of misuse and potential gaming by the registrant. Periodically, the app store will refine the list of categories available and potentially reassign the apps. However, it has been observed that the mismatch between the description of the app and the category it belongs to, continues to persist. Although some common mechanisms (e.g., a complaint-driven or manual checking) exist, they limit the response time to detect miscategorized apps and still open the challenge on categorization. We introduce FRAC+: (FR)amework for (A)pp (C)ategorization. FRAC+ has the following salient features: (i) it is based on a data-driven topic model and automatically suggests the categories appropriate for the app store, and (ii) it can detect miscategorizated apps. Extensive experiments attest to the performance of FRAC+. Experiments on GOOGLE Play shows that FRAC+'s topics are more aligned with GOOGLE's new categories and 0.35-1.10 percent game apps are detected to be miscategorized. Didi Surian, Suranga Seneviratne, Aruna Seneviratne, Sanjay Chawla |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Spam Mobile Apps: Characteristics, Detection, and in the Wild AnalysisabstractThe increased popularity of smartphones has attracted a large number of developers to offer various applications for the different smartphone platforms via the respective app markets. One consequence of this popularity is that the app markets are also becoming populated with spam apps. These spam apps reduce the users’ quality of experience and increase the workload of app market operators to identify these apps and remove them. Spam apps can come in many forms such as apps not having a specific functionality, those having unrelated app descriptions or unrelated keywords, or similar apps being made available several times and across diverse categories. Market operators maintain antispam policies and apps are removed through continuous monitoring. Through a systematic crawl of a popular app market and by identifying apps that were removed over a period of time, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. We build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy of over 95% with precision varying between 85% and 95% and recall varying between 38% and 98%. We further show that a limited number of features, in the range of 10--30, generated from app metadata is sufficient to achieve a satisfactory level of performance. On a set of 180,627 apps that were present at the app market during our crawl, our classifier predicts 2.7% of the apps as potential spam. Finally, we perform additional manual verification and show that human reviewers agree with 82% of our classifier predictions. Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra |
ACM Trans. Web | 1 |
| 2016 | An Analysis of the Privacy and Security Risks of Android VPN Permission-enabled Apps
Muhammad Ikram 0001, Narseo Vallina-Rodriguez, Suranga Seneviratne, Mohamed Ali Kâafar, Vern Paxson |
Internet Measurement Conference | 3 |
| 2015 | SSIDs in the wild: Extracting semantic information from WiFi SSIDsabstractWiFi networks are becoming increasingly ubiquitous. In addition to providing network connectivity, WiFi finds applications in areas such as indoor and outdoor localisation, home automation, and physical analytics. In this paper, we explore the semantics of one key attribute of a WiFi network, SSID name. Using a dataset of approximately 120,000 WiFi access points and their corresponding geo-locations, we use a set of similarity metrics to relate SSID names to known business venues such as cafes, theatres, and shopping centres. Such correlations can be exploited by an adversary who has access to smartphone users preferred networks lists to build an accurate profile of the user and thus can be a potential privacy risk to the users. Suranga Seneviratne, Fangzhou Jiang, Mathieu Cunche, Aruna Seneviratne |
LCN | 1 |
| 2015 | A measurement study of tracking in paid mobile applicationsabstractSmartphone usage is tightly coupled with the use of apps that can be either free or paid. Numerous studies have investigated the tracking libraries associated with free apps. Only a limited number of these have focused on paid apps. As expected, these investigations indicate that tracking is happening to a lesser extent in paid apps, yet there is no conclusive evidence. This paper provides the first large-scale study of paid apps. We analyse top paid apps obtained from four different countries: Australia, Brazil, Germany, and US, and quantify the level of tracking taking place in paid apps in comparison to free apps. Our analysis shows that 60% of the paid apps are connected to trackers that collect personal information compared to 85%--95% in free apps. We further show that approximately 20% of the paid apps are connected to more than three trackers. With tracking being pervasive in both free and paid apps, we then quantify the aggregated privacy leakages associated with individual users. Using the data of user installed apps of over 300 smartphone users, we show that 50% of the users are exposed to more than 25 trackers which can result in significant leakages of privacy. Suranga Seneviratne, Harini Kolamunna, Aruna Seneviratne |
WISEC | 1 |
| 2015 | Early Detection of Spam Mobile AppsabstractIncreased popularity of smartphones has attracted a large number of developers to various smartphone platforms. As a result, app markets are also populated with spam apps, which reduce the users' quality of experience and increase the workload of app market operators. Apps can be "spammy" in multiple ways including not having a specific functionality, unrelated app description or unrelated keywords and publishing similar apps several times and across diverse categories. Market operators maintain anti-spam policies and apps are removed through continuous human intervention. Through a systematic crawl of a popular app market and by identifying a set of removed apps, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. Finally, we build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy over 95% with precision varying between 85%-95% and recall varying between 38%-98%. By applying the classifier on a set of apps present at the app market during our crawl, we estimate that at least 2.7% of them are spam apps. Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra |
WWW | 1 |