EDBT 2026 Demo / reviewers in the wild / expert
Arshad Jhumka
dblp:45/4421
· DBLP profile ↗
70ranked-venue papers
14as first author
23since 2021 · last 2025
0000-0003-0540-2845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 25 · 4 first-author · 10 since 2021Systems, architecture and hardware · 23 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 13 · 5 first-author · 2 since 2021Computer networks · 11 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism DetectionabstractThe detection of sexism in online content remains an open problem, as harmful language disproportionately affects women and marginalized groups.While automated systems for sexism detection have been developed, they still face two key challenges: data sparsity and the nuanced nature of sexist language.Even in large, well-curated datasets like the Explainable Detection of Online Sexism (EDOS), severe class imbalance hinders model generalization.Additionally, the overlapping and ambiguous boundaries of fine-grained categories introduce substantial annotator disagreement, reflecting the difficulty of interpreting nuanced expressions of sexism.To address these challenges, we propose two prompt-based data augmentation techniques: Definition-based Data Augmentation (DDA), which leverages categoryspecific definitions to generate semanticallyaligned synthetic examples, and Contextual Semantic Expansion (CSE), which targets systematic model errors by enriching examples with task-specific semantic features.To further improve reliability in fine-grained classification, we introduce an ensemble strategy that resolves prediction ties by aggregating complementary perspectives from multiple language models.Our experimental evaluation on the EDOS dataset demonstrates state-of-the-art performance across all tasks, with notable improvements of macro F1 by 1.5 points for binary classification (Task A) and 4.1 points for fine-grained classification (Task C) 1 .Warning: This paper includes examples that might be offensive and upsetting.1 Code and resources publicly available at: https://github.com/Sahrish42/explaining_matters_ sexism_detection_acl2025 Variation 1: "They lack the mental toughness to handle high-stakes leadership roles" Variation 2: "You can't trust women with important decisions?they're too caught up in their emotions."Variation 3: "Women just let their feelings get in the way of making logical choices." VariationsAs part of a fine-grained sexism detection project aimed at categorizing harmful language into distinct sexist categories: Ø Generate 3 Variations of the following Example Ø Reflect the Sexist Intent: Ensure each variation clearly embodies the sexist intent and nature described in the vector and definition.Ø Match the Tone: The tone should align with the emotional intensity and aggression described in definition.Ø Mimic Informal Social Media Language: Craft statements that resemble informal language typical of social media platforms like Reddit or Gab, using slang, abbreviations, or casual expressions as appropriate.• Vector: 2.1 Descriptive Attacks• Definition: Characterizing or describing women in a derogatory manner.This could include negative generalizations about women's abilities, appearance, sexual behavior, intellect, character, or morals.• Example: "Women are too soft and emotional to make rational decisions'' Sahrish Khan, Arshad Jhumka, Gabriele Pergola |
ACL (1) | 2 |
| 2025 | RIFLES: Resource-effIcient Federated LEarning via SchedulingabstractFederated Learning (FL) is a privacy-preserving machine learning technique that allows decentralized collaborative model training across a set of distributed clients, by avoiding raw data exchange. A fundamental component of FL is the selection of a subset of clients in each round for model training by a central server. Current selection strategies are myopic in nature in that they are based on past or current interactions, often leading to inefficiency issues such as straggling clients. In this paper, we address this serious shortcoming by proposing the RIFLES approach that builds a novel availability forecasting layer to support the client selection process. We make the following contributions: (i) we formalise the sequential selection problem and reduce it to a scheduling problem and show that the problem is NP-complete, (ii) leveraging heartbeat messages from clients, RIFLES build an availability prediction layer to support (long term) selection decisions, (iii) we propose a novel adaptive selection strategy to support efficient learning and resource usage. To circumvent the inherent exponential complexity of the client selection problem, RIFLES leverages clients’ historical availability data to predict future availability by using a CNN-LSTM time series forecasting model, allowing the server to predict the optimal participation times of clients, thereby enabling informed selection decisions. By comparing against other FL techniques, we show that RIFLES provide significant improvement by between 10%-50% on a variety of metrics such as accuracy and test loss. To the best of our knowledge, it is the first work to investigate FL as a scheduling problem. Sara Alosaime, Arshad Jhumka |
MASS | 2 |
| 2025 | Deep learning-based prediction of major page faults in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, Aladdin Ayesh |
CCF Trans. High Perform. Comput. | 2 |
| 2025 | Deep learning-based prediction of reflection attacks using NetFlow data
Edward Chuah, Arshad Jhumka, Aladdin Ayesh |
Comput. Secur. | 2 |
| 2024 | CheckIn: Efficiently Checkpointing Intermittent Sensor-Based Internet of Things (IoT) NetworksabstractOne of the major limitations of current Internet of Things (IoT) sensor networks is the finite energy supply available for computation and communication. A potential solution that has been proposed is energy harvesting, which enables embedded devices to mitigate their dependency on traditional battery-driven power sources. However, energy supply due to energy harvesting varies both spatially and temporally, leading to nodes crashing due to energy exhaustion, with application(s) losing their operational state. Efficient state checkpointing in non-volatile memory (NVM) has been proposed to enable forward progress, albeit at the expense of significant overheads (viz., energy and time), thus a trade-off is necessary. It has previously been shown that, while checkpointing has a positive impact on the performance of some applications in transiently-powered nodes, checkpointing may adversely affect the efficiency of other applications. Our objective in this paper is to study checkpointing, to better understand how to boost its performance. We thus make the following contributions: (i) we introduce the CheckIn problem, which formalises the checkpointing problem in intermittent (transiently-powered) networks as a task scheduling problem, (ii) we prove that CheckIn is NP-complete, (iii) to circumvent the computational complexity, we propose an adaptive heuristic and compare its efficiency against known checkpointing techniques. Results show that our heuristic significantly outperforms other checkpointing techniques. Jawaher Alharbi, Arshad Jhumka |
PRDC | 2 |
| 2024 | FLARE: Availability Awareness for Resource-Efficient Federated LearningabstractAs privacy concerns increase, machine learning (ML) is undergoing a paradigm shift whereby training occurs at the end users rather than centrally in data centers, in a process called federated learning (FL). As training becomes distributed, challenges such as data heterogeneity and device availability can affect the performance of FL. Focusing on device selection, existing FL schemes either use random or full selection to ensure fairness and model accuracy or smart selection for resource efficiency. Smart selection often assumes complete knowledge of client availability. However, in reality, there can be discrepancies between the actual device availabilities and the advertised ones, leading to availability faults. We conjecture that such faults will impact of the performance of FL. In this paper, we propose FLARE (Federated Learning for Availability and Resource Efficiency), a novel fault-tolerant framework for FL that tolerates availability faults. In FLARE, FL occurs at the end devices while an availability model is built centrally for smart client selection. We demonstrate the viability of FLARE by comparing it against another state of the art technique. We show that (i) the client selection technique of FLARE outperforms current ones, enabling better resource usage efficiency, (ii) the convergence time of FLARE can be bounded and (iii) the model accuracy of FLARE is comparable to the case when full availability data of client is sent to the server. Sara Alosaime, Arshad Jhumka |
PRDC | 2 |
| 2023 | Time Machine: Generative Real-Time Model for Failure (and Lead Time) Prediction in HPC SystemsabstractHigh Performance Computing (HPC) systems generate a large amount of unstructured/alphanumeric log messages that capture the health state of their components. Due to their design complexity, HPC systems often undergo failures that halt applications (e.g., weather prediction, aerodynamics simulation) execution. However, existing failure prediction methods, which typically seek to extract some information theoretic features, fail to scale both in terms of accuracy and prediction speed, limiting their adoption in real-time production systems. In this paper, differently from existing work and inspired by current transformer-based neural networks which have revolutionized the sequential learning in the natural language processing (NLP) tasks, we propose a novel scalable log-based, self-supervised model (i.e., no need for manual labels), called Time Machine11A Time Machine allows us to travel into the future to observe the health state of HPC system and report back. Here, we travel into the log extension to report an upcoming failure., that predicts (i) forthcoming log events (ii) the upcoming failure and its location and (iii) the expected lead time to failure. Time Machine is designed by combining two stacks of transformer-decoders, each employing the self-attention mechanism. The first stack addresses the failure location by predicting the sequence of log events and then identifying if a failure event is part of that sequence. The lead time to predicted failure is addressed by the second stack. We evaluate Time Machine on four real-world HPC log datasets and compare it against three state-of-the-art failure prediction approaches. Results show that Time Machine significantly outperforms the related works on Bleu, Rouge, MCC, and F1-score in predicting forthcoming events, failure location, failure lead-time, with higher prediction speed. Khalid Ayedh Alharthi, Arshad Jhumka, Sheng Di, Lin Gui 0003, Franck Cappello, Simon McIntosh-Smith |
DSN | 2 |
| 2023 | To Checkpoint or Not to Checkpoint: That is the Question
Jawaher Alharbi, Arshad Jhumka, Daniele Palossi |
EWSN | 2 |
| 2023 | Poster Abstract: Checkpointing in Transiently Powered IoT NetworksabstractOne of the major shortcomings in IoT/sensor networks is the finite energy supply available for computation and communication. To circumvent this issue, energy harvesting has been proposed to enable embedded devices to mitigate their dependency on traditional battery-driven power source. However, energy supply due to energy harvesting often varies, leading to nodes crashing due to energy exhaustion, with application(s) losing their state. Efficient state checkpointing in non-volatile memory (NMV) has been proposed to enable forward progress, albeit at the expense of significant overhead (viz., energy and time). In this poster, we show preliminary results that, for a certain class of applications, state checkpointing may adversely affect the performance of the applications. This is different from checkpointing in traditional distributed systems, where the network topology is generally assumed to be stable. Jawaher Alharbi, Arshad Jhumka |
IPSN | 2 |
| 2023 | Addressing a Malicious Tampering Attack on the Default Isolation Level in DBMSabstractThere exists a plethora of online transaction processing (OLTP) applications, such as banking and inventory management. Database management systems (DBMS) use concurrency control mechanisms to manage OLTP transactions and their conflicts simultaneously. We introduce a DBMS attack, namely default isolation level manipulation (DILM), and investigate the security problem related to the malicious alteration of the default isolation level and suggest precautions and a technique to alleviate its effects. We assume an attacker with administrative privileges who can observe traffic and subsequently alter isolation levels, and we discuss the potential consequences of concurrency-related attacks. The algorithm is able to detect all the anomalies according to its objective function with a high accuracy rate. Abdullah Alhajri, Arshad Jhumka |
TrustCom | 2 |
| 2023 | Towards Understanding Checkpointing in Transiently Powered IoT NetworksabstractThe finite energy supply available for computation and communication is one of the major shortcomings in IoT/sensor networks. To circumvent this issue, energy harvesting has been proposed to enable embedded devices to mitigate their dependency on traditional battery-driven power sources. However, energy supply due to energy harvesting often varies, leading to nodes crashing due to energy exhaustion, with the application(s) losing their state. Efficient state checkpointing in non-volatile memory (NVM) has been proposed to enable forward progress, albeit at the expense of significant overheads (viz., energy and time). This work is based on the observation that, as the network evolves, IoT applications adapt/refresh to ensure they have up-to-date state of information, thereby reducing the need for checkpointing. However, little is known about (i) whether checkpointing is actually needed for any application, (ii) if needed, what to checkpoint for a given application and (iii) when checkpointing is needed in an application. In this paper, we address the first two problems and make the following novel contributions: (i) we formalise a variant of the checkpointing problem for IoT networks and show that it is NP-complete, (ii) we introduce a theory of state checkpointing and develop two types of checkpointing: (a) local checkpointing and (b) network checkpointing, (iii) we show that there is no general checkpointing strategy for the network-based application that will improve efficiency and (iv) we briefly survey related works and show that most if not all, works on checkpointing of IoT-based applications that lead to increased efficiency have focused on node-centric (i.e., local) programs. We run experiments on an actual testbed to confirm our expectations, and our results show that checkpointing can actually have a negative impact on applications when wrongly used. Thus, compared to existing works, we make major gains: No checkpointing is needed for network-based applications. Jawaher Alharbi, Adam P. Chester, Arshad Jhumka |
TrustCom | 3 |
| 2023 | An empirical study of major page faults for failure diagnosis in cluster systems
Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy |
J. Supercomput. | 2 |
| 2022 | Clairvoyant: a log-based transformer-decoder for failure prediction in large-scale systemsabstractSystem failures are expected to be frequent in the exascale era such as current Petascale systems. The health of such systems is usually determined from challenging analysis of large amounts of unstructured & redundant log data. In this paper, we leverage log data and propose Clairvoyant, a novel self-supervised (i.e., no labels needed) model to predict node failures in HPC systems based on a recent deep learning approach called transformer-decoder and the self-attention mechanism. Clairvoyant predicts node failures by (i) predicting a sequence of log events and then (ii) identifying if a failure is a part of that sequence. We carefully evaluate Clairvoyant and another state-of-the-art failure prediction approach - Desh, based on two real-world system log datasets. Experiments show that Clairvoyant is significantly better: e.g., it can predict node failures with an average Bleu, Rouge, and MCC scores of 0.90, 0.78, and 0.65 respectively while Desh scores only 0.58, 0.58, and 0.25. More importantly, this improvement is achieved with faster training and prediction time, with Clairvoyant being about 25X and 15X faster than Desh respectively. Khalid Ayedh Alharthi, Arshad Jhumka, Sheng Di, Franck Cappello |
ICS | 2 |
| 2022 | Information management for trust computation on resource-constrained IoT devicesabstractResource-constrained Internet of Things (IoT) devices are executing increasingly sophisticated applications that may require computational or memory intensive tasks to be executed. Due to their resource constraints, IoT devices may be unable to compute these tasks and will offload them to more powerful resource-rich edge nodes. However, as edge nodes may not necessarily behave as expected, an IoT device needs to be able to select which edge node should execute its tasks. This selection problem can be addressed by using a measure of behavioural trust of the edge nodes delivering a correct response, based on historical information about past interactions with edge nodes that are stored in memory. However, due to their constrained memory capacity, IoT devices will only be able to store a limited amount of trust information, thereby requiring an eviction strategy when its memory is full of which there has been limited investigation in the literature. To address this, we develop the concept of the memory profile of an agent and that profile’s utility. We formalise the profile eviction problem in a unified profile memory model and show it is NP-complete. To circumvent the inherent complexity, we study the performance of eviction algorithms in a partitioned profile memory model using our utility metric. Our results show that localised eviction strategies which only consider one specific type of information do not perform well. Thus we propose a novel eviction strategy that globally considers all types of trust information stored and we show that it outperforms local eviction strategies for the majority of memory sizes and agent behaviours. In this paper, we develop a concept of information utility to a trust model and formalise the problem of information eviction, which we prove to be NP-complete. We then investigate the usefulness of different eviction strategies to maximise the utility of information stored to enable trust-based task offloading. Matthew Bradbury, Arshad Jhumka, Tim Watson |
Future Gener. Comput. Syst. | 2 |
| 2022 | Quantifying Source Location Privacy Routing Performance via Divergence and Information LossabstractSource location Privacy (SLP) is an important property for security critical applications deployed over a wireless sensor network. This property specifies that the location of the source of messages needs to be kept secret from an eavesdropping adversary that is able to move around the network. Most previous work on SLP has focused on developing protocols to enhance the SLP imparted to the network under various attacker models and other conditions. Other works have focused on analysing the level of SLP being imparted by a specific protocol. In this paper, we introduce the notion of a routing matrix which captures when messages arefirstreceived. We then introduce a novel approach where an optimal SLP routing matrix is derived. In this approach, the attacker’s movement is modelled as a Markov chain where measures of conditional entropy and divergence are used to compare routing matrices and quantify if they provide high levels of SLP. We propose the notion of aproperly competing pathsthat causes an attacker todivertwhen moving towards the source. This concept provides the basis for developing aperturbation model, similar to those used in privacy-preserving data mining. We formally prove that properly competing paths are both necessary and sufficient in ensuring the existence of an SLP-aware routing matrix and show their usage in developing an SLP-aware routing matrix. Further, we show how different SLP-aware routing matrices can be obtained through different instantiations of the framework. Those instantiations are obtained based on a notion of information loss achieved through the use of the perturbation model proposed. Matthew Bradbury, Arshad Jhumka |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Threat-modeling-guided Trust-based Task Offloading for Resource-constrained Internet of ThingsabstractThere is an increasing demand for Internet of Things (IoT) networks consisting of resource-constrained devices executing increasingly complex applications. Due to these resource constraints, IoT devices will not be able to execute expensive tasks. One solution is to offload expensive tasks to resource-rich edge nodes, which requires a framework that facilitates the selection of suitable edge nodes to perform task offloading. Therefore, in this article, we present a novel trust-model-driven system architecture , based on behavioral evidence , that is suitable for resource-constrained IoT devices and supports computation offloading. We demonstrate the viability of the proposed architecture with an example deployment of the Beta Reputation System trust model on real hardware to capture node behaviors. The open environment of edge-based IoT networks means that threats against edge nodes can lead to deviation from expected behavior. Hence, we perform a threat modeling to identify such threats. The proposed system architecture includes threat handling mechanisms that provide security properties such as confidentiality, authentication, and non-repudiation of messages in required scenarios and operate within the resource constraints. We evaluate the efficacy of the threat handling mechanisms and identify future work for the standards used. Matthew Bradbury, Arshad Jhumka, Tim Watson, Denys A. Flores, Jonathan Burton, Matthew Butler 0004 |
ACM Trans. Sens. Networks | 2 |
| 2021 | Sit Here: Placing Virtual Machines Securely in Cloud EnvironmentsabstractA Cloud Computing Environment (CCE) leverages the advantages offered by virtualisation to enable virtual machines (VMs) within the same physical machine (PM) to share physical resources. Cloud service providers (CSPs) accommodate the fluctuating resource demands of cloud users dynamically, through elastic resource provisioning. CSPs use VM allocation techniques such as VM placement and VM migration to optimise the use of shared physical resources in the CCE. However, these techniques are exposed to potential security threats that can lead to the problem of malicious co-residency between VMs. This threat happens when a malicious VM is co-located with a critical (or target) VM on the same PM. Hence, the VM allocation techniques need to be made secure. While earlier works propose specific solutions to address this malicious co-residency problem, our work here proposes to investigate the allocation patterns that are more likely to lead to a secure allocation. Furthermore, we introduce a s ecurity-aware VM allocation algorithm (SRS) that aims to allocate the VMs securely, to reduce the potential for co-residency between malicious and target VMs. Our study shows: (i) our SRS algorithm outperforms all state-of-the-art allocation algorithms and (ii) algorithms that adopt stacking-based behaviours are more likely to return secure allocations than those with spreading or random behaviours. Mansour Aldawood, Arshad Jhumka, Suhaib A. Fahmy |
CLOSER | 2 |
| 2021 | Sentiment Analysis based Error Detection for Large-Scale SystemsabstractToday's large-scale systems such as High Performance Computing (HPC) Systems are designed/utilized towards exascale computing, inevitably decreasing its reliability due to the increasing design complexity. HPC systems conduct extensive logging of their execution behaviour. In this paper, we leverage the inherent meaning behind the log messages and propose a novel sentiment analysis-based approach for the error detection in large-scale systems, by automatically mining the sentiments in the log messages. Our contributions are four-fold. (1) We develop a machine learning (ML) based approach to automatically build a sentiment lexicon, based on the system log message templates. (2) Using the sentiment lexicon, we develop an algorithm to detect system errors. (3) We develop an algorithm to identify the nodes and components with erroneous behaviors, based on sentiment polarity scores. (4) We evaluate our solution vs. other state-of-the-art machine/deep learning algorithms based on three representative supercomputers' system logs. Experiments show that our error detection algorithm can identify error messages with an average MCC score and f-score of 91% and 96% respectively, while state of the art ML/deep learning model (LSTM) obtains only 67% and 84%. To the best of our knowledge, this is the first work leveraging the sentiments embedded in log entries of large-scale systems for system health analysis. Khalid Ayedh Alharthi, Arshad Jhumka, Sheng Di, Franck Cappello, Edward Chuah |
DSN | 2 |
| 2021 | Trust Trackers for Computation Offloading in Edge-Based IoT NetworksabstractWireless Internet of Things (IoT) devices will be deployed to enable applications such as sensing and actuation. These devices are typically resource-constrained and are unable to perform resource-intensive computations. Therefore, these jobs need to be offloaded to resource-rich nodes at the edge of the IoT network for execution. However, the timeliness and correctness of edge nodes may not be trusted (such as during high network load or attack). In this paper, we look at the applicability of trust for successful offloading. Traditionally, trust is computed at the application level, with suitable mechanisms to adjust for factors such as recency. However, these do not work well in IoT networks due to resource constraints. We propose a novel device called Trust Tracker (denoted by Σ) that provides higher-level applications with up-to-date trust information of the resource-rich nodes. We prove impossibility results regarding computation offloading and show that Σ is necessary and sufficient for correct offloading. We show that, Σ cannot be implemented even in a synchronous network and we compute the probability of offloading to a bad node, which we show to be negligible when a majority of nodes are correct. We perform a small-scale deployment to demonstrate our approach. Matthew Bradbury, Arshad Jhumka, Tim Watson |
INFOCOM | 2 |
| 2021 | Challenges in Identifying Network Attacks Using Netflow DataabstractLarge networks often encounter attacks that can affect the network availability. While multiple techniques exist to detect network attacks, a comprehensive understanding of how an attack occurs considering the various layers and components of the network software stack, can be an important element to help improve network security. By performing correlation analysis on contemporary unlabeled Netflow data, this paper conducts a comprehensive study of network flow events to identify communication patterns that may precede an attack, thereby providing potentially useful attack signatures to network administrators. Our work shows that, surprisingly, the Netflow data is not strongly correlated to network attacks. We observe that while spoof requests trigger reflection attacks, only a small percentage of the network packets are associated with the attack. Furthermore, lead time enhancements are feasible for reflection attacks that show long dwell times. Our study on network event correlations highlights empirical observations that could facilitate better attack handling in large networks. Edward Chuah, Neeraj Suri, Arshad Jhumka, Samantha Alt |
NCA | 3 |
| 2021 | Fault-Tolerant Ant Colony Based-Routing in Many-to-Many IoT Sensor NetworksabstractWireless IoT Sensor Networks have a wide range of applications in many areas of modern life, including environmental monitoring where data is sent to a sink. In such networks, it is likely for node failures to occur during the course of normal operation, e.g., when nodes run out of battery power or they have crashed due to defective hardware. Increasingly sophisticated applications, such as fire sprinkler systems, however deploy multiple sources and multiple sinks, in what is called many-many IoT networks. For these critical applications, it is necessary to develop a fault-tolerant routing protocol that is able to route messages around failed nodes, without a significant overhead. Focusing on many-many IoT networks, we present a novel distributed fault-tolerant routing protocol for wireless IoT sensor networks based on ant colony optimisation, that is able to route from multiple sources to multiple sinks. Our results show that our protocol is able to achieve more than 80% delivery ratio with 5% node failures. Our approach is scalable, compared to several approaches that require periodic topology maintenance to work. Jasmine Grosso, Arshad Jhumka |
NCA | 2 |
| 2021 | Secure Allocation for Graph-Based Virtual Machines in Cloud EnvironmentsabstractCloud computing systems (CCSs) enable the sharing of physical computing resources through virtualisation, where a group of virtual machines (VMs) can share the same physical resources of a given machine. However, this sharing can lead to a so-called side-channel attack (SCA), widely recognised as a potential threat to CCSs. Specifically, malicious VMs can capture information from (target) VMs, i.e., those with sensitive information, by merely co-located with them on the same physical machine. As such, a VM allocation algorithm needs to be cognizant of this issue and attempts to allocate the malicious and target VMs onto different machines, i.e., the allocation algorithm needs to be security-aware. This paper investigates the allocation patterns of VM allocation algorithms that are more likely to lead to a secure allocation. A driving objective is to reduce the number of VM migrations during allocation. We also propose a graph-based secure VMs allocation algorithm (GbSRS) to minimise SCA threats. Our results show that algorithms following a stacking-based behaviour are more likely to produce secure VMs allocation than those following spreading or random behaviours. Mansour Aldawood, Arshad Jhumka |
PST | 2 |
| 2021 | Reliable Logging in Wireless IoT Networks in the Presence of Byzantine FaultsabstractWireless IoT networks are becoming increasingly prevalent for sensing and actuating in a range of sophisticated applications, such as smart homes and smart cities. For various reasons, including auditability, debugging and performance tracking, event logging is essential in such networks. However, most current research has assumed that nodes always behave correctly. Yet, nodes in open IoT networks may misbehave due to the presence of attackers. Hence, an understanding of the problem of event logging in the presence of Byzantine nodes is essential. In this paper, we propose two versions of event logging: strong and weak logging. We show a number of impossibility results and propose a distributed and centralised version of the weak event logging. We show that only a weak logging can be achieved probabilistically, and we propose, and prove correct, an algorithm to solve the problem. Sara Alhajaili, Arshad Jhumka |
TrustCom | 2 |
| 2020 | A Spatial Source Location Privacy-aware Duty Cycle for Internet of Things Sensor NetworksabstractSource Location Privacy (SLP) is an important property for monitoring assets in privacy-critical sensor network and Internet of Things applications. Many SLP-aware routing techniques exist, with most striking a tradeoff between SLP and other key metrics such as energy (due to battery power). Typically, the number of messages sent has been used as a proxy for the energy consumed. Existing work (for SLP against a local attacker) does not consider the impact of sleeping via duty cycling to reduce the energy cost of an SLP-aware routing protocol. Therefore, two main challenges exist: (i) how to achieve a low duty cycle without loss of control messages that configure the SLP protocol and (ii) how to achieve high SLP without requiring a long time spent awake. In this article, we present a novel formalisation of a duty cycling protocol as a transformation process. Using derived transformation rules, we present the first duty cycling protocol for an SLP-aware routing protocol for a local eavesdropping attacker . Simulation results on grids demonstrate a duty cycle of 10%, while only increasing the capture ratio of the source by 3 percentage points, and testbed experiments on FlockLab demonstrate an 80% reduction in the average current draw. Matthew Bradbury, Arshad Jhumka, Carsten Maple |
ACM Trans. Internet Things | 2 |
| 2019 | Scheduling Dependent Tasks in Edge NetworksabstractIn this paper, we focus on the problem of offloading of jobs composed of dependent tasks in a mobile edge network (MEN). We formalise the problem and develop a heuristic lower the completion time of a job executing in MEN than when executing on a mobile device. We show the viability of the heuristic through our simulation results. Mohammed Maray, Arshad Jhumka, Adam P. Chester, Mohamed F. Younis |
IPCCC | 2 |
| 2019 | Auditability: An Approach to Ease Debugging of Reliable Distributed SystemsabstractThere are various means of providing dependability, one of which is fault removal and debugging is one popular example of a fault removal technique. During debugging, the system is executed, and execution data is collected, to be analysed later, to determine if the execution has satisfied the system specification. In a distributed system, the data collection is done centrally via a monitor, and the assumption typically is that all the nodes are correct. However, in open distributed systems such as the Internet of Things (IoT), there is no central authority to enforce this assumption and nodes may behave arbitrarily by violating protocol steps, making processes such as debugging very challenging. We call the data collection process for such processes auditing and a program that reliably records such data as being auditable. In this context, we make the following novel contributions towards auditability enforcement: (i) we define the auditability problem and (ii) identify a necessary condition for a program to be auditable. We then provide examples of auditable programs. Subsequently, (iii) we show an impossibility result for strong auditability. To circumvent this impossibility, we study a weaker problem and discuss the ramifications of certain implementations. Finally, we show that auditability is at least as difficult as the problem of fair exchange. This is the first formal work towards the design of reliable systems through audibility. Sara Alhajaili, Arshad Jhumka |
PRDC | 2 |
| 2019 | Phantom walkabouts: A customisable source location privacy aware routing protocol for wireless sensor networksabstractSummary Source location privacy (SLP) is an important property for a large class of security‐critical wireless sensor network (WSN) applications such as monitoring and tracking. In the seminal work on SLP, phantom routing was proposed as a viable approach to address SLP. However, recent work has shown some limitations of phantom routing such as poor data yield and low SLP. In this paper, we propose phantom walkabouts, a novel and more general version of phantom routing, which performs phantom routes of variable lengths. Through extensive simulations, we show that phantom walkabouts provides high SLP level than phantom routing under specific network configuration. Chen Gu, Matthew Bradbury, Arshad Jhumka |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Automation of fault-tolerant graceful degradation
Yiyan Lin, Sandeep S. Kulkarni, Arshad Jhumka |
Distributed Comput. | 3 |
| 2019 | Towards comprehensive dependability-driven resource use and message log-analysis for HPC systems diagnosis
Edward Chuah, Arshad Jhumka, Samantha Alt, Daniel Balouek-Thomert, James C. Browne, Manish Parashar |
J. Parallel Distributed Comput. | 2 |
| 2018 | Towards optimal source location privacy-aware TDMA schedules in wireless sensor networks
Jack Kirton, Matthew Bradbury, Arshad Jhumka |
Comput. Networks | 3 |
| 2018 | A decision theoretic framework for selecting source location privacy aware routing protocols in wireless sensor networksabstractSource location privacy (SLP) is becoming an important property for a large class of security-critical wireless sensor network applications such as monitoring and tracking. Many routing protocols have been proposed that provide SLP, all of which provide a trade-off between SLP and energy. Experiments have been conducted to gauge the performance of the proposed protocols under different network parameters such as noise levels . As that there exists a plethora of protocols which contain a set of possibly conflicting performance attributes, it is difficult to select the SLP protocol that will provide the best trade-offs across them for a given application with specific requirements. In this paper, we propose a methodology where SLP protocols are first profiled to capture their performance under various protocol configurations. Then, we present a novel decision theoretic procedure for selecting the most appropriate SLP routing algorithm for the application and network under investigation. We show the viability of our approach through different case studies . Chen Gu, Matthew Bradbury, Jack Kirton, Arshad Jhumka |
Future Gener. Comput. Syst. | 4 |
| 2018 | Hybrid online protocols for source location privacy in wireless sensor networksabstractWireless sensor networks (WSNs) will form the building blocks of many novel applications such as asset monitoring. These applications will have to guarantee that the location of the occurrence of specific events is kept private from attackers, in what is called the source location privacy (SLP) problem. Fake sources have been used in numerous techniques, however, the solution’s efficiency is typically achieved by fine-tuning parameters at compile time. This is undesirable as WSN conditions may change. In this paper, we first present an SLP algorithm – Dynamic – that estimates the relevant parameters at runtime and show that it provides a high level of SLP, albeit at the expense of a high number of messages. To address this, we provide a hybrid online algorithm – DynamicSPR – that uses directed random walks for the fake sources allocation strategy to reduce energy usage. We perform simulations of the various protocols we present and our results show that DynamicSPR provides a similar level of SLP as when parameters are optimised at compile-time, with a lower number of messages sent. Matthew Bradbury, Arshad Jhumka, Matthew Leeke |
J. Parallel Distributed Comput. | 2 |
| 2017 | Enabling Dependability-Driven Resource Use and Message Log-Analysis for Cluster System DiagnosisabstractRecent work have used both failure logs and resource use data separately (and together) to detect system failure-inducing errors and to diagnose system failures. System failure occurs as a result of error propagation and the (unsuccessful) execution of error recovery mechanisms. Knowledge of error propagation patterns and unsuccessful error recovery is important for more accurate and detailed failure diagnosis, and knowledge of recovery protocols deployment is important for improving system reliability. This paper presents the CORRMEXT framework which carries failure diagnosis another significant step forward by analyzing and reporting error propagation patterns and degrees of success and failure of error recovery protocols. CORRMEXT uses both error messages and resource use data in its analyses. Application of CORRMEXT to data from the Ranger supercomputer have produced new insights. CORRMEXT has: (i) identified correlations between resource use counters that capture recovery attempts after an error, (ii) identified correlations between error events to capture error propagation patterns within the system, (iii) identified error propagation and recovery paths during system execution to explain system behaviour, (iv) showed that the earliest times of change in system behaviour can only be identified by analyzing both the correlated resource use counters and correlated errors. CORRMEXT will be installed on the HPC clusters at the Texas Advanced Computing Center in Autumn 2017. Edward Chuah, Arshad Jhumka, Samantha Alt, Theodoros Damoulas, Nentawe Gurumdimma, Marie-Christine Sawley, William L. Barth, Tommy Minyard, James C. Browne |
HiPC | 2 |
| 2017 | Source Location Privacy-Aware Data Aggregation Scheduling for Wireless Sensor NetworksabstractSource location privacy (SLP) is an important property for the class of asset monitoring problems in wireless sensor networks (WSNs). SLP aims to prevent an attacker from finding a valuable asset when a WSN node is broadcasting information due to the detection of the asset. Most SLP techniques focus at the routing level, with typically high message overhead. The objective of this paper is to investigate the novel problem of developing a TDMA MAC schedule that can provide SLP. We make a number of important contributions: (i) we develop a novel formalisation of a class of eavesdropping attackers and provide novel formalisations of SLP-aware data aggregation schedules (DAS), (ii) we present a decision procedure to verify whether a DAS schedule is SLP-aware, that returns a counterexample if the schedule is not, similar to model checking, and (iii) we develop a 3-stage distributed algorithm that transforms an initial DAS algorithm into a corresponding SLP-aware schedule against a specific class of eavesdroppers. Our simulation results show that the resulting SLP-aware DAS protocol reduces the capture ratio by 50% at the expense of negligable message overhead. Jack Kirton, Matthew Bradbury, Arshad Jhumka |
ICDCS | 3 |
| 2017 | Understanding source location privacy protocols in sensor networks via perturbation of time seriesabstractSource location privacy (SLP) is becoming an important property for a large class of security-critical wireless sensor network applications such as monitoring and tracking. Much of the previous work on SLP has focused on the development of various protocols to enhance the level of SLP imparted to the network, under various attacker models and other conditions. Other work has focused on analysing the level of SLP being imparted by a specific protocol. In this paper, we adopt a different approach where we model the attacker movement as a time series and use information theoretic concepts to infer the properties of a routing protocol that imparts high levels of SLP. We propose the notion of a properly competing path that causes an attacker to “stall” when moving towards the source. This concept provides the basis for developing a perturbation model, similar to those in privacy-preserving data mining. We then show how to use properly competing paths to develop properties of an SLP-aware routing protocol. Further, we show how different SLP-aware routing protocols can be obtained through different instantiations of the framework. Those instantiations are obtained based on a notion of information loss achieved through the use of the perturbation model proposed. Matthew Bradbury, Arshad Jhumka |
INFOCOM | 2 |
| 2017 | Detection of Recovery Patterns in Cluster Systems Using Resource Usage DataabstractThe failure of large-scale distributed systems such as cluster systems has adverse effects on the performance of high-performance computing applications such as scientific applications. Techniques to handle these failures, such as checkpointing, typically incur a prohibitively high computational cost. To reduce or prevent the occurrences of such failures, system administrators have employed a divide and conquer approach to diagnosing the root-cause of such failures, in order to take corrective or preventive measures. Most times, event logs are the main sources of information about the failures. However, it is also important to be able to predict when the system is recovering to avoid such costly error handling. To this end, we present a novel technique, based on system resource usage information, to detect recovery runs. Our approach uses an unsupervised learning technique, namely change point detection, to predict recovery. We run our approach on data from Ranger Supercomputer System and the results are positive: our approach have an F-measure of 64%. Nentawe Gurumdimma, Arshad Jhumka |
PRDC | 2 |
| 2017 | Many-to-many data aggregation scheduling in wireless sensor networks with two sinks
Sain Saginbekov, Arshad Jhumka |
Comput. Networks | 2 |
| 2016 | Using Message Logs and Resource Use Data for Cluster Failure DiagnosisabstractFailure diagnosis for large compute clusters using only message logs is known to be incomplete. Recent availability of resource use data provides another potentially useful source of data for failure detection and diagnosis. Early work combining message logs and resource use data for failure diagnosis has shown promising results. This paper describes the CRUMEL framework which implements a new approach to combining rationalized message logs and resource use data for failure diagnosis. CRUMEL identifies patterns of errors and resource use and correlates these patterns by time with system failures. Application of CRUMEL to data from the Ranger supercomputer has yielded improved diagnoses over previous research. CRUMEL has: (i) showed that more events correlated with system failures can only be identified by applying different correlation algorithms, (ii) confirmed six groups of errors, (iii) identified Lustre I/O resource use counters which are correlated with occurrence of Lustre faults which are potential flags for online detection of failures, (iv) matched the dates of correlated error events and correlated resource use with the dates of compute node hang-ups and (v) identified two more error groups associated with compute node hang-ups. The pre-processed data will be put on the public domain in September, 2016. Edward Chuah, Arshad Jhumka, James C. Browne, Nentawe Gurumdimma, Sai Narasimhamurthy, William L. Barth |
HiPC | 2 |
| 2016 | CRUDE: Combining Resource Usage Data and Error Logs for Accurate Error Detection in Large-Scale Distributed SystemsabstractThe use of console logs for error detection in large scale distributed systems has proven to be useful to system administrators. However, such logs are typically redundant and incomplete, making accurate detection very difficult. In an attempt to increase this accuracy, we complement these incomplete console logs with resource usage data, which captures the resource utilisation of every job in the system. We then develop a novel error detection methodology, the CRUDE approach, that makes use of both the resource usage data and console logs. We thus make the following specific technical contributions: we develop (i) a clustering algorithm to group nodes with similar behaviour, (ii) an anomaly detection algorithm to identify jobs with anomalous resource usage, (iii) an algorithm that links jobs with anomalous resource usage with erroneous nodes. We then evaluate our approach using console logs and resource usage data from the Ranger Supercomputer. Our results are positive: (i) our approach detects errors with a true positive rate of about 80%, and (ii) when compared with the well-known Nodeinfo error detection algorithm, our algorithm provides an average improvement of around 85% over Nodeinfo, with a best-case improvement of 250%. Nentawe Gurumdimma, Arshad Jhumka, Maria Liakata, Edward Chuah, James C. Browne |
SRDS | 2 |
| 2016 | A Multidimension Taxonomy of Insider Threats in Cloud ComputingabstractSecurity is considered a significant deficiency in cloud computing, and insider threats problem exacerbate security concerns in the cloud. In addition to that, cloud computing is very complex by itself, because it encompasses numerous technologies and concepts. Apparently, overcoming these challenges requires substantial efforts from information security researchers to develop powerful mitigation solutions for this emerging problem. This entails developing a taxonomy of insider threats in cloud environments encompassing all potential abnormal activities in the cloud and can be useful for conducting security assessment. This article describes the first phase of an ongoing research to develop a framework for mitigating insider threats in cloud computing environments. Primarily, it presents a multidimensional taxonomy of insider threats in cloud computing and demonstrates its viability. The taxonomy provides a fundamental understanding for this complicated problem by identifying five dimensions; it also supports security engineers in identifying hidden paths, thus determining proper countermeasures, and presents a guidance that covers all bounders of insiders’ threats issue in clouds; hence, it facilitates researchers’ endeavours in tackling this problem. For instance, according to the hierarchical taxonomy, clearly many significant issues exist in public cloud, while conventional insider mitigation solutions can be used in private clouds. Finally, the taxonomy assists in identifying future research directions in this emerging area. Mohannad Alhanahnah, Arshad Jhumka, Sahel Alouneh |
Comput. J. | 2 |
| 2016 | Neighborhood View Consistency in Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are characterized by localized interactions , that is, protocols are often based on message exchanges within a node’s direct radio range. We recognize that for these protocols to work effectively, nodes must have consistent information about their shared neighborhoods. Different types of faults, however, can affect this information, severely impacting a protocol’s performance. We factor this problem out of existing WSN protocols and argue that a notion of neighborhood view consistency (NVC) can be embedded within existing designs to improve their performance. To this end, we study the problem from both a theoretical and a system perspective. We prove that the problem cannot be solved in an asynchronous system using any of Chandra and Toueg’s failure detectors. Because of this, we introduce a new software device called pseudocrash failure detector (PCD), study its properties, and identify necessary and sufficient conditions for solving NVC with PCDs. We prove that, in the presence of transient faults, NVC is impossible to solve with any PCDs, thus define two weaker specifications of the problem. We develop a global algorithm that satisfies both specifications in the presence of unidirectional links, and a localized algorithm that solves the weakest specification in networks of bidirectional links. We implement the latter atop two different WSN operating systems, integrate our implementations with four different WSN protocols, and run extensive micro-benchmarks and full-stack experiments on a real 90-node WSN testbed. Our results show that the performance significantly improves for NVC-equipped protocols; for example, the Collection Tree Protocol (CTP) halves energy consumption with higher data delivery. Arshad Jhumka, Luca Mottola |
ACM Trans. Sens. Networks | 1 |
| 2015 | An Investigation of the Impact of Double Bit-Flip Error Variants on Program Execution
Fatimah Adamu-Fika, Arshad Jhumka |
ICA3PP (4) | 2 |
| 2015 | Assessing the Performance of Phantom Routing on Source Location Privacy in Wireless Sensor NetworksabstractAs wireless sensor networks (WSNs) have been applied across a spectrum of application domains, the problem of source location privacy (SLP) has emerged as a significant issue, particularly in safety-critical situations. In seminal work on SLP, phantom routing was proposed as an approach to addressing the issue. However, results presented in support of phantom routing have not included considerations for practical network configurations, omitting simulations and analyses with larger network sizes. This paper addresses this shortcoming by conducting an in-depth investigation of phantom routing under various network configurations. The results presented demonstrate that previous work in phantom routing does not generalise well to different network configurations. Specifically, under certain configurations, it is shown that the afforded SLP is reduced by a factor of up to 75. Chen Gu, Matthew Bradbury, Arshad Jhumka, Matthew Leeke |
PRDC | 3 |
| 2015 | Fake source-based source location privacy in wireless sensor networksabstractSummary The development of novel wireless sensor network (WSN) applications, such as asset monitoring, has led to novel reliability requirements. One such property is source location privacy (SLP). The original SLP problem is to protect the location of a source node in a WSN from a singledistributed eavesdropperattacker. Several techniques have been proposed to address the SLP problem, and most of them use some form of traffic analysis and engineering to provide enhanced SLP. The use of fake sources is considered to be promising for providing SLP, and several works have investigated the effectiveness of the fake sources approach under various attacker models. However, very little work has been done to understand the theoretical underpinnings of the fake source technique. In this paper, we (i) provide a novel formalisation of the fake sources selection problem; (ii) prove the fake sources selection problem to be NP‐complete; (iii) provide parametric heuristics for three different network configurations; and (iv) show that these heuristics provide (near) optimal levels of SLP under appropriate parameterisation. Our results show that fake sources can provide a high level of SLP. Our work is the first to investigate the theoretical underpinnings of the fake source technique. Copyright © 2014 John Wiley & Sons, Ltd. Arshad Jhumka, Matthew Bradbury, Matthew Leeke |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Towards Efficient Stabilizing Code Dissemination in Wireless Sensor NetworksabstractOne important component of network reprogramming is code dissemination (CD), when the updated program code is distributed to the relevant nodes. Very few CD protocols tolerate transient faults that corrupt the state and these faults can cause the old code to disseminate in the network. We propose two protocols called BestEffort-Repair and Consistent-Repair that transform fault-intolerant CD protocols into non-masking fault-tolerant protocols where, eventually, all nodes obtain the new code. We conduct experiments with both protocols on TelosB-like motes and over TOSSIM simulations to show their correctness and also their performance. We conduct a case study whereby both protocols are added to a state-of-the-art CD protocol, namely Varuna to evaluate their impact on Varuna. Our results show that (i) Varuna, which is fault-intolerant, is transformed into a stabilizing CD protocol; (ii) they induce low overhead on Varuna, and cause all nodes to eventually receive the new code. BestEffort-Repair is biased towards fast recovery, whereas Consistent-Repair attempts to reduce the number of erroneous downloads in the network. Our main contribution is the first corrector protocols that correct CD in the presence of transient faults. Sain Saginbekov, Arshad Jhumka |
Comput. J. | 2 |
| 2014 | Efficient code dissemination in wireless sensor networks
Sain Saginbekov, Arshad Jhumka |
Future Gener. Comput. Syst. | 2 |
| 2014 | Efficient fault-tolerant collision-free data aggregation scheduling for wireless sensor networks
Arshad Jhumka, Matthew Bradbury, Sain Saginbekov |
J. Parallel Distributed Comput. | 1 |
| 2013 | Linking Resource Usage Anomalies with System Failures from Cluster Log DataabstractBursts of abnormally high use of resources are thought to be an indirect cause of failures in large cluster systems, but little work has systematically investigated the role of high resource usage on system failures, largely due to the lack of a comprehensive resource monitoring tool which resolves resource use by job and node. The recently developed TACC_Stats resource use monitor provides the required resource use data. This paper presents the ANCOR diagnostics system that applies TACC_Stats data to identify resource use anomalies and applies log analysis to link resource use anomalies with system failures. Application of ANCOR to first identify multiple sources of resource anomalies on the Ranger supercomputer, then correlate them with failures recorded in the message logs and diagnosing the cause of the failures, has identified four new causes of compute node soft lockups. ANCOR can be adapted to any system that uses a resource use monitor which resolves resource use by job. Edward Chuah, Arshad Jhumka, Sai Narasimhamurthy, John L. Hammond, James C. Browne, William L. Barth |
SRDS | 2 |
| 2013 | Manipulating convention emergence using influencer agents
Henry Franks, Nathan Griffiths, Arshad Jhumka |
Auton. Agents Multi Agent Syst. | 3 |
| 2013 | Towards the Design of Efficient Error Detection Mechanisms for Transient Data ErrorsabstractA dependable software system must contain two dependability components: (i) error detection mechanisms (EDMs) and (ii) error recovery mechanisms. Currently, EDMs are generally designed based on some system specification or based on the experience of software engineers, with their efficiency typically being measured using fault injection and software measures such as coverage and latency. In contrast to finite-state programs, for which efficient EDMs can be obtained by design, no systematic design approach exists for real-world software systems. In this paper, we bridge this gap by developing an approach for the design of highly efficient error detection predicates for EDMs for such software systems. Our approach is based on the use of data mining techniques to classify states as safe or failure-inducing. The results presented, under a transient data value fault model, demonstrate the viability of the approach for the development of efficient EDMs, as the EDMs generated yield a true positive rate of nearly 100% and a false positive rate close to 0% for the detection of failure-inducing states. Matthew Leeke, Arshad Jhumka, Sarabjot S. Anand |
Comput. J. | 2 |
| 2012 | Towards Understanding Source Location Privacy in Wireless Sensor Networks through Fake SourcesabstractSource location privacy is becoming an increasingly important property in wireless sensor network applications, such as asset monitoring. The original source location problem is to protect the location of a source in a wireless sensor network from a single distributed eavesdropper attack. Several techniques have been proposed to address the source location problem, where most of these apply some form of traffic analysis and engineering to provide enhanced privacy. One such technique, namely fake sources, has proved to be promising for providing source location privacy. Recent research has concentrated on investigating the efficiency of fake source approaches under various attacker models. In this paper, we (i) provide a novel formalisation of the source location privacy problem, (ii) prove the source location privacy problem to be NP-complete, and (iii) provide a heuristic that yields an optimal level of privacy under appropriate parameterisation. Crucially, the results presented show that fake sources can provide a high, sometimes optimal, level of privacy. Arshad Jhumka, Matthew Bradbury, Matthew Leeke |
TrustCom | 1 |
| 2012 | Fast and Efficient Information Dissemination in Event-Based Wireless Sensor NetworksabstractGiven the dynamic nature of the mission of a wireless sensor network (WSN), and of the environment in which it is usually deployed, network reprogramming is an important activity that enables the WSN to adapt to the mission and/or environment. One important component of a reprogramming protocol is code dissemination and maintenance, during which new code is propagated to relevant WSN nodes. Several dissemination protocols have been proposed, each with a specific objective in mind. Protocols such as Trickle minimise dissemination latency by periodically broadcasting advertisement messages at the expense of energy consumption, while protocols, e.g., Varuna, reduce energy usage by broadcasting advertisement only when needed. In certain type of WSNs, such as event-based WSNs, Varuna has high code dissemination latency, while the energy consumption of Trickle does not improve in such WSNs. In this paper, we propose a new code dissemination protocol, called Triva, for event based WSNs, by leveraging the properties of Trickle and Varuna. Our results show that, for event-based WSNs, Triva outperforms Trickle and Varuna in terms of energy consumption and code dissemination latency respectively. We also show that Triva provides excellent results during bursty traffic in event-based WSNs. Triva is the first information dissemination protocol for event-based WSNs. Sain Saginbekov, Arshad Jhumka |
TrustCom | 2 |
| 2011 | A methodology for the generation of efficient error detection mechanismsabstractA dependable software system must contain error detection mechanisms and error recovery mechanisms. Software components for the detection of errors are typically designed based on a system specification or the experience of software engineers, with their efficiency typically being measured using fault injection and metrics such as coverage and latency. In this paper, we introduce a methodology for the design of highly efficient error detection mechanisms. The proposed methodology combines fault injection analysis and data mining techniques in order to generate predicates for efficient error detection mechanisms. The results presented demonstrate the viability of the methodology as an approach for the development of efficient error detection mechanisms, as the predicates generated yield a true positive rate of almost 100% and a false positive rate very close to 0% for the detection of failure-inducing states. The main advantage of the proposed methodology over current state-of-the-art approaches is that efficient detectors are obtained by design, rather than by using specification-based detector design or the experience of software engineers. Matthew Leeke, Saima Arif, Arshad Jhumka, Sarabjot S. Anand |
DSN | 3 |
| 2011 | The Early Identification of Detector Locations in Dependable SoftwareabstractThe dependability properties of a software system are usually assessed and refined towards the end of the software development lifecycle. Problems pertaining to software dependability may necessitate costly system redesign. Hence, early insights into the potential for error propagation within a software system would be beneficial. Further, the refinement of the dependability properties of software involves the design and location of dependability components called detectors and correctors. Recently, a metric, called spatial impact, has been proposed to capture the extent of error propagation in a software system, providing insights into the location of detectors and correctors. However, the metric only provides insights towards the end of the software development life cycle. In this paper, our objective is to investigate whether spatial impact can enable the early identification of locations for detectors. To achieve this we first hypothesise that spatial impact is correlated with module coupling, a metric that can be evaluated early in the software development life cycle, and show this relationship to hold. We then evaluate module coupling for the modules of a complex software system, identifying modules with high coupling values as potential locations for detectors. We then enhanced these modules with detectors and perform fault-injection analysis to determine the suitability of these locations. The results presented demonstrate that our approach can permit the early identification of possible detector locations. Arshad Jhumka, Matthew Leeke |
ISSRE | 1 |
| 2011 | On the Use of Fake Sources for Source Location Privacy: Trade-Offs Between Energy and PrivacyabstractWireless sensor networks have enabled novel applications such as monitoring, where security is invariably a requirement. One aspect of security, namely source location privacy, is becoming an increasingly important property of some wireless sensor network applications. The fake source technique has been proposed as an efficient technique to handle the source location privacy problem. However, there are several factors that limit the usefulness of current results: (i) the selection of fake sources is dependent on sophisticated nodes, (ii) fake sources are known a priori and (iii) the selection of fake sources is based on a prohibitively expensive pre-configuration phase. In this paper, we investigate the privacy enhancement and energy efficiency of different implementations of the fake source technique that circumvents these limitations. Our results show that the fake source technique is indeed effective in enhancing privacy. Specifically, one implementation achieves near-perfect privacy when there is at least one fake source in the network, at the expense of increased energy consumption. In the presence of multiple attackers, the same implementation yields only a 30% decrease in capture ratio with respect to flooding. To address this problem, we propose a hybrid technique which achieves a corresponding 50% reduction in the capture ratio and a near-perfect privacy whenever at least one fake source exists in the network. Arshad Jhumka, Matthew Leeke, Sambid Shrestha |
Comput. J. | 1 |
| 2010 | Crash-Tolerant Collision-Free Data Aggregation Scheduling for Wireless Sensor NetworksabstractData aggregation scheduling, or converge cast, is a fundamental pattern of communication in wireless sensor networks (WSNs), where sensor nodes aggregate and relay data to a sink node. For WSN applications that require fast response time, it is imperative that the data reaches the sink as fast as possible. For such timeliness guarantees, TDMA-based scheduling can be used to assign time slots to nodes in which they can transmit messages. However, any slot assignment approach needs to be cognisant of the fact that crash failures can occur (e.g., due to battery exhaustion, defective hardware). In this paper, we study the design of such data aggregation scheduling (converge cast) protocols. We make the following contributions: (i) we identify a necessary condition to solve the converge cast problem, (ii) we introduce two versions of the converge cast problem, namely (a) a strong version, and (b) a weak version , (iii) we show that the strong converge cast problem cannot be solved, (iv) we show that deterministic weak converge cast cannot be solved in presence of crash failures, (v) we show that there is no d-local algorithm that solves stabilising weak converge cast in presence of crash failures, (vi) we provide a modular d-local algorithm that solves stabilising weak converge cast in presence of crash failures where d is the network radius, and (vii) we show how specific instantiations of parameters can lead to an d-local algorithm that achieves more efficient stabilization. Our contributions are novel: (i) the first contribution (necessary condition) provides the theoretical basis which explains the structure of existing converge cast algorithms, and (ii) the study of converge cast in presence of crash failures has not previously been studied. Arshad Jhumka |
SRDS | 1 |
| 2010 | A software integration approach for designing and assessing dependable embedded systems
Neeraj Suri, Arshad Jhumka, Martin Hiller, András Pataricza, Shariful Islam, Constantin Sârbu |
J. Syst. Softw. | 2 |
| 2009 | Issues on the Design of Efficient Fail-Safe Fault ToleranceabstractThe design of a fault-tolerant program is known to be an inherently difficult task. Decisions taken during the design process will invariably have an impact on the efficiency of the resulting fault-tolerant program. In this paper, we focus on two such decisions, namely (i) the class of faults the program is to tolerate, and (ii) the variables that can be read and written. The impact these design issues have on the overall fault tolerance of the system needs to be well-understood, failure of which can lead to costly redesigns. For the case of understanding the impact of fault classes on the efficiency of fail-safe fault tolerance, we show that, under the assumption of a general fault model, it is impossible to preserve the original behavior of the fault-intolerant program. For the second problem of read and write constraints of variables, we again show that it is impossible to preserve the original behavior of the fault-intolerant program. We analyze the reasons that lead to these impossibility results, and suggest possible ways of circumventing them. Arshad Jhumka, Matthew Leeke |
ISSRE | 1 |
| 2009 | Evaluating the Use of Reference Run Models in Fault Injection AnalysisabstractFault injection (FI) has been shown to be an effective approach to assessing the dependability of software systems. To determine the impact of faults injected during FI, a given oracle is needed. Oracles can take a variety of forms, including (i) specifications, (ii) error detection mechanisms and (iii) golden runs. Focusing on golden runs, in this paper we show that there are classes of software which a golden run based approach can not be used to analyse. Specifically, we demonstrate that a golden run based approach can not be used in the analysis of systems which employ a main control loop with an irregular period. Further, we show how a simple model, which has been refined using FI experiments, can be employed as an oracle in the analysis of such a system. Matthew Leeke, Arshad Jhumka |
PRDC | 2 |
| 2009 | On Consistent Neighborhood Views in Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are characterized by localized interactions. Indeed, several WSN algorithms and protocols work in a decentralized fashion by coordinating nodes within the wireless communication range, e.g., localization algorithms and MAC protocols. Nevertheless, most often these mechanisms do not address faults that may affect the way wireless neighborhoods are recognized by nodes, e.g., as in the case of data corruption. As the operation of these mechanisms is rooted in the use of topology information, these faults may be a significant detriment to correct and efficient system operation.In this paper, we argue that the above issues are particular instances of a general problem of consistent neighborhood view. We present three increasingly weaker specifications of the problem. Next, we prove the impossibility of solving the two stronger specifications, and provide an algorithm to solve the weakest specification. In addition, we implement our algorithm in a commonly used WSN network stack, and assess its performance both in simulation and in a real-world testbed. The results show that, when possible, our mechanisms efficiently solve the problem of consistent neighborhood view, providing higher-level mechanisms with a re-usable building block to leverage off. Arshad Jhumka, Luca Mottola |
SRDS | 1 |
| 2007 | Global Predicate Detection in Distributed Systems with Small Faults
Felix C. Freiling, Arshad Jhumka |
SSS | 2 |
| 2005 | A Dependability-Driven System-Level Design Approach for Embedded SystemsabstractThe paper introduces dependability as an optimization criterion in the system-level design process of embedded systems. Given the pervasiveness of embedded systems, especially in the area of highly dependable and safety-critical systems, it is imperative to consider dependability in the system level design process directly. This naturally leads to a multi-objective optimization problem, as cost and time have to be considered too. The paper proposes a genetic algorithm to solve this multi-objective optimization problem and to determine a set of Pareto optimal design alternatives in a single optimization run. Based on these alternatives, the designer can choose his best solution, finding the desired tradeoff between cost, schedulability, and dependability. Arshad Jhumka, Stephan Klaus, Sorin A. Huss |
DATE | 1 |
| 2005 | Designing Efficient Fail-Safe Multitolerant Systems
Arshad Jhumka, Neeraj Suri |
FORTE | 1 |
| 2005 | Putting Detectors in Their PlaceabstractIn this paper, we address the problem of locating detectors in a given program under resource constraints. A detector is a program component that asserts the validity of a predicate in a program. The detector location problem is to identify which program actions need to be monitored by detectors such that certain given dependability properties are met. In this paper, we focus on the following dependability properties: (i) high detection coverage, (ii) low detection latency, and (iii) low false alarms rate. Our main contributions are: (i) we first provide a formal definition of the detector location problem under resource constraints, and (ii) we subsequently show that the problem is NP-complete, (iii) we investigate a special case of the detector location problem that can be solved in polynomial time, and present a sound and complete algorithm that solves the problem. We present an example to show the applicability of our approach, which is intended in the area of dependable embedded systems. Arshad Jhumka, Martin Hiller |
SEFM | 1 |
| 2004 | EPIC: Profiling the Propagation and Effect of Data Errors in SoftwareabstractWe present an approach for analyzing the propagation and effect of data errors in modular software enabling the profiling of the vulnerabilities of software to find 1) the modules and signals most likely exposed to propagating errors and 2) the modules and signals which, when subjected to error, tend to cause more damage than others from a systems operation point-of-view. We discuss how to use the obtained profiles to identify where dependability structures and mechanisms will likely be the most effective, i.e., how to perform a cost-benefit analysis for dependability. A fault-injection-based method for estimation of the various measures is described and the software of a real embedded control system is profiled to show the type of results obtainable by the analysis framework. Martin Hiller, Arshad Jhumka, Neeraj Suri |
IEEE Trans. Computers | 2 |
| 2003 | A Framework for the Design and Validation of Efficient Fail-Safe Fault-Tolerant Programs
Arshad Jhumka, Neeraj Suri, Martin Hiller |
SCOPES | 1 |
| 2002 | On the Placement of Software Mechanisms for Detection of Data ErrorsabstractAn important aspect in the development of dependable software is to decide where to locate mechanisms for efficient error detection and recovery. We present a comparison between two methods for selecting locations for error detection mechanisms, in this case executable assertions (EAs), in black-box, modular software. Our results show that by placing EAs based on error propagation analysis one may reduce the memory and execution time requirements as compared to experience- and heuristic-based placement while maintaining the obtained detection coverage. Further, we show the sensitivity of the EA-provided coverage estimation on the choice of the underlying error model. Subsequently, we extend the analysis framework such that error-model effects are also addressed and introduce measures for classifying signals according to their effect on system output when errors are present. The extended framework facilitates profiling of software systems from varied dependability perspectives and is also less susceptible to the effects of having different error models for estimating detection coverage. Martin Hiller, Arshad Jhumka, Neeraj Suri |
DSN | 2 |
| 2002 | PROPANE: an environment for examining the propagation of errors in softwareabstractIn order to produce reliable software, it is important to have knowledge on how faults and errors may affect the software. In particular, designing efficient error detection mechanisms requires not only knowledge on which types of errors to detect but also the effect these errors may have on the software as well as how they propagate through the software. This paper presents the Propagation Analysis Environment (PROPANE) which is a tool for profiling and conducting fault injection experiments on software running on desktop computers. PROPANE supports the injection of both software faults (by mutation of source code) and data errors (by manipulating variable and memory contents). PROPANE supports various error types out-of-the-box and has support for user-defined error types. For logging, probes are provided for charting the values of variables and memory areas as well as for registering events during execution of the system under test. PROPANE has a flexible design making it useful for development of a wide range of software systems, e.g., embedded software, generic software components, or user-level desktop applications. We show examples of results obtained using PROPANE and how these can guide software developers to where software error detection and recovery could increase the reliability of the software system. Martin Hiller, Arshad Jhumka, Neeraj Suri |
ISSTA | 2 |
| 2001 | An Approach for Analysing the Propagation of Data Errors in SoftwareabstractWe present a novel approach for analysing the propagation of data errors in software. The concept of error permeability is introduced as a basic measure upon which we define a set of related measures. These measures guide us in the process of analysing the vulnerability of software to find the modules that are most likely exposed to propagating errors. Based on the analysis performed with error permeability and its related measures, we describe how to select suitable locations for error detection mechanisms (EDMs) and error recovery mechanisms (ERMs). A method for experimental estimation of error permeability, based on fault injection, is described and the software of a real embedded control system analysed to show the type of results obtainable by the analysis framework. The results show that the developed framework is very useful for analysing error propagation and software vulnerability and for deciding where to place EDMs and ERMs. Martin Hiller, Arshad Jhumka, Neeraj Suri |
DSN | 2 |
| 2001 | Assessing Inter-Modular Error Propagation in Distributed SoftwareabstractWith the functionality of most embedded systems based on software (SW), interactions amongst SW modules arise, resulting in error propagation across them. During SW development, it would be helpful to have a framework that clearly demonstrates the error propagation and containment capabilities of the different SW components. In this paper, we assess the impact of inter-modular error propagation. Adopting a white-box SW approach, we make the following contributions: (a) we study and characterize the error propagation process and derive a set of metrics that quantitatively represents the inter-modular SW interactions, (b) we use a real embedded target system used in an aircraft arrestment system to perform fault-injection experiments to obtain experimental values for the metrics proposed, (c) we show how the set of metrics can be used to obtain the required analytical framework for error propagation analysis. We find that the derived analytical framework establishes a very close correlation between the analytical and experimental values obtained. The intent is to use this framework to be able to systematically develop SW such that inter-modular error propagation is reduced by design. Arshad Jhumka, Martin Hiller, Neeraj Suri |
SRDS | 1 |