EDBT 2026 Demo / reviewers in the wild / expert
Neeraj Suri
dblp:s/NeerajSuri
· DBLP profile ↗
158ranked-venue papers
7as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 66 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 35 · 2 first-author · 2 since 2021Systems, architecture and hardware · 33 · 4 first-author · 2 since 2021Computer networks · 16 · 4 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning Forecasting and GAN-Based Scenario Control for EV Charging and PV IntegrationabstractThe increasing integration of electric vehicles (EVs) and photovoltaic (PV) generation introduces significant uncertainty into modern distribution grids. This paper presents a dual-stage, AI-driven framework for resilient energy management that combines machine learning-based forecasting with generative modelling for scenario-based control. The first stage uses a hybrid forecasting architecture: long short-term memory (LSTM) and convolutional LSTM (ConvLSTM) models for EV demand prediction and eXtreme gradient boosting (XGBoost) for PV generation forecasting. The second stage employs a generative adversarial network (GAN) to produce realistic EV and PV scenarios, capturing both typical variability and a wide range of operating conditions. The framework is validated on a modified IEEE 33-bus distribution system with integrated EV charging and stationary storage. Results show that the dual-model forecasting approach achieves high accuracy across diverse temporal patterns, while GAN-based scenario generation improves the adaptability of control decisions. Scenario-based optimisation enhances performance under uncertainty, especially at high-variance nodes, and offers greater flexibility than deterministic control in balancing energy cost and demand satisfaction. Fatemeh Nasr Esfahani, Neeraj Suri, Xiandong Ma |
IECON | 2 |
| 2025 | Game Theory Empowered Carbon-Intelligent Federated Multiedge Caching for Industrial Internet of ThingsabstractTo navigate the carbon emission and functional challenges associated with edge caching within heterogeneous Industrial Internet of Things (IIoT) spanning energy use, cache hit rate, and bandwidth usage, this paper proposes a novel Game Theory Empowered Carbon-Intelligent Federated Multi-Edge Caching framework (GT-FMC). The proposed framework enables distributed collaborative caching by intelligently coordinating edge nodes to optimize content decisions while efficiently integrating content providers (CPs), edge nodes, and users with energy-aware strategies. In GT-FMC, a lightweight federated content popularity prediction method based on Temporal Convolutional Networks (TCN) is introduced to collaboratively learn global content popularity while reducing prediction energy cost. The energy-aware utilities of the three involved parties are jointly formulated as a coupled non-linear optimization problem. To address this challenge, a two-stage game-theoretic algorithm is designed. Experimental results on a real-world testbed show that GT-FMC achieves up to 77.9% of Oracle in cache hit rate and 10.6%–32.4% reduction in transmission energy consumption compared to baseline methods. Complementary evaluations also validate the game-theoretic design’s effectiveness. Zhengxin Yu, Haris Pervaiz, Guhan Zheng, Neeraj Suri |
IEEE Internet Things J. | 5 |
| 2024 | Self-supervised Representation Learning for Adversarial Attack Detection
Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
ECCV (60) | 3 |
| 2024 | Federated Adversarial Learning for Robust Autonomous Landing Runway Detection
Yi Li 0047, Plamen Angelov 0001, Zhengxin Yu, Alvaro Lopez Pellicer, Neeraj Suri |
ICANN (6) | 5 |
| 2024 | Rethinking Self-supervised Learning for Cross-domain Adversarial Sample RecoveryabstractAdversarial attacks can cause misclassification in machine learning pipelines, posing a significant safety risk in critical applications such as autonomous systems or medical applications. Supervised learning-based methods for adversarial sample recovery rely heavily on large volumes of labeled data, which often results in substantial performance degradation when applying the trained model to new domains. In this paper, differing from conventional self-supervised learning techniques such as data augmentation, we present a novel two-stage self-supervised representation learning framework for the task of adversarial sample recovery, aimed at overcoming these limitations. In the first stage, we employ a clean image autoencoder (CAE) to learn representations of clean images. Subsequently, the second stage utilizes an adversarial image autoencoder (AAE) to learn a shared latent space that captures the relationships between the representations acquired by CAE and AAE. It is noteworthy that the input clean images in the first stage and adversarial images in the second stage are cross-domain and not paired. To the best of our knowledge, this marks the first instance of self-supervised adversarial sample recovery work that operates without the need for labeled data. Our experimental evaluations, spanning a diverse range of images, consistently demonstrate the superior performance of the proposed method compared to conventional adversarial sample recovery methods. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IJCNN | 3 |
| 2024 | UNICAD: A Unified Approach for Attack Detection, Noise Reduction and Novel Class IdentificationabstractAs the use of Deep Neural Networks (DNNs) becomes pervasive, their vulnerability to adversarial attacks and limitations in handling unseen classes poses significant challenges. The state-of-the-art offers discrete solutions aimed to tackle individual issues covering specific adversarial attack scenarios, classification or evolving learning. However, real-world systems need to be able to detect and recover from a wide range of adversarial attacks without sacrificing classification accuracy and to flexibly act in unseen scenarios. In this paper, UNICAD, is proposed as a novel framework that integrates a variety of techniques to provide an adaptive solution.For the targeted image classification, UNICAD achieves accurate image classification, detects unseen classes, and recovers from adversarial attacks using Prototype and Similarity-based DNNs with denoising autoencoders. Our experiments performed on the CIFAR-10 dataset highlight UNICAD’s effectiveness in adversarial mitigation and unseen class classification, outperforming traditional models. Alvaro Lopez Pellicer, Kittipos Giatgong, Yi Li 0047, Neeraj Suri, Plamen Angelov 0001 |
IJCNN | 4 |
| 2024 | An empirical study of reflection attacks using NetFlow dataabstractAbstract Reflection attacks are one of the most intimidating threats organizations face. A reflection attack is a special type of distributed denial-of-service attack that amplifies the amount of malicious traffic by using reflectors and hides the identity of the attacker. Reflection attacks are known to be one of the most common causes of service disruption in large networks. Large networks perform extensive logging of NetFlow data, and parsing this data is an advocated basis for identifying network attacks. We conduct a comprehensive analysis of NetFlow data containing 1.7 billion NetFlow records and identified reflection attacks on the network time protocol (NTP) and NetBIOS servers. We set up three regression models including the Ridge, Elastic Net and LASSO. To the best of our knowledge, there is no work that studied different regression models to understand patterns of reflection attacks in a large network. In this paper, we (a) propose an approach for identifying correlations of reflection attacks, and (b) evaluate the three regression models on real NetFlow data. Our results show that (a) reflection attacks on the NTP servers are not correlated, (b) reflection attacks on the NetBIOS servers are not correlated, (c) the traffic generated by those reflection attacks did not overwhelm the NTP and NetBIOS servers, and (d) the dwell times of reflection attacks on the NTP and NetBIOS servers are too small for predicting reflection attacks on these servers. Our work on reflection attacks identification highlights recommendations that could facilitate better handling of reflection attacks in large networks. Edward Chuah, Neeraj Suri |
Cybersecur. | 2 |
| 2024 | Enabling Multi-Layer Threat Analysis in Dynamic Cloud EnvironmentsabstractMost Threat Analysis (TA) techniques analyze threats to targeted assets (e.g., components, services) by considering static interconnections among them. However, in dynamic environments, e.g., the Cloud, resources can instantiate, migrate across physical hosts, or decommission to provide rapid resource elasticity to its users. Existing TA techniques are not capable of addressing such requirements. Moreover, complex multi-layer/multi-asset attacks on Cloud systems are increasing, e.g., the Equifax data breach; thus, TA approaches must be able to analyze them. This paper proposes ThreatPro, which supports dynamic interconnections and analysis of multi-layer attacks in the Cloud. ThreatPro facilitates threat analysis by developing a technology-agnostic information flow model, representing the Cloud's functionality through conditional transitions. The model establishes the basis to capture the multi-layer and dynamic interconnections during the life cycle of a Virtual Machine. ThreatPro contributes to (1) enabling the exploration of a threat's behavior and its propagation across the Cloud, and (2) assessing the security of the Cloud by analyzing the impact of multiple threats across various operational layers/assets. Using public information on threats from the National Vulnerability Database, we validate ThreatPro's capabilities, i.e., identify and trace actual Cloud attacks and speculatively postulate alternate potential attack paths. Salman Manzoor, Antonios Gouglidis, Matthew Bradbury, Neeraj Suri |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | Adversarial Attack Detection via Fuzzy PredictionsabstractImage processing using neural networks act as a tool to speed up predictions for users, specifically on large-scale image samples. To guarantee the clean data for training accuracy, various deep learning-based adversarial attack detection techniques have been proposed. These crisp set-based detection methods directly determine whether an image is clean or attacked, while, calculating the loss is nondifferentiable and hinders training through normal back-propagation. Motivated by the recent success in fuzzy systems, in this work, we present an attack detection method to further improve detection performance, which is suitable for any pretrained neural network classifier. Subsequently, the fuzzification network is used to obtain feature maps to produce fuzzy sets of difference degree between clean and attacked images. The fuzzy rules control the intelligence that determines the detection boundaries. Different from previous fuzzy systems, we propose a fuzzy mean-intelligence mechanism with new support and confidence functions to improve fuzzy rule's quality. In the defuzzification layer, the fuzzy prediction from the intelligence is mapped back into the crisp model predictions for images. The loss between the prediction and label controls the rules to train the fuzzy detector. We show that the fuzzy rule-based network learns rich feature information than binary outputs and offer to obtain an overall performance gain. Experiment results show that compared to various benchmark fuzzy systems and adversarial attack detection methods, our fuzzy detector achieves better detection performance over a wide range of images. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IEEE Trans. Fuzzy Syst. | 3 |
| 2023 | Domain Generalization and Feature Fusion for Cross-domain Imperceptible Adversarial Attack DetectionabstractDeep learning-based imperceptible adversarial attack detection methods have recently seen significant progress. However, the accuracy, latency, and computational cost of previous methods remain insufficient. Particularly, trained attack detection models can potentially be applied in previously unseen conditions, such as new datasets or attacks for real-world applications. Therefore, to improve domain generalization performance, we propose a new method for cross-domain imperceptible adversarial attack detection by leveraging domain generalization, where we train the model's feature extractor or detector with a partner well-tuned for different domains. Different from conventional domain generalization methods, we use the global loss and local loss to train each feature extractor or detector. Moreover, to efficiently re-use high-resolution feature maps from the feature extractor, we propose a feature fusion network, which exploits feature maps from images that are attacked with different error rates and helps extract rich features to further improve the attack detection accuracy. Extensive experiments on four public datasets are used to demonstrate the efficacy of the proposed method. The source code of the proposed method is available at https://github.com/Yukino-3/DTAD. Yi Li 0047, Plamen Angelov 0001, Neeraj Suri |
IJCNN | 3 |
| 2023 | Replication: 20 Years of Inferring Interdomain Routing PoliciesabstractIn 2003, Wang and Gao [67] presented an algorithm to infer and characterize routing policies as this knowledge could be valuable in predicting and debugging routing paths. They used their algorithm to measure the phenomenon of selectively announced prefixes, in which, ASes would announce their prefixes to specific providers to manipulate incoming traffic. Since 2003, the Internet has evolved from a hierarchical graph, to a flat and dense structure. Despite 20 years of extensive research since that seminal work, the impact of these topological changes on routing policies is still blurred. Savvas Kastanakis, Vasileios Giotsas, Ioana Livadariu, Neeraj Suri |
IMC | 4 |
| 2023 | Privacy-preserving Decentralized Federated Learning over Time-varying Communication GraphabstractEstablishing how a set of learners can provide privacy-preserving federated learning in a fully decentralized (peer-to-peer, no coordinator) manner is an open problem. We propose the first privacy-preserving consensus-based algorithm for the distributed learners to achieve decentralized global model aggregation in an environment of high mobility, where participating learners and the communication graph between them may vary during the learning process. In particular, whenever the communication graph changes, the Metropolis-Hastings method [ 69 ] is applied to update the weighted adjacency matrix based on the current communication topology. In addition, the Shamir’s secret sharing (SSS) scheme [ 61 ] is integrated to facilitate privacy in reaching consensus of the global model. The article establishes the correctness and privacy properties of the proposed algorithm. The computational efficiency is evaluated by a simulation built on a federated learning framework with a real-world dataset. Zhengxin Yu, Neeraj Suri |
ACM Trans. Priv. Secur. | 3 |
| 2022 | Poster: Effectiveness of Moving Target Defense Techniques to Disrupt Attacks in the CloudabstractMoving Target Defense (MTD) can eliminate the asymmetric advantage that attackers have in terms of time to explore a static system by changing a system's configuration dynamically to reduce the efficacy of reconnaissance and increase uncertainty and complexity for attackers. To this extent, a variety of MTDs have been proposed for specific aspects of a system. However, deploying MTDs at different layers/components of the Cloud and assessing their effects on the overall security gains for the entire system is still challenging since the Cloud is a complex system entailing physical and virtual resources, and there exists a multitude of attack surfaces that an attacker can target. Thus, we explore the combination of MTDs, and their deployment at different components (belonging to various operational layers) to maximize the security gains offered by the MTDs.We also propose a quantification mechanism to evaluate the effectiveness of the MTDs against the attacks in the Cloud. Salman Manzoor, Antonios Gouglidis, Matthew Bradbury, Neeraj Suri |
CCS | 4 |
| 2022 | Poster: Multi-Layer Threat Analysis of the CloudabstractA variety of Threat Analysis (TA) techniques exist that typically target exploring threats to discrete assets (e.g., services, data, etc.) and reveal potential attacks pertinent to these assets. Furthermore, these techniques assume that the interconnection among the assets is static. However, in the Cloud, resources can instantiate or migrate across physical hosts at run-time, thus making the Cloud a dynamic environment. Additionally, the number of attacks targeting multiple assets/layers emphasizes the need for threat analysis approaches developed for Cloud environments. Therefore, this proposal presents a novel threat analysis approach that specifically addresses multi-layer attacks. The proposed approach facilitates threat analysis by developing a technology-agnostic information flow model. It contributes to exploring a threat's propagation across the operational stack of the Cloud and, consequently, holistically assessing the security of the Cloud. Salman Manzoor, Antonios Gouglidis, Matthew Bradbury, Neeraj Suri |
CCS | 4 |
| 2022 | Understanding the confounding factors of inter-domain routing modelingabstractThe Border Gateway Protocol (BGP) is a policy-based protocol, which enables Autonomous Systems (ASes) to independently define their routing policies with little or no global coordination. AS-level topology and AS-level paths inference have been long-standing problems for the past two decades, yet, an important question remains open: "which elements of Internet routing affect the AS-path inference accuracy and how much do they contribute to the error?". In this work, we: (1) identify the confounding factors behind Internet routing modeling, and (2) quantify their contribution on the inference error. Our results indicate that by solving the first-hop inference problem, we can increase the exact-path score from 33.6% to 84.1%, and, by taking geolocation into consideration, we can refine the accuracy up to 94.6%. Savvas Kastanakis, Vasileios Giotsas, Neeraj Suri |
IMC | 3 |
| 2022 | SlowCoach: Mutating Code to Simulate Performance BugsabstractPerformance bugs are unnecessarily inefficient code chunks in software codebases that cause prolonged execution times and degraded computational resource utilization. For performance bug diagnostics, tools that aid in the identification of said bugs, such as benchmarks and profilers, are commonly employed. However, due to factors such as insufficient workloads or ineffective benchmarks, software defects related to code inefficiencies are inherently difficult to diagnose. Hence, the capabilities of performance bug diagnostic tools are limited and performance bug instances may be missed. Traditional mutation testing (MT) is a technique for quantifying a test suite's ability to find functional bugs by mutating the code of the test subject. Similarly, we adopt performance mutation testing (PMT) to evaluate performance bug diagnostic tools and identify where improvements need to be made to a performance testing methodology. We carefully investigate the different performance bug fault models and how synthesized performance bugs based on these models can evaluate benchmarks and workload selection to help improve performance diagnostics. In this paper, we present the design of our PMT framework, SLOWCOACH, and evaluate it with over 1600 mutants from 4 real-world software projects. Oliver Schwahn, Roberto Natella, Matthew Bradbury, Neeraj Suri |
ISSRE | 5 |
| 2022 | PPFM: An Adaptive and Hierarchical Peer-to-Peer Federated Meta-Learning FrameworkabstractWith the advancement in Machine Learning (ML) techniques, a wide range of applications that leverage ML have emerged across research, industry, and society to improve application performance. However, existing ML schemes used within such applications struggle to attain high model accuracy due to the heterogeneous and distributed nature of their generated data, resulting in reduced model performance. In this paper we address this challenge by proposing PPFM: an adaptive and hierarchical Peer-to-Peer Federated Meta-learning framework. Instead of leveraging a conventional static ML scheme, PPFM uses multiple learning loops to dynamically self-adapt its own architecture to improve its training effectiveness for different generated data characteristics. Such an approach also allows for PPFM to remove reliance on a fixed centralized server in a distributed environment by utilizing peer-to-peer Federated Learning (FL) framework. Our results demonstrate PPFM provides significant improvement to model accuracy across multiple datasets when compared to contemporary ML approaches. Zhengxin Yu, Plamen Angelov 0001, Neeraj Suri |
MSN | 4 |
| 2021 | Fast Kernel Error Propagation Analysis in Virtualized EnvironmentsabstractAssessing operating system dependability remains a challenging problem, particularly in monolithic systems. Component interfaces are not well-defined and boundaries are not enforced at runtime. This allows faults in individual components to arbitrarily affect other parts of the system. Software fault injection (SFI) can be used to experimentally assess the resilience of such systems in the presence of faulty components. However, applying SFI to complex, monolithic operating systems poses challenges due to long test latencies and the difficulty of detecting corruptions in the internal state of the operating system.In this paper, we present a novel approach that leverages static and dynamic analysis alongside modern operating system and virtual machine features to reduce SFI test latencies for operating system kernel components while enabling efficient and accurate detection of internal state corruptions.We demonstrate the feasibility of our approach by applying it to multiple widely used Linux file systems. Nicolas Coppik, Oliver Schwahn, Neeraj Suri |
ICST | 3 |
| 2021 | Challenges in Identifying Network Attacks Using Netflow DataabstractLarge networks often encounter attacks that can affect the network availability. While multiple techniques exist to detect network attacks, a comprehensive understanding of how an attack occurs considering the various layers and components of the network software stack, can be an important element to help improve network security. By performing correlation analysis on contemporary unlabeled Netflow data, this paper conducts a comprehensive study of network flow events to identify communication patterns that may precede an attack, thereby providing potentially useful attack signatures to network administrators. Our work shows that, surprisingly, the Netflow data is not strongly correlated to network attacks. We observe that while spoof requests trigger reflection attacks, only a small percentage of the network packets are associated with the attack. Furthermore, lead time enhancements are feasible for reflection attacks that show long dwell times. Our study on network event correlations highlights empirical observations that could facilitate better attack handling in large networks. Edward Chuah, Neeraj Suri, Arshad Jhumka, Samantha Alt |
NCA | 2 |
| 2021 | PCaaD: Towards automated determination and exploitation of industrial systemsabstractOver the last decade, Programmable Logic Controllers (PLCs) have been increasingly targeted by attackers to obtain control over industrial processes that support critical services.Such targeted attacks typically require detailed knowledge of system-specific attributes, including hardware configurations, adopted protocols, and PLC control-logic, i.e., process comprehension.The consensus from both academics and practitioners suggests stealthy process comprehension obtained from a PLC alone, to execute targeted attacks, is impractical.In contrast, we assert that current PLC programming practices open the door to a new vulnerability class, affording attackers an increased level of process comprehension.To support this, we propose the concept of Process Comprehension at a Distance (PCaaD), as a novel methodological and automatable approach towards the system-agnostic identification of PLC library functions.This leads to the targeted exfiltration of operational data, manipulation of control-logic behavior, and establishment of covert command and control channels through unused memory.We validate PCaaD on widely used PLCs through its practical application. Benjamin Green 0001, Richard Derbyshire, Marina Krotofil, William Knowles, Daniel Prince, Neeraj Suri |
Comput. Secur. | 6 |
| 2020 | TraceSanitizer - Eliminating the Effects of Non-Determinism on Error Propagation AnalysisabstractModern computing systems typically relax execution determinism, for instance by allowing the CPU scheduler to inter- leave the execution of several threads. While beneficial for performance, execution non-determinism affects programs' execution traces and hampers the comparability of repeated executions. We present TraceSanitizer, a novel approach for execution trace comparison in Error Propagation Analyses (EPA) of multi-threaded programs. TraceSanitizer can identify and compensate for non- determinisms caused either by dynamic memory allocation or by non-deterministic scheduling. We formulate a condition under which TraceSanitizer is guaranteed to achieve a 0% false positive rate, and automate its verification using Satisfiability Modulo Theory (SMT) solving techniques. TraceSanitizer is comprehensively evaluated using execution traces from the PARSEC and Phoenix benchmarks. In contrast with other approaches, Trace- Sanitizer eliminates false positives without increasing the false negative rate (for a specific class of programs), with reasonable performance overheads. Habib Saissi, Stefan Winter 0001, Oliver Schwahn, Karthik Pattabiraman, Neeraj Suri |
DSN | 5 |
| 2020 | Decentralized Runtime Monitoring Approach Relying on the Ethereum Blockchain InfrastructureabstractCloud computing offers a model where resources (storage, applications, etc.) are abstracted and provided “as-aservice” in a remotely accessible manner. Although there are numerous claimed benefits of the Cloud to ensure confidentiality, integrity, and availability of the stored data, the number of security breaches is still on the rise. The lack of security assurance and transparency prevented customers/enterprises from trusting the Cloud Service Providers (CSPs). Unless the customer's security requirements are identified and documented by the CSPs, customers can not be assured that the CSPs will satisfy their requirements. Furthermore, the customer's compensation upon a violation is a manual time intensive process.In this paper we address the aforementioned challenges by proposing a decentralized customer-based monitoring approach running over Ethereum blockchain. The proposed approach allows the customer(s) to validate the compliance of CSP(s) to the contracted services in the Service Level Agreements (SLAs) and “autonomsly” compensate customers in case of security breaches. At the same time, the proposed approach prevents customers from misreporting for financial gain. The approach builds upon the Ethereum blockchain infrastructure in order to securely store monitoring logs and incorporate SLAs as smart contracts. The compliance validation framework is implemented and its functionality is evaluated on Amazon EC2 and Ethereum Blockchain. Ahmed Taha 0002, Ahmed Zakaria, Dong Seong Kim 0001, Neeraj Suri |
IC2E | 4 |
| 2020 | Extracting safe thread schedules from incomplete model checking resultsabstractAbstract Model checkers frequently fail to completely verify a concurrent program, even if partial-order reduction is applied. The verification engineer is left in doubt whether the program is safe and the effort toward verifying the program is wasted. We present a technique that uses the results of such incomplete verification attempts to construct a (fair) scheduler that allows the safe execution of the partially verified concurrent program. This scheduler restricts the execution to schedules that have been proven safe (and prevents executions that were found to be erroneous). We evaluate the performance of our technique and show how it can be improved using partial-order reduction. While constraining the scheduler results in a considerable performance penalty in general, we show that in some cases our approach—somewhat surprisingly—even leads to faster executions. Patrick Metzler, Neeraj Suri, Georg Weissenbacher |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2020 | Analyzing the Effects of Bugs on Software InterfacesabstractCritical systems that integrate software components (e.g., from third-parties) need to address the risk of residual software defects in these components. Software fault injection is an experimental solution to gauge such risk. Many error models have been proposed for emulating faulty components, such as by injecting error codes and exceptions, or by corrupting data with bit-flips, boundary values, and random values. Even if these error models have been able to find breaches in fragile systems, it is unclear whether these errors are in fact representative of software faults. To pursue this open question, we propose a methodology to analyze how software faults in C/C++ software components turn into errors at components' interfaces (interface error propagation), and present an experimental analysis on what, where, and when to inject interface errors. The results point out that the traditional error models, as used so far, do not accurately emulate software faults, but that richer interface errors need to be injected, by: injecting both fail-stop behaviors and data corruptions; targeting larger amounts of corrupted data structures; emulating silent data corruptions not signaled by the component; combining bit-flips, boundary values, and data perturbations. Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
IEEE Trans. Software Eng. | 4 |
| 2019 | MemFuzz: Using Memory Accesses to Guide FuzzingabstractFuzzing is a form of random testing that is widely used for finding bugs and vulnerabilities. State of the art approaches commonly leverage information about the control flow of prior executions of the program under test to decide which inputs to mutate further. By relying solely on control flow information to characterize executions, such approaches may miss relevant differences. We propose augmenting evolutionary fuzzing by additionally leveraging information about memory accesses performed by the target program. The resulting approach can leverage more sophisticated information about the execution of the target program, enhancing the effectiveness of the evolutionary fuzzing. We implement our approach as a modification of the widely used AFL fuzzer and evaluate our implementation on three widely used target applications. We find distinct crashes from those detected by AFL for all three targets in our evaluation. Nicolas Coppik, Oliver Schwahn, Neeraj Suri |
ICST | 3 |
| 2019 | Inferring Performance Bug Patterns from Developer CommitsabstractPerformance bugs, i.e., program source code that is unnecessarily inefficient, have received significant attention by the research community in recent years. A number of empirical studies have investigated how these bugs differ from "ordinary" bugs that cause functional deviations and several approaches to aid their detection, localization, and removal have been proposed. Many of these approaches focus on certain subclasses of performance bugs, e.g., those resulting from redundant computations or unnecessary synchronization, and the evaluation of their effectiveness is usually limited to a small number of known instances of these bugs. To provide researchers working on performance bug detection and localization techniques with a larger corpus of performance bugs to evaluate against, we conduct a study of more than 700 performance bug fixing commits across 13 popular open source projects written in C and C++ and investigate the relative frequency of bug types as well as their complexity. Our results show that many of these fixes follow a small set of bug patterns, that they are contributed by experienced developers, and that the number of lines needed to fix performance bugs is highly project dependent. Stefan Winter 0001, Neeraj Suri |
ISSRE | 3 |
| 2019 | Assessing the state and improving the art of parallel testing for CabstractThe execution latency of a test suite strongly depends on the degree of concurrency with which test cases are executed. However, if test cases are not designed for concurrent execution, they may interfere, causing result deviations compared to sequential execution. To prevent this, each test case can be provided with an isolated execution environment, but the resulting overheads diminish the merit of parallel testing. Our large-scale analysis of the Debian Buster package repository shows that existing test suites in C projects make limited use of parallelization. We present an approach to (a) analyze the potential of C test suites for safe concurrent execution, i.e., result invariance compared to sequential execution, and (b) execute tests concurrently with different parallelization strategies using processes or threads if it is found to be safe. Applying our approach to 9 C projects, we find that most of them cannot safely execute tests in parallel due to unsafe test code or unsafe usage of shared variables or files within the program code. Parallel test execution shows a significant acceleration over sequential execution for most projects. We find that multi-threading rarely outperforms multi-processing. Finally, we observe that the lack of a common test framework for C leaves make as the standard driver for running tests, which introduces unnecessary performance overheads for test execution. Oliver Schwahn, Nicolas Coppik, Stefan Winter 0001, Neeraj Suri |
ISSTA | 4 |
| 2019 | Extracting Safe Thread Schedules from Incomplete Model Checking Results
Patrick Metzler, Neeraj Suri, Georg Weissenbacher |
SPIN | 2 |
| 2019 | Gyro: A Modular Scale-Out Layer for Single-Server DBMSsabstractScaling out database management systems (DBMSs) requires distributed coordination, which can easily become a bottleneck. Recent work on speeding up distributed transactions has addressed this problem by proposing scale-out techniques that are deeply integrated with the concurrency control mechanism of the DBMS. This paper explores the design of modular coordination layers, which encapsulate all scale-out logic and can be applied to scale out any unmodified single-server DBMS. It proposes Gyro, a modular coordination layer that runs on top of a collection of single-server DBMS instances and interacts with them only through their client interface. Gyro distributes the load by ensuring that as many requests as possible are executed by only one DBMS instance. Our experiments show that modular distributed coordination is practically viable and that it can be much faster than traditional distributed transaction protocols using two-phase commit. Habib Saissi, Marco Serafini, Neeraj Suri |
SRDS | 3 |
| 2019 | Security Requirements Engineering in Safety-Critical Railway Signalling NetworksabstractSecuring a safety-critical system is a challenging task, because safety requirements have to be considered alongside security controls. We report on our experience to develop a security architecture for railway signalling systems starting from the bare safety-critical system that requires protection. We use a threat-based approach to determine security risk acceptance criteria and derive security requirements. We discuss the executed process and make suggestions for improvements. Based on the security requirements, we develop a security architecture. The architecture is based on a hardware platform that provides the resources required for safety as well as security applications and is able to run these applications of mixed-criticality (safety-critical applications and other applications run on the same device). To achieve this, we apply the MILS approach, a separation-based high-assurance security architecture to simplify the safety case and security case of our approach. We describe the assurance requirements of the separation kernel subcomponent, which represents the key component of the MILS architecture. We further discuss the security measures of our architecture that are included to protect the safety-critical application from cyberattacks. Markus Heinrich, Tsvetoslava Vateva-Gurova, Tolga Arul, Stefan Katzenbeisser 0001, Neeraj Suri, Henk Birkholz, Andreas Fuchs 0002, Christoph Krauß, Maria Zhdanova, Don Kuzhiyelil, Sergey Tverdyshev, Christian Schlehuber |
Secur. Commun. Networks | 5 |
| 2019 | Cross-Domain Noise Impact Evaluation for Black Box Two-Level Control CPSabstractControl Cyber-Physical Systems (CPSs) constitute a major category of CPS. In control CPSs, in addition to the well-studied noises within the physical subsystem, we are interested in evaluating the impact of cross-domain noise : the noise that comes from the physical subsystem, propagates through the cyber subsystem, and goes back to the physical subsystem. Impact of cross-domain noise is hard to evaluate when the cyber subsystem is a black box, which cannot be explicitly modeled. To address this challenge, this article focuses on the two-level control CPS, a widely adopted control CPS architecture, and proposes an emulation based evaluation methodology framework. The framework uses hybrid model reachability to quantify the cross-domain noise impact, and exploits Lyapunov stability theories to reduce the evaluation benchmark size. We validated the effectiveness and efficiency of our proposed framework on a representative control CPS testbed. Particularly, 24.1% of evaluation effort is saved using the proposed benchmark shrinking technology. Liansheng Liu, Stefan Winter 0001, Qixin Wang 0001, Neeraj Suri, Lei Bu, Yu Peng 0002, Xue (Steve) Liu, Xiyuan Peng |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2019 | Proofs of Writing for Robust StorageabstractExisting Byzantine fault tolerant (BFT) storage solutions that achieve strong consistency and high availability, are costly compared to solutions that tolerate simple crashes. This cost is one of the main obstacles in deploying BFT storage in practice. In this paper, we present PoWerStore, a robust and efficient data storage protocol. PoWerStore's robustness comprises tolerating network outages, maximum number of Byzantine storage servers, any number of Byzantine readers and crash-faulty writers, and guaranteeing high availability (wait-freedom) and strong consistency (linearizability) of read/write operations. PoWerStore's efficiency stems from combining lightweight cryptography, erasure coding and metadata write-backs, where readers write-back only metadata to achieve strong consistency. Central to PoWerStore is the concept of “Proofs of Writing” (PoW), a novel data storage technique inspired by commitment schemes. PoW rely on a 2-round write procedure, in which the first round writes the actual data and the second round only serves to “prove” the occurrence of the first round. PoW enable efficient implementations of strongly consistent BFT storage through metadata write-backs and low latency reads. We implemented PoWerStore and show its improved performance when compared to state of the art robust storage protocols, including protocols that tolerate only crash faults. Dan Dobre, Ghassan Karame, Wenting Li 0001, Matthias Majuntke, Neeraj Suri, Marko Vukolic |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Flashlight: A Novel Monitoring Path Identification Schema for Securing Cloud ServicesabstractCloud monitoring is an essential mechanism for helping secure cloud services. Thus, a plethora of monitoring schemas have been proposed in recent years. Particularly, a newly proposed indirect monitoring mechanism outperforms others with the unique merit of addressing scenarios where the information of the monitoring target is not directly accessible. To conduct indirect cloud security monitoring, a key prerequisite is to obtain a special set of monitoring data termed "monitoring path". However, how to ascertain the monitoring path is still an open issue. Heng Zhang 0009, Jesus Luna, Neeraj Suri, Rubén Trapero |
ARES | 3 |
| 2018 | Threat Modeling and Analysis for the Cloud EcosystemabstractAs the usage of the Cloud proliferates, the need for security evaluation of the Cloud also grows. The process of threat modeling and analysis is advocated to assess potential vulnerabilities that can undermine the Cloud security goals. However, given the plethora of distinct services involved in the Cloud ecosystem and the varied attack surfaces entailed in the Cloud-specific architectures, performing threat analysis for the Cloud is a challenging task. Consequently, contemporary Cloud threat analysis approaches, typically using relational security models (e.g., attack graphs, trees...), primarily focus on specific services/layers of the Cloud. Also, these schemes often fail to include the variants of the identified vulnerabilities in their analysis. Hence, a comprehensive threat analysis approach is required that can (a) model and analyze threats across the multilayer Cloud operational stack, and (b) include variants of the vulnerabilities in the threat analysis procedure. We target achieving a holistic Cloud threat analysis by designing a novel multi-layer Cloud model, using Petri Nets, to comprehensively profile the operational behavior of the services involved in the Cloud operations. We subsequently conduct threat modeling to identify threats within and across the different layers of the Cloud operations. Our proposed threat analysis approach also investigates the variants of the potential vulnerabilities to comprehensively infer the Cloud attack surface. Salman Manzoor, Heng Zhang 0009, Neeraj Suri |
IC2E | 3 |
| 2018 | Monitoring Path Discovery for Supporting Indirect Monitoring of Cloud ServicesabstractCloud monitoring is an established support mechanism for securing Cloud services. Over the last few years, various Cloud monitoring mechanisms have been proposed with different level of monitoring efficacy. To achieve effective monitoring, the monitoring mechanism is required to address the basic challenge that the information of the target-to-monitor may be inaccessible due to (i) access controls, (ii) privacy protection, or (iii) technical difficulties. Therefore, the indirect monitoring mechanism is particularly proposed for addressing the challenge. Specifically, the indirect mechanism makes inferences about the inaccessible information of monitoring targets with the help of the "monitoring path" that contains a special set of accessible monitoring data underpinning the inference task. However, the process to effectively discover the monitoring path is an open issue. To address this problem, we present our preliminary studies in this paper. Firstly, we review existing work and reveal the limitations for discovering monitoring paths. Secondly, we analyze real cases to provide insights for designing monitoring paths. Finally, we propose a framework and planned research tasks for developing a novel monitoring path discovery mechanism that facilitates performing indirect Cloud monitoring as the expected contribution of our research. Heng Zhang 0009, Salman Manzoor, Neeraj Suri |
IC2E | 3 |
| 2018 | FastFI: Accelerating Software Fault InjectionsabstractSoftware Fault Injection (SFI) is a widely used technique to experimentally assess the dependability of software systems. To provide a comprehensive view on the dependability of a software under test, SFI typically requires large numbers of experiments, which leads to long test latencies. In order to reduce the overall test duration for SFI, we propose FASTFI, which (1) avoids redundant executions of common path prefixes for faults in the same injection location, (2) avoids test executions for faults that do not get activated, and (3) utilizes parallel processors by executing SFI tests concurrently. FASTFI takes patch files that specify source code mutations as an input, conducts an automated source code analysis to identify the function they target, and then automatically parallelizes the execution of all mutants that target the same function. Our evaluation of FASTFI on four PARSEC benchmarks shows a SFI test latency reduction of up to a factor of 26. Oliver Schwahn, Nicolas Coppik, Stefan Winter 0001, Neeraj Suri |
PRDC | 4 |
| 2018 | Exploring the Relationship Between Dimensionality Reduction and Private Data ReleaseabstractIt is important to facilitate data sharing between data owners and data analysts as data owners do not always have the ability to process and analyze data. For example, governments around the world are starting to release collected data to the public to leverage data analysis competence of the crowd. However, some privacy leakage incidents have made the public to rediscover the importance of privacy protection, leading to new privacy regulations. In existing researches dimensionality reduction has played an important role in private data release mechanisms to improve utility but its influence on privacy protection has never been examined. In this study, we perform a series of experiments and found that dimensionality reduction could provide similar privacy protection effects as K-anonymity mechanisms, and it could work as a preprocessor of K-anonymity process to it to reduce the generalization and suppression needed. Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, Neeraj Suri, Pang-Chieh Wang |
PRDC | 4 |
| 2018 | InfoLeak: Scheduling-Based Information LeakageabstractCovert-and side-channel attacks, typically enabled by the usage of shared resources, pose a serious threat to complex systems such as the Cloud. While their exploitation in the real world depends on properties of the execution environment (e.g., scheduling), the explicit consideration of these factors is often neglected. This paper introduces InfoLeak, an information leakage model that establishes the crucial role of the scheduler for exploiting core-private caches as covert channels. We show, formally and empirically, how the availability of these channels and the corresponding attack feasibility are affected by scheduling. Moreover, our model allows security experts to assess the related threat, posed by core-private cache covert channels for a particular system by considering solely the scheduling information. To validate the utility of InfoLeak, we deploy a covert-channel attack and correlate its success ratio to the scheduling of the attacker processes in the target system. We demonstrate the applicability of the InfoLeak model for analyzing the scheduling information for possible information leakage and also provide an example on its usage. Tsvetoslava Vateva-Gurova, Salman Manzoor, Yennun Huang, Neeraj Suri |
PRDC | 4 |
| 2018 | On the Detection of Side-Channel AttacksabstractThreats posed by side-channel and covert-channel attacks exploiting the CPU cache to compromise the confidentiality of a system raise serious security concerns. This applies especially to systems offering shared hardware or resources to their customers. As eradicating this threat is practically impeded due to performance implications or financial cost of the current mitigation approaches, a detection mechanism might enhance the security of such systems. In the course of this work, we propose an approach towards side-channel attacks detection, considering the specificity of cache-based SCAs and their implementations. Tsvetoslava Vateva-Gurova, Neeraj Suri |
PRDC | 2 |
| 2018 | Whetstone: Reliable Monitoring of Cloud ServicesabstractCloud services have become powerful enablers for a variety of smart computing solutions supporting multimedia, social networking, e-commerce and critical infrastructures among others. Consequently, as we increasingly depend on the cloud, the need exists to ensure its effective role as a trustworthy services platform. Towards this objective, a plethora of cloud monitoring mechanisms have been proposed which typically assume that the collected monitoring information is reliably correct. In reality, the information collected by cloud monitors is often susceptible to reliability issues (e.g., monitor malfunctions, data corruptions, or data tampering), and obtaining reliable cloud monitoring information is still an open issue. We propose Whetstone as a novel approach to address the gap where an efficient approach of ascertaining reliable values from a set of collected monitoring data is required. To this end, Whetstone first introduces a statistical approach to filter defective data from the collected data set. Next, Whetstone develops an optimization approach to quantify the reliability of the collected data by leveraging the value deviation of the collected data. Finally, Whetstone devises a weighted aggregation approach for generating the reliable value based on the obtained information. We evaluate the proposed approach with different experimental configurations. The experimental results demonstrate the efficacy of our approach for successfully generating the maximum likelihood reliable value for raw data sets. Heng Zhang 0009, Jesus Luna, Rubén Trapero, Neeraj Suri |
SMARTCOMP | 4 |
| 2018 | A Detection Mechanism for Internal Attacks on Pull-Based P2P Streaming SystemsabstractOnline streaming is a popular service for data-intensive applications such as video streaming. P2P-based streaming solutions are advocated to help reduce costs for both providers and users. Yet, involving users over data dissemination entails security risks including a variety of denial-of-service attacks. While extensive research exists on mitigating varied attack types, their effectiveness is limited if the attacker can infer information about the topology such as the identity of nodes that have direct connections to the source. The attacker can then leverage the gained insights to place malicious participants in prominent positions. By dropping chunks that should be forwarded, the malicious peers degrade the performance in a stealthy way that does not raise suspicion. We first demonstrate the feasibility of conducting such attacks. Accordingly, we propose a detection mechanism that identifies the attack and removes potential malicious peers from their disruptive positions. We ascertain, theoretically and through simulations, that malicious peers cannot misuse the detection mechanism to gain influence. Our simulation-based study indicates that the proposed detection mechanism is able to detect malicious peers with up to 80-90% accuracy while inducing a small overhead of approximately 8%. Hatem Ismail, Stefanie Roos, Neeraj Suri |
WOWMOM | 3 |
| 2018 | How to Fillet a Penguin: Runtime Data Driven Partitioning of Linux CodeabstractIn many modern operating systems (OSs), there exists no isolation between different kernel components, i.e., the failure of one component can affect the whole kernel. While microkernel OSs introduce address space separation for large parts of the OS, their improved fault isolation comes at the cost of performance. Despite significant improvements in modern microkernels, monolithic OSs like Linux are still prevalent in many systems. To achieve fault isolation in addition to high performance and code reuse in these systems, approaches to move only fractions of kernel code into user mode have been proposed. These approaches solely rely on static code analyses for deciding which code to isolate, neglecting dynamic properties like invocation frequencies. We propose to augment static code analyses with runtime data to achieve better estimates of dynamic properties for common case operation. We assess the impact of runtime data on the decision what code to isolate and the impact of that decision on the performance of such “microkernelized” systems. We extend an existing tool chain to implement automated code partitioning for existing monolithic kernel code and validate our approach in a case study of two widely used Linux device drivers and a file system. Oliver Schwahn, Stefan Winter 0001, Nicolas Coppik, Neeraj Suri |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2017 | C'mon: Monitoring the Compliance of Cloud Services to Contracted PropertiesabstractThe usage of computing resources "as a service" makes cloud computing an attractive solution for enterprises with fluctuating needs for information processing. As security aspects play an important role when cloud computing is applied for business-critical tasks, security service level agreements (secSLAs) have been proposed to specify the security properties of a provided cloud service. Soha Alboghdady, Stefan Winter 0001, Ahmed Taha 0002, Heng Zhang 0009, Neeraj Suri |
ARES | 5 |
| 2017 | Towards DDoS Attack Resilient Wide Area Monitoring SystemsabstractThe traditional physical power grid is evolving into a cyber-physical Smart Grid (SG) that links the cyber communication and computational elements with the physical control functions to dynamically integrate varied and geographically distributed energy producers/consumers. In the SG, the cyber elements of Wide Area Measurement Systems (WAMS) are deployed to provide the critical monitoring of the state of power transmission and distribution to accomplish real-time control of the grid. Unfortunately, the increasing adoption of such computing/communication cyber-technologies essential to providing the SG operations also opens the risk of the SG being vulnerable to cyberattacks. In particular, attacks such as Denial-of-Service (DoS) and Distributed DoS (DDoS) are of primary concern for WAMS where such attacks can compromise its safety-critical accuracy and responsiveness characteristics. Kubilay Demir 0001, Neeraj Suri |
ARES | 2 |
| 2017 | IPA: Error Propagation Analysis of Multi-Threaded Programs Using Likely InvariantsabstractError Propagation Analysis (EPA) is a technique forunderstanding how errors affect a program's execution and resultin program failures. For this purpose, EPA usually compares thetraces of a fault-free (golden) run with those from a faulty run ofthe program. This makes existing EPA approaches brittle for multithreadedprograms, which do not typically have a deterministicgolden run. In this paper, we study the use of likely invariantsgenerated by automated approaches as alternatives for goldenrun based EPA in multithreaded programs. We present InvariantPropagation Analysis (IPA), an approach and a framework forautomatically deriving invariants for multithreaded programs, and using the invariants for EPA. We evaluate the invariantsderived by IPA in terms of their coverage for different faulttypes across six representative programs through fault injectionexperiments. We find that stable invariants can be inferred in allsix programs, although their coverage of faults depends on theapplication and the fault type. Abraham Chan, Stefan Winter 0001, Habib Saissi, Karthik Pattabiraman, Neeraj Suri |
ICST | 5 |
| 2017 | TrEKer: tracing error propagation in operating system kernelsabstractModern operating systems (OSs) consist of numerous interacting components, many of which are developed and maintained independently of one another. In monolithic systems, the boundaries of and interfaces between such components are not strictly enforced at runtime. Therefore, faults in individual components may directly affect other parts of the system in various ways. Software fault injection (SFI) is a testing technique to assess the resilience of a software system in the presence of faulty components. Unfortunately, SFI tests of OSs are inconclusive if they do not lead to observable failures, as corruptions of the internal software state may not be visible at its interfaces and, yet, affect the subsequent execution of the OS beyond the duration of the test. In this paper we present TrEKer, a fully automated approach for identifying how faulty OS components affect other parts of the system. TrEKer combines static and dynamic analyses to achieve efficient tracing on the granularity of memory accesses. We demonstrate TrEKer's ability to support SFI oracles by accurately tracing the effects of faults injected into three widely used Linux kernel modules. Nicolas Coppik, Oliver Schwahn, Stefan Winter 0001, Neeraj Suri |
ASE | 4 |
| 2017 | Quick verification of concurrent programs by iteratively relaxed schedulingabstractThe most prominent advantage of software verification over testing is a rigorous check of every possible software behavior. However, large state spaces of concurrent systems, due to non-deterministic scheduling, result in a slow automated verification process. Therefore, verification introduces a large delay between completion and deployment of concurrent software. This paper introduces a novel iterative approach to verification of concurrent programs that drastically reduces this delay. By restricting the execution of concurrent programs to a small set of admissible schedules, verification complexity and time is drastically reduced. Iteratively adding admissible schedules after their verification eventually restores non-deterministic scheduling. Thereby, our framework allows to find a sweet spot between a low verification delay and sufficient execution time performance. Our evaluation of a prototype implementation on well-known benchmark programs shows that after verifying only few schedules of the program, execution time overhead is competitive to existing deterministic multi-threading frameworks. Patrick Metzler, Habib Saissi, Péter Bokor, Neeraj Suri |
ASE | 4 |
| 2017 | SeReCP: A Secure and Reliable Communication Platform for the Smart GridabstractThe management of a complex cyber-physical system such as the Smart Grid (SG) requires responsive, scalable and high-bandwidth communication, which is often beyond the capabilities of the classical closed communication networks of the power grid. Consequently, the use of scalable public IP-based networks is increasingly being advocated. However, a direct consequence of the use of public networks is the exposure of the SG to varied reliability/security risks, e.g., distributed denial-of-service (DDoS). Thus the need exists for new lightweight mechanisms that can provide both cost-effective communication along with proactive DDoS attack protection. We fill this gap by proposing a novel approach termed as SeReCP, which leverages: (1) a semi-trusted P2P-based publish-subscribe (pub-sub) system providing a proactive countermeasure for DDoS attacks and secure group communications by aid of a group key management system, (2) a data diffusion mechanism that sustains the network availability in the case of both randomly sweeping and targeted DDoS attacks on pub-sub brokers, and (3) a multi-homing-based fast recovery mechanism for detecting and requesting the dropped packets, thus paving the way for meeting the stringent laency requirements os SG applications. Our evaluation on a real testbed demonstrates that SG applications. Our evaluation on a real testbed demonstrates that SeReCP provides the required security and availability of SG applications with up to 30% failures of the pnb-snb brokers. Overall, we show that SeReCP helps enable the secure use of public network based communication for safety-critical cyber-physical systems such as the SG. Kubilay Demir 0001, Neeraj Suri |
PRDC | 2 |
| 2017 | A Security Architecture for Railway Signalling
Christian Schlehuber, Markus Heinrich, Tsvetoslava Vateva-Gurova, Stefan Katzenbeisser 0001, Neeraj Suri |
SAFECOMP | 5 |
| 2017 | P2P routing table poisoning: A quorum-based sanitizing approach
Hatem Ismail, Daniel Germanus, Neeraj Suri |
Comput. Secur. | 3 |
| 2017 | A novel approach to manage cloud security SLA incidents
Rubén Trapero, Jolanda Modic, Miha Stopar, Ahmed Taha 0002, Neeraj Suri |
Future Gener. Comput. Syst. | 5 |
| 2017 | Robust QoS-aware communication in the smart distribution grid
Kubilay Demir 0001, Daniel Germanus, Neeraj Suri |
Peer-to-Peer Netw. Appl. | 3 |
| 2017 | Quantitative Reasoning about Cloud Security Using Service Level AgreementsabstractWhile the economic and technological advantages of cloud computing are apparent, its overall uptake has been limited, in part, due to the lack of security assurance and transparency on the Cloud Service Provider (CSP). Although, the recent efforts on specification of security using Service Level Agreements, also known as “Security Level Agreements” or secSLAs is a positive development multiple technical and usability issues limit the adoption of Cloud secSLA's in practice. In this paper we develop two evaluation techniques, namely QPT and QHP, for conducting the quantitative assessment and analysis of the secSLA based security level provided by CSPs with respect to a set of Cloud Customer security requirements. These proposed techniques help improve the security requirements specifications by introducing a flexible and simple methodology that allows Customers to identify and represent their specific security needs. Apart from detailing guidance on the standalone and collective use of QPT and QHP, these techniques are validated using two use case scenarios and a prototype, leveraging actual real-world CSP secSLAdata derived from the Cloud Security Alliance's Security, Trust and Assurance Registry. Jesus Luna, Ahmed Taha 0002, Rubén Trapero, Neeraj Suri |
IEEE Trans. Cloud Comput. | 4 |
| 2016 | Efficient Verification of Program Fragments: Eager POR
Patrick Metzler, Habib Saissi, Péter Bokor, Robin Hesse, Neeraj Suri |
ATVA | 5 |
| 2016 | Identifying and Utilizing Dependencies Across Cloud Security ServicesabstractSecurity concerns are often mentioned amongst the reasons why organizations hesitate to adopt Cloud computing. Given that multiple Cloud Service Providers (CSPs) offer similar security services (e.g., "encryption key management") albeit with different capabilities and prices, the customers need to comparatively assess the offered security services in order to select the best CSP matching their security requirements. However, the presence of both explicit and implicit dependencies across security related services add further challenges for Cloud customers to (i) specify their security requirements taking service dependencies into consideration and (ii) to determine which CSP can satisfy these requirements. We present a framework to address these challenges. For challenge (i), our framework automatically detects conflicts resulting from inconsistent customer requirements. Moreover, our framework provides an explanation for the detected conflicts allowing customers to resolve these conflicts. To tackle challenge (ii), our framework assesses the security level provided by various CSPs and ranks the CSPs according to the desired customer requirements. We demonstrate the framework's effectiveness with real-world CSP case studies derived from the Cloud Security Alliance's Security, Trust and Assurance Registry. Ahmed Taha 0002, Patrick Metzler, Rubén Trapero, Jesus Luna, Neeraj Suri |
AsiaCCS | 5 |
| 2016 | Novel efficient techniques for real-time cloud security assessment
Jolanda Modic, Rubén Trapero, Ahmed Taha 0002, Jesus Luna, Miha Stopar, Neeraj Suri |
Comput. Secur. | 6 |
| 2016 | Run Time Application Repartitioning in Dynamic Mobile Cloud EnvironmentsabstractAs mobile computing increasingly interacts with the cloud, a number of approaches, e.g., MAUI and CloneCloud, have been proposed, aiming to offload parts of the mobile application execution to the cloud. To achieve a good performance by using these approaches, they particularly focus on the application partitioning problem, i.e., to decide which parts of an application should be offloaded to the cloud and which parts should be executed on mobile devices such that the execution cost is minimized. Most works on this problem assume that the offloading cost of each part of the application remains the same as the application is running. Unfortunately, this assumption does not hold in dynamic mobile cloud environments, where the device and network connection status may fluctuate, and thus affects the offloading cost. With the varying offloading cost, the one time partitioning of the application may yield significant performance degradations. In this paper, we study application repartitioning problem which considers updating the partition periodically during the course of application execution. We first propose a framework for run time application repartitioning in dynamic mobile cloud environments. Based on this framework, we take the dynamic network connection to clouds as a case study, and design an online solution, Foreseer, to solve the mobile cloud application repartitioning problem. We evaluate our solution based on real world data traces that are collected in a campus WiFi hotspot testbed. The result shows that our method can achieve significantly shorter completion time over previous approaches. Lei Yang 0024, Jiannong Cao 0001, Shaojie Tang 0001, Di Han 0002, Neeraj Suri |
IEEE Trans. Cloud Comput. | 5 |
| 2015 | PBMC: Symbolic Slicing for the Verification of Concurrent Programs
Habib Saissi, Péter Bokor, Neeraj Suri |
ATVA | 3 |
| 2015 | Detecting and Mitigating P2P Eclipse AttacksabstractPeer-to-Peer (P2P) protocols increasingly constitute the foundations for many large-scale applications as the inherently distributed nature of P2P easily supports both scalability and fault-tolerance. However, the decentralized design of P2P also exposes it to a variety of distributed threats with Eclipse Attacks (EAs) being a prominent type to impact P2P functionality. While the basic technique of divergent lookups has been demonstrated for suitability to mitigate EA, it can only (effectively) address limited variants of EAs. This paper investigates both the detection and mitigation potential of enhanced divergent lookups for handling complex EA scenarios. In addition, we propose an approach that can identify malicious peers with a high degree of accuracy. Our simulations have shown EA mitigation rates of up to 96% in case 25% of the peers are malicious. Also, our approach allows for anonymity-fostering, fully decentralized usage, and facilitating downstream mechanisms such as malicious peer removal. Hatem Ismail, Daniel Germanus, Neeraj Suri |
ICPADS | 3 |
| 2015 | No PAIN, No Gain? The Utility of PArallel Fault INjectionsabstractSoftware Fault Injection (SFI) is an established technique for assessing the robustness of a software under test by exposing it to faults in its operational environment. Depending on the complexity of this operational environment, the complexity of the software under test, and the number and type of faults, a thorough SFI assessment can entail (a) numerous experiments and (b) long experiment run times, which both contribute to a considerable execution time for the tests. In order to counteract this increase when dealing with complex systems, recent works propose to exploit parallel hardware to execute multiple experiments at the same time. While Parallel fault Injections (PAIN) yield higher experiment throughput, they are based on an implicit assumption of non-interference among the simultaneously executing experiments. In this paper we investigate the validity of this assumption and determine the trade-off between increased throughput and the accuracy of experimental results obtained from PAIN experiments. Stefan Winter 0001, Oliver Schwahn, Roberto Natella, Neeraj Suri, Domenico Cotroneo |
ICSE (1) | 4 |
| 2015 | Mitigating Timing Error Propagation in Mixed-Criticality Automotive SystemsabstractFor mixed-criticality automotive systems, the functional safety standard ISO 26262 stipulates freedom from interference, i.e., Errors should not propagate from low to high criticality tasks. To prevent the propagation of timing errors, the automotive software standard AUTOSAR provides monitor-based timing protection, which detects and confines task timing errors. As current monitors are unaware of a criticality concept, the effective protection of a critical task requires to monitor all tasks that constitute a potential source of propagating errors, thereby causing overhead for worst-case execution time analysis, configuration and monitoring. Differing from the indirect protection of critical tasks facilitated by existing mechanisms, we propose a novel monitoring scheme that directly protects critical tasks from interference, by providing them with execution time guarantees. Overall, our approach provides efficient low-overhead interference protection, while also adding transient timing error ride-through capabilities. Thorsten Piper, Stefan Winter 0001, Oliver Schwahn, Suman Bidarahalli, Neeraj Suri |
ISORC | 5 |
| 2015 | FTDE: Distributed Fault Tolerance for WSN Data Collection and Compression SchemesabstractWireless Sensor Networks (WSNs), being power and communication capacity constrained, often employ data compression schemes to reduce the data volumes. At the same time, various factors, such as low node reliability and communication faults, can compromise the core objective of accurate collection and delivery of the sensor data. However, contemporary compression schemes often fail to consider operational faults, and the resulting data errors arising from low node reliability (node crashes), and communication faults (data and links corruption) result in erroneous data being collected. This paper proposes a novel model-based fault-tolerance technique applicable to varied compression schemes. Being fully distributed, our scheme corrects data errors at the point of origin (faulty sensor nodes), and thus avoids costly transmissions of corrupt data, decreasing message cost, and enhances compression effectiveness by correcting the erroneous samples. Azad Ali, Abdelmajid Khelil, Neeraj Suri |
SRDS | 3 |
| 2015 | PASS: An Address Space Slicing Framework for P2P Eclipse Attack MitigationabstractThe decentralized design of Peer-to-Peer (P2P) protocols inherently provides for fault tolerance to non-malicious faults. However, the base P2P scalability and decentralization requirements often result in design choices that negatively impact their robustness to varied security threats. A prominent vulnerability are Eclipse attacks that aim at information hiding and consequently perturb a P2P overlay's reliable service delivery. Divergent lookups constitute an advocated mitigation technique but are size-limited to overlay networks with tens of thousands of peers. In this work, building upon divergent lookups, we propose a novel and scalable P2P address space slicing strategy (PASS) to efficiently mitigate attacks in overlays that host hundreds of thousands of peers. Moreover, we integrate and evaluate diversely designed lookup variants to assess their network overhead and mitigation rates. The proposed PASS approach shows mitigation rates reaching up to 100%. Daniel Germanus, Hatem Ismail, Neeraj Suri |
SRDS | 3 |
| 2015 | In the Compression Hornet's Nest: A Security Study of Data Compression in Network Services
Giancarlo Pellegrino, Davide Balzarotti, Stefan Winter 0001, Neeraj Suri |
USENIX Security Symposium | 4 |
| 2015 | Adaptive Hybrid Compression for Wireless Sensor NetworksabstractWireless Sensor Networks (WSNs) are often deployed to sample the desired environmental attributes and deliver the acquired samples to a central station, termed as the sink, for processing as needed by the application. Many applications stipulate high granularity and data accuracy that results in high data volumes. However, sensor nodes are battery powered, and sending the requested large amounts of data rapidly depletes their energy. Fortunately, environmental attributes (e.g., temperature, pressure) often exhibit spatial and temporal correlations. Moreover, a large class of applications such as scientific analysis and simulations tolerate high latency for sensor data collection. Hence, we exploit the spatiotemporal correlation of sensor readings while benefiting from possible data delivery latency tolerance to minimize the amount of data to be transported to the sink. Accordingly, we develop a fully distributed adaptive hybrid compression scheme that exploits both spatial and temporal data redundancies and fuses both temporal and spatial compression for maximal data compression with accuracy guarantees. We present two main contributions: (i) an adaptive modeling technique that allows frugal and maximized temporal compression on resource-constraint sensor nodes by exploiting the data collection latency, and (ii) a novel model-based hierarchical clustering technique that allows for maximized spatial compression resulting into a hybrid compression scheme. Compared to the existing spatiotemporal compression approaches, our approach is fully decentralized and the proposed clustering scheme is based on sensor data models rather than instantaneous sensor data values, which allows merging nearby nodes with similar models into large clusters over a longer period of time rather than specific time instances. The analysis for computation and message overheads, the analysis for theoretical compressibility, and simulations using real-world data demonstrate that our proposed scheme can provide significant communication/energy savings without sacrificing the accuracy of collected data. Azad Ali, Abdelmajid Khelil, Neeraj Suri, Mohammadreza Mahmudimanesh |
ACM Trans. Sens. Networks | 3 |
| 2015 | A Lease Based Hybrid Design Pattern for Proper-Temporal-Embedding of Wireless CPS InterlockingabstractCyber-Physical Systems (CPS) integrate discrete-time computing and continuous-time physical-world entities, which are often wirelessly interlinked. The use of wireless safety-critical CPS requires safety guarantees despite communication faults. This paper focuses on one important set of such safety rules: Proper-Temporal-Embedding (PTE), where distributed CPS entities must enter/leave risky states according to properly nested temporal pattern and certain duration spacing. Our solution introduces hybrid automata to formally describe and analyze CPS design patterns. We propose a novel leasing based design pattern, along with closed-form configuration constraints, to guarantee PTE safety rules under arbitrary wireless communication faults. We propose a formal procedure to transform the design pattern hybrid automata into specific wireless CPS designs. This procedure can effectively isolate physical world parameters from affecting the PTE safety of the resultant specific designs. We conduct two wireless CPS case studies, one on medicine and the other on control, to show that the resulted system is safe against communication failures. We also compare our approach with a polling based approach. Both approaches support PTE under arbitrary communication failures. The polling approach performs better under severely adverse wireless medium conditions; while ours performs better under benign or moderately adverse wireless medium conditions. Yufei Wang 0004, Qixin Wang 0001, Lei Bu, Neeraj Suri |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2014 | Efficient Agile Sink Selection in Wireless Sensor Networks Based on Compressed SensingabstractCollection of the sensed data in a wireless sensor network at one or more sink (s) is a well studied problem and there are a lot of efficient solutions for a variety of wireless sensor network configurations and application requirements. These methods are often optimized towards collection of the sensed data at a predetermined base station or sink. This inherently reduces the agility of the wireless sensor network as the flow of information is not easily changeable after the establishment of the routing and data collection algorithms. This paper presents an efficient data dissemination method based on the compressed sensing theory that allows each sensor node to take the role of a sink. Agile sink selection is especially advantageous in scenarios where the sink or the end user of the wireless sensor network is mobile. The proposed method allows availing the global state of the environment by fetching a small set of data from any arbitrary node. Our evaluations prove the better performance of our technique over existing methods. Also a comparison with an oracle-based approach gives sufficient experimental evidences of a nearly optimal performance of our method. Mohammadreza Mahmudimanesh, Amir Naseri, Neeraj Suri |
DCOSS | 3 |
| 2014 | Robust Compressive Data Gathering in Wireless Sensor Networks with Linear TopologyabstractWireless Sensor Networks (WSNs) are deployed in a variety of topologies and configurations depending on specific applications and requirements. In this paper, we study a simple and yet very important class of the WSN topologies, the linear or chain topology in which the Sensor Nodes (SNs) are connected in a series and gather the sensed data at a single base station or sink at the end of the chain. WSNs with linear topology have many practical applications, e.g., in infrastructure monitoring and surveillance of civil constructions. There is a large body of research on efficient data gathering techniques to transmit the sensed data over WSNs' limited communication bandwidth. In particular for linear topology, data collection technique has to put a balanced load on all SNs to avoid breakage of the chain at the exhausted nodes. Compressed Sensing (CS) is an efficient data collection technique for WSNs that fulfills these requirements. In a failure-free scenario, CS avoids exhausted nodes by balancing the communication and processing load on the SNs. In this paper, we examine the performance of a special implementation of CS for WSNs, called Compressive Data Gathering (CDG) when a SN in a chain WSN encounters a failure and cannot forward messages to the next hop. We propose a method to enhance the robustness of CDG in such failure scenarios by transmitting the messages to the next healthy node and excluding the failed samples from CS signal recovery mechanism. Evaluations show that our method effectively withstands the failures without sacrificing the accuracy of the collected data. Mohammadreza Mahmudimanesh, Neeraj Suri |
DCOSS | 2 |
| 2014 | An empirical study of injected versus actual interface errorsabstractThe reuse of software components is a common practice in commercial applications and increasingly appearing in safety critical systems as driven also by cost considerations. This practice puts dependability at risk, as differing operating conditions in different reuse scenarios may expose residual software faults in the components. Consequently, software fault injection techniques are used to assess how residual faults of reused software components may affect the system, and to identify appropriate counter-measures. As fault injection in components’ code suffers from a number of practical disadvantages, it is often replaced by error injection at the component interface level. However, it is still an open issue, whether such injected errors are actually representative of the effects of residual faults. To this end, we propose a method for analyzing how software faults turn into interface errors, with the ultimate aim of supporting more representative interface error injection experiments. Our analysis in the context of widely used software libraries reveals that existing interface error models are not suitable for emulating software faults, and provides useful insights for improving the representativeness of interface error injection. Anna Lanzaro, Roberto Natella, Stefan Winter 0001, Domenico Cotroneo, Neeraj Suri |
ISSTA | 5 |
| 2014 | Towards a Framework for Benchmarking Privacy-ABC Technologies
Fatbardh Veseli, Tsvetoslava Vateva-Gurova, Ioannis Krontiris, Kai Rannenberg, Neeraj Suri |
SEC | 5 |
| 2014 | Agile sink selection in wireless sensor networksabstractThe conventional setup of a wireless sensor network is composed of several sensor nodes and one or more sinks. The network topology and data collection techniques are then optimized towards efficient collection of the sensed data at the sink(s). In this paper, we present a novel network coding technique based on Compressed Sensing that allows each node to operate as a sink. Our network coding technique efficiently disseminates a number of linear combinations of the sensed data. After the dissemination phase, the entire sensed data is available by querying any node of the network. This is especially useful for distributed control using a wireless sensor and actuator network or in scenarios where the end user needs to access the global state of the environment from any node in its vicinity; e.g., when the end user is mobile. Mohammadreza Mahmudimanesh, Neeraj Suri |
SECON | 2 |
| 2014 | Towards a Framework for Assessing the Feasibility of Side-channel Attacks in Virtualized EnvironmentsabstractPhysically co-located virtual machines should be securely isolated from one another, as well as from the underlying layers in a virtualized environment. In particular the virtualized environment is supposed to guarantee the impossibility of an adversary to attack a virtual machine e.g., by exploiting a side-channel stemming from the usage of shared physical or software resources. However, this is often not the case and the lack of sufficient logical isolation is considered a key concern in virtualized environments. In the academic world this view has been reinforced during the last years by the demonstration of sophisticated side-channel attacks (SCAs). In this paper we argue that the feasibility of executing a SCA strongly depends on the actual context of the execution environment. To reflect on these observations, we propose a feasibility assessment framework for SCAs using cache based systems as an example scenario. As a proof of concept we show that the feasibility of cache-based side-channel attacks can be assessed following the proposed approach. Tsvetoslava Vateva-Gurova, Jesus Luna, Giancarlo Pellegrino, Neeraj Suri |
SECRYPT | 4 |
| 2014 | AHP-Based Quantitative Approach for Assessing and Comparing Cloud SecurityabstractWhile Cloud usage increasingly involves security considerations, there is still a conspicuous lack of techniques for users to assess/ensure that the security level advertised by the Cloud Service Provider (CSP) is actually delivered. Recent efforts have proposed extending existing Cloud Service Level Agreements (SLAs) to the security domain, by creating Security SLAs (SecLAs) along with attempts to quantify and reason about the security assurance provided by CSPs. However, both technical and usability issues limit their adoption in practice. In this paper we introduce a new technique for conducting quantitative and qualitative analysis of the security level provided by CSPs. Our methodology significantly improves upon contemporary security assessment approaches by creating a novel decision making technique based on the Analytic Hierarchy Process (AHP) that allows the comparison and benchmarking of the security provided by a CSP based on its SecLA. Furthermore, our technique improves security requirements specifications by introducing a flexible and simple methodology that allows users to identify their specific security needs. The proposed technique is demonstrated with real-world CSP data obtained from the Cloud Security Alliance's Security, Trust and Assurance Registry. Ahmed Taha 0002, Rubén Trapero, Jesus Luna, Neeraj Suri |
TrustCom | 4 |
| 2014 | Assessing the security of internet-connected critical infrastructuresabstractABSTRACT Because the Internet of Things (IoT) pervasively extends to all facets of life, the “things” are increasingly extending to include the interconnection of the Internet to critical infrastructures (CIs) such as telecommunication, power grid, transportation, e‐commerce systems, and others. The objective of this paper is twofold: (i) addressing IoT from a CI protection (CIP) and connectivity viewpoint, and (ii) highlighting the need for security quantification to improve the quality of protection (QoP) of CIs. Using a financial infrastructure as an example, a CIP and trust quantification perspective is built up. To this end, we are developing a novel security metrics‐based approach to assess and thereon enhance the CIP. We focus on the communication level of the CI where IoT is playing an increasingly important role with respect to sensing and communication across CI elements. Determining the security and dependability level of the communication over the CI constitutes a basic precondition for assessing the QoP of the whole CI, which is needed for any efforts to improve this QoP. Because metrics play a central role for such quantification, this paper develops their QoP use from an IoT perspective, and a reference implementation along with experimental results is presented. Copyright © 2012 John Wiley & Sons, Ltd. Hamza Ghani, Abdelmajid Khelil, Neeraj Suri, György Csertán, László Gönczy, Gábor Urbanics, James Clarke |
Secur. Commun. Networks | 3 |
| 2013 | PoWerStore: proofs of writing for efficient and robust storageabstractExisting Byzantine fault tolerant (BFT) storage solutions that achieve strong consistency and high availability, are costly compared to solutions that tolerate simple crashes. This cost is one of the main obstacles in deploying BFT storage in practice. Dan Dobre, Ghassan Karame, Wenting Li 0001, Matthias Majuntke, Neeraj Suri, Marko Vukolic |
CCS | 5 |
| 2013 | Negotiating and Brokering Cloud Resources based on Security Level Agreements
Jesus Luna, Tsvetoslava Vateva-Gurova, Neeraj Suri, Massimiliano Rak, Loredana Liccardo |
CLOSER | 3 |
| 2013 | Security as a Service Using an SLA-Based Approach via SPECSabstractThe cloud offers attractive options to migrate corporate applications, without any implication for the corporate security manager to manage or to secure physical resources. While this ease of migration is appealing, several security issues arise: can the validity of corporate legal compliance regulations still be ensured for remote data storage? How is it possible to assess the Cloud Service Provider (CSP) ability to meet corporate security requirements? Can one monitor and enforce the agreed cloud security levels? Unfortunately, no comprehensive solutions exist for these issues. In this context, we introduce a new approach, named SPECS. It aims to offer mechanisms to specify cloud security requirements and to assess the security features offered by CSPs, and to integrate the desired security services (e.g., credential and access management) into cloud services with a Security-as-a-Service approach. Furthermore, SPECS intends to provide systematic approaches to negotiate, to monitor and to enforce the security parameters specified in Service Level Agreements (SLA), to develop and to deploy security services that are cloud SLA-aware and are implemented as an open-source Platform-as-a-Service (PaaS). This paper introduces the main concepts of SPECS. Massimiliano Rak, Neeraj Suri, Jesus Luna, Dana Petcu, Valentina Casola, Umberto Villano |
CloudCom (2) | 2 |
| 2013 | Predictive vulnerability scoring in the context of insufficient information availabilityabstractMultiple databases and repositories exist for collecting known vulnerabilities for different systems and on different levels. However, it is not unusual that extensive time elapses, in some cases more than a year, in order to collect all information needed to perform/publish vulnerability scoring calculations for security management groups to assess and prioritize vulnerabilities for remediation. Scoring a vulnerability also requires broad knowledge about its characteristics, which is not always provided. As an alternative, this paper targets the quantitative understanding of security vulnerabilities in the context of insufficient vulnerability information. We propose a novel approach for the predictive assessment of security vulnerabilities, taking into consideration the relevant scenarios, e.g., zero day vulnerabilities, for which there is limited or no information to perform a typical vulnerability scoring. We propose a new analytical model, the Vulnerability Assessment Model (VAM), which is inspired by the Linear Discriminant Analysis and uses publicly available vulnerability databases such as the National Vulnerability Database (NVD) as a training data set. To demonstrate the applicability of our approach, we have developed a publicly available web application, the VAM Calculator. The experimental results obtained using real-world vulnerability data from the three most widely used Internet browsers show that by reducing the amount of required vulnerability information by around 50%, we can maintain the misclassification rate at approximately 5%. Hamza Ghani, Jesus Luna, Abdelmajid Khelil, Najib Alkadri, Neeraj Suri |
CRiSIS | 5 |
| 2013 | Quantitative assessment of software vulnerabilities based on economic-driven security metricsabstractVulnerability exploits cost organizations large amounts of resources, mainly due to disruption of ICT services, and thus loss of confidentiality, integrity and availability. As security managers in the industry usually have to operate with limited budgets allocated to information security, they need to prioritize their investment efforts regarding the response mechanisms to the existing vulnerabilities. The utilization of quantitative security vulnerability assessment methods enables efficient prioritization of security efforts and investments to mitigate the discovered vulnerabilities and thus an opportunity to lower expected losses. State of the art approaches for vulnerability assessment such as the Common Vulnerability Scoring System (CVSS), which is the de facto standard quantifying the severity of vulnerabilities, do not consider the economic impact in case of a vulnerability exploit. To this end, our paper targets the quantitative understanding of vulnerability severity taking into account the potential economic damage a successful vulnerability exploit can cause. We propose a novel approach for a systematic consideration of the relevant cost units (associated costs) for the economic damage estimation of vulnerability exploits. Our approach utilizes Multiple Criteria Decision Analysis (MCDA) methods to perform a prioritization of the existing vulnerabilities within the target system. The evaluation results show the potential cost savings w.r.t. the mitigation costs using our approach. Our method supports managers and decision makers in the process of prioritizing security investments to mitigate the discovered vulnerabilities. Hamza Ghani, Jesus Luna, Neeraj Suri |
CRiSIS | 3 |
| 2013 | Guaranteeing Proper-Temporal-Embedding safety rules in wireless CPS: A hybrid formal modeling approachabstractCyber-Physical Systems (CPS) integrate discrete-time computing and continuous-time physical-world entities, which are often wirelessly interlinked. The use of wireless safety critical CPS (control, healthcare etc.) requires safety guarantees despite communication faults. This paper focuses on one important set of such safety rules: Proper-Temporal-Embedding (PTE). Our solution introduces hybrid automata to formally describe and analyze CPS design patterns. We propose a novel lease based design pattern, along with closed-form configuration constraints, to guarantee PTE safety rules under arbitrary wireless communication faults. We propose a formal methodology to transform the design pattern hybrid automata into specific wireless CPS designs. This methodology can effectively isolate physical world parameters from affecting the PTE safety of the resultant specific designs. We conduct a case study on laser tracheotomy wireless CPS to show that the resulting system is safe and can withstand communication disruptions. Yufei Wang 0004, Qixin Wang 0001, Lei Bu, Rong Zheng 0001, Neeraj Suri |
DSN | 6 |
| 2013 | simFI: From single to simultaneous software fault injectionsabstractSoftware-implemented fault injection (SWIFI) is an established experimental technique to evaluate the robustness of software systems. While a large number of SWIFI frameworks exist, virtually all are based on a single-fault assumption, i.e., interactions of simultaneously occurring independent faults are not investigated. As software systems containing more than a single fault often are the norm than an exception [1] and current safety standards require the consideration of “multi-point faults” [2], the validity of this single-fault assumption is at question for contemporary software systems. To address the issue and support simultaneous SWIFI (simFI), we analyze how independent faults can manifest in a generic software composition model and extend an existing SWIFI tool to support some characteristic simultaneous fault types. We implement three simultaneous fault models and demonstrate their utility in evaluating the robustness of the Windows CE kernel. Our findings indicate that simultaneous fault injections prove highly efficient in triggering robustness vulnerabilities. Stefan Winter 0001, Michael Tretter, Benjamin Sattler, Neeraj Suri |
DSN | 4 |
| 2013 | Efficient Verification of Distributed Protocols Using Stateful Model CheckingabstractThis paper presents efficient model checking of distributed software. Key to the achieved efficiency is a novel stateful model checking strategy that is based on the decomposition of states into a relevant and an auxiliary part. We formally show this strategy to be sound, complete, and terminating for general finite-state systems. As a case study, we implement the proposed strategy within Basset/MP-Basset, a model checker for message-passing Java programs. Our evaluation with actual deployed fault-tolerant message-passing protocols shows that the proposed stateful optimization is able to reduce model checking time and memory by up to 69% compared to the naive stateful search, and 39% compared to partial-order reduction. Habib Saissi, Péter Bokor, Can Arda Muftuoglu, Neeraj Suri, Marco Serafini |
SRDS | 4 |
| 2013 | GMTC: A Generalized Commit Approach for Hybrid Mobile EnvironmentsabstractMobile environments increasingly require distributed atomic transactions to support the growing diversity of financial, gaming, social networking and many other applications. The underlying mobile infrastructure is correspondingly evolving with increasingly diverse wired and wireless elements and also with increasing exposure to a variety of operational perturbations at the mobile elements and communication levels. Consequently, the challenge is not only in providing efficient nonblocking mobile commit (as a fundamental basis behind consistent mobile transactions) but to also provide efficient perturbation-resilient atomic commit in the heterogeneous mobile space. The contribution of this paper is in developing a perturbation-resilient mobile commit protocol that efficiently provides for and preserves strict atomicity for transactional applications. The protocol does not necessarily require access to the powerful communication/computation elements of the wired infrastructure during transaction execution. However, in case access to a wired network becomes possible, it then adapts to utilize this to 1) increase the resilience to network perturbations achieving higher commit rates, and 2) reduce the wireless message overhead and the blocking of transaction participants leading to higher transactions throughput. In contrast, existing solutions are often tailored either for 1) infrastructure-based mobile environments, or 2) infrastructure-less ad hoc networks. To our knowledge, there is no existing commit protocol that can adapt across diverse infrastructure communication modes. The proposed perturbation-resilient generalized mobile transaction commit (GMTC) protocol represents the first atomic commit protocol for hybrid mobile environments which 1) takes advantage of accessing infrastructures, by choosing reliable infrastructure nodes for coordination of transactions and for replication of commit data of mobile participants to tolerate network disconnections, and 2) tolerates network partitioning and delivers best-effort resultsâin terms of transaction commit rate, message complexity, and commit/abort decision time (latency)âif the access to wired infrastructure is unavailable. The protocol performance simulations (covering transaction commit rate, message complexity, and commit/abort decision time) demonstrate the effectiveness of the developed protocol in generalized mobile environments. Brahim Ayari, Abdelmajid Khelil, Neeraj Suri |
IEEE Trans. Mob. Comput. | 3 |
| 2012 | Privacy-by-design based on quantitative threat modelingabstractWhile the general concept of “Privacy-by-Design (PbD)” is increasingly a popular one, there is considerable paucity of either rigorous or quantitative underpinnings supporting PbD. Drawing upon privacy-aware modeling techniques, this paper proposes a quantitative threat modeling methodology (QTMM) that can be used to draw objective conclusions about different privacy-related attacks that might compromise a service. The proposed QTMM has been empirically validated in the context of the EU project ABC4Trust, where the end-users actually elicited security and privacy requirements of the so-called privacy-Attribute Based Credentials (privacy-ABCs) in a real-world scenario. Our overall objective, is to provide architects of privacy-respecting systems with a set of quantitative and automated tools to help decide across functional system requirements and the corresponding trade-offs (security, privacy and economic), that should be taken into account before the actual deployment of their services. Jesus Luna, Neeraj Suri, Ioannis Krontiris |
CRiSIS | 2 |
| 2012 | Instrumenting AUTOSAR for dependability assessment: A guidance frameworkabstractThe AUTOSAR standard guides the development of component-based automotive software. As automotive software typically implements safety-critical functions, it needs to fulfill high dependability requirements, and the effort put into the quality assurance of these systems is correspondingly high. Testing, fault injection (FI), and other techniques are employed for the experimental dependability assessment of these increasingly software-intensive systems. Having flexible and automated support for instrumentation is key in making these assessment techniques efficient. However, providing a usable, customizable and performant instrumentation for AUTOSAR is non-trivial due to the varied abstractions and high complexity of these systems. This paper develops a dependability assessment guidance framework tailored towards AUTOSAR that helps identify the applicability and effectiveness of instrumentation techniques at (a) varied levels of software abstraction and granularity, (b) at varied software access levels - black-box, grey-box, white-box, and (c) the application of interface wrappers for conducting FI. Thorsten Piper, Stefan Winter 0001, Paul Manns, Neeraj Suri |
DSN | 4 |
| 2012 | Balanced spatio-temporal compressive sensing for multi-hop wireless sensor networksabstractCompressive Sampling (CS) is a powerful sampling technique that allows accurately reconstructing a compressible signal from a few random linear measurements. CS theory has applications in sensory systems where acquiring individual samples is either expensive or infeasible. A Wireless Sensor Network (WSN) is a distributed sensory system comprised of resource-limited sensor nodes. Transferring all the recorded samples in a WSN can easily result in data traffic that can exceed the network capacity. There are ongoing attempts to devise efficient and accurate compression schemes for WSNs and CS has proved to be a key sampling method compared to many other existing techniques. In this paper, specifically targeting the dominant WSN deployments of multi-hop WSNs, we develop a novel CS-based concept of sampling window as an efficient spatio-temporal signal acquisition/compression technique. We show that much higher energy-efficient signal acquisition is possible, if composite temporal and spatial correlations are considered. Our model is also capable of abnormal event detection which is a crucial feature in WSNs. It guarantees balanced energy consumption by the sensor nodes in a multi-hop topology to prevent overloaded nodes and network partitioning. Mohammadreza Mahmudimanesh, Abdelmajid Khelil, Neeraj Suri |
MASS | 3 |
| 2012 | Quantitative Assessment of Cloud Security Level Agreements - A Case Study
Jesus Luna, Hamza Ghani, Tsvetoslava Vateva-Gurova, Neeraj Suri |
SECRYPT | 4 |
| 2012 | Susceptibility Analysis of Structured P2P Systems to Localized Eclipse AttacksabstractPeer-to-Peer (P2P) protocols are susceptible to Localized Eclipse Attacks (LEA), i.e., attacks where a victim peer's environment is masked by malicious peers which are then able to instigate progressively insidious security attacks. To obtain effective placement of malicious peers, LEAs significantly benefit from overlay topology-awareness. Hence, we propose heuristics for Chord, Pastry and Kademlia to assess the protocols' LEA susceptibility based on their topology characteristics and overlay routing mechanisms. As a result, our method can be used for P2P protocol parameter tuning in order to substantially mitigate LEAs. We present evaluations highlighting LEA's impact on contemporary P2P protocols. Our proposed heuristics are abstract in nature, making them applicable plus customizable for many other structured P2P protocols. We validate our model's accuracy through a simulation case study. Daniel Germanus, Robert Langenberg, Abdelmajid Khelil, Neeraj Suri |
SRDS | 4 |
| 2012 | Brief Announcement: MP-State: State-Aware Software Model Checking of Message-Passing Systems
Can Arda Muftuoglu, Péter Bokor, Neeraj Suri |
SSS | 3 |
| 2011 | Efficient model checking of fault-tolerant distributed protocolsabstractTo aid the formal verification of fault-tolerant distributed protocols, we propose an approach that significantly reduces the costs of their model checking. These protocols often specify atomic, process-local events that consume a set of messages, change the state of a process, and send zero or more messages. We call such events quorum transitions and leverage them to optimize state exploration in two ways. First, we generate fewer states compared to models where quorum transitions are expressed by single-message transitions. Second, we refine transitions into a set of equivalent, finer-grained transitions that allow partial-order algorithms to achieve better reduction. We implement the MP-Basset model checker, which supports refined quorum transitions. We model check protocols representing core primitives of deployed reliable distributed systems, namely: Paxos consensus, regular storage, and Byzantine-tolerant multicast. We achieve up to 92% memory and 85% time reduction compared to model checking with standard unrefined single-message transitions. Péter Bokor, Johannes Kinder, Marco Serafini, Neeraj Suri |
DSN | 4 |
| 2011 | The impact of fault models on software robustness evaluationsabstractFollowing the design and in-lab testing of software, the evaluation of its resilience to actual operational perturbations in the field is a key validation need. Software-implemented fault injection (SWIFI) is a widely used approach for evaluating the robustness of software components. Recent research [24, 18] indicates that the selection of the applied fault model has considerable influence on the results of SWIFI-based evaluations, thereby raising the question how to select appropriate fault models (i.e. that provide justified robustness evidence). This paper proposes several metrics for comparatively evaluating fault models's abilities to reveal robustness vulnerabilities. It demonstrates their application in the context of OS device drivers by investigating the influence (and relative utility) of four commonly used fault models, i.e. bit flips (in function parameters and in binaries), data type dependent parameter corruptions, and parameter fuzzing. We assess the efficiency of these models at detecting robustness vulnerabilities during the SWIFI evaluation of a real embedded operating system kernel and discuss application guidelines for our metrics alongside. Stefan Winter 0001, Constantin Sârbu, Neeraj Suri, Brendan Murphy |
ICSE | 3 |
| 2011 | Assessing the comparative effectiveness of map construction protocols in wireless sensor networksabstractMap-based visualization of live spatial sensor data collected from a Wireless Sensor Networks (WSN) is a promising application. This application requires samples from each sensor node. The specific WSN properties such as energy imply that the data collection and/or in-network processing should carefully address the trade-off between the energy/network overhead and the map accuracy. Accordingly, several map construction approaches have been proposed. Their efficiency/accuracy balance has been proved in the corresponding papers for a set of carefully selected parameters. To date there exists no quantitative comparative study that considers a wide range of protocol and network parameters in order to provide a comprehensive evaluation of existing strategies. In this paper, we present the first study to evaluate the comparative effectiveness of the available map construction approaches. Abdelmajid Khelil, Hanbin Chang, Neeraj Suri |
IPCCC | 3 |
| 2011 | Supporting domain-specific state space reductions through local partial-order reductionabstractModel checkers offer to automatically prove safety and liveness properties of complex concurrent software systems, but they are limited by state space explosion. Partial-Order Reduction (POR) is an effective technique to mitigate this burden. However, applying existing notions of POR requires to verify conditions based on execution paths of unbounded length, a difficult task in general. To enable a more intuitive and still flexible application of POR, we propose local POR (LPOR). LPOR is based on the existing notion of statically computed stubborn sets, but its locality allows to verify conditions in single states rather than over long paths. As a case study, we apply LPOR to message-passing systems. We implement it within the Java Pathfinder model checker using our general Java-based LPOR library. Our experiments show significant reductions achieved by LPOR for model checking representative message-passing protocols and, maybe surprisingly, that LPOR can outperform dynamic POR. Péter Bokor, Johannes Kinder, Marco Serafini, Neeraj Suri |
ASE | 4 |
| 2011 | An adaptive and composite spatio-temporal data compression approach for wireless sensor networksabstractWireless Sensor Networks (WSN) are often deployed to sample the desired environmental attributes and deliver the acquired samples to the sink for processing, analysis or simulations as per the application needs. Many applications stipulate high granularity and data accuracy that results in high data volumes. Sensor nodes are battery powered and sending the requested large amount of data rapidly depletes their energy. Fortunately, the environmental attributes (e.g., temperature, pressure) often exhibit spatial and temporal correlations. Moreover, a large class of applications such as scientific measurement and forensics tolerate high latencies for sensor data collection. Accordingly, we develop a fully distributed adaptive technique for spatial and temporal in-network data compression with accuracy guarantees. We exploit the spatio-temporal correlation of sensor readings while benefiting from possible data delivery latency tolerance to further minimize the amount of data to be transported to the sink. Using real data, we demonstrate that our proposed scheme can provide significant communication/energy savings without sacrificing the accuracy of collected data. In our simulations, we achieved data compression of up to 95% on the raw data requiring around 5% of the original data to be transported to the sink. Azad Ali, Abdelmajid Khelil, Piotr Szczytowski, Neeraj Suri |
MSWiM | 4 |
| 2011 | Fork-Consistent Constructions from Registers
Matthias Majuntke, Dan Dobre, Christian Cachin, Neeraj Suri |
OPODIS | 4 |
| 2011 | The complexity of robust atomic storageabstractWe study the time-complexity of robust atomic read/write storage from fault-prone storage components in asynchronous message-passing systems. Robustness here means wait-free tolerating the largest possible number t of Byzantine storage component failures (optimal resilience) without relying on data authentication. We show that no single-writer multiple-reader (SWMR) robust atomic storage implementation exists if (a) read operations complete in less than four communication round-trips (rounds), and (b) the time complexity of write operations is constant. More precisely, we present two lower bounds. The first is a read lower bound stating that three rounds of communication are necessary to read from a SWMR robust atomic storage. The second is a write lower bound, showing that Ω(log(t)) write rounds are necessary to read in three rounds from such a storage. Applied to known results, our lower bounds close a fundamental gap: we show that time-optimal robust atomic storage can be obtained using well-known transformations from regular to atomic storage and existing time-optimal regular storage implementations. © 2011 ACM. Dan Dobre, Rachid Guerraoui, Matthias Majuntke, Neeraj Suri, Marko Vukolic |
PODC | 4 |
| 2011 | Fork-consistent constructions from registersabstractSo far, all implementations providing fork-consistent semantics are based on objects with read-modify-write capabilities (also termed servers). We propose constructions of fork-consistent shared objects from single-writer multiple-reader(SWMR) read/write base registers, that are strictly weaker than servers. Our shared object constructions provide linearizability if all base registers behave correctly, and gracefully degrade to either fork-linearizability or weak fork-linearizability if any number of registers fails Byzantine. We make the following contributions: (a) A fork-linearizable construction of a universal type where operations are allowed to abort under concurrency, and (b) a weak fork-linearizable implementation of a shared memory that ensures wait-freedom when the registers are correct. Matthias Majuntke, Dan Dobre, Neeraj Suri |
PODC | 3 |
| 2011 | TOM: Topology oriented maintenance in sparse Wireless Sensor NetworksabstractThe physical number of sensor nodes constitutes a major cost factor for Wireless Sensor Networks (WSN) deployments. Hence, a natural goal is to minimize the number of sensor nodes to be deployed, while still maintaining the desired properties of the WSN. However, sparse networks even while connected, usually suffer from topology irregularities that negatively impact the network lifetime and responsiveness, i.e., sensor data delivery reliability and latency. In addition, sensor node failures easily complicate/enforce/aggravate these irregularities. Valuable efforts have been conducted to discover topology specific anomalies such as coverage holes or critical/bottleneck nodes. Unfortunately, these efforts suffer from at least one of the following drawbacks: (a) They are centralized and consequently inefficient in large-scale networks, (b) they are tailored to one class of anomalies, or (c) do not propose how to remedy the identified anomaly. In this paper, we focus on sparse WSN which usually show varied topology irregularities and propose an in-network and localized strategy that efficiently (i) discovers generic topology irregularities, and (ii) identifies locations for minimal number of new augmented sensor deployments to remedy topology irregularities and sustain the desired operational requirements. We show the effectiveness and efficiency of the solution through a set of extensive simulations. Piotr Szczytowski, Abdelmajid Khelil, Azad Ali, Neeraj Suri |
SECON | 4 |
| 2011 | A Security Metrics Framework for the Cloud
Jesus Luna, Hamza Ghani, Daniel Germanus, Neeraj Suri |
SECRYPT | 4 |
| 2011 | Application-Level Diagnostic and Membership Protocols for Generic Time-Triggered SystemsabstractWe present online tunable diagnostic and membership protocols for generic time-triggered (TT) systems to detect crashes, send/receive omission faults, and network partitions. Compared to existing diagnostic and membership protocols for TT systems, our protocols do not rely on the single-fault assumption and also tolerate non-fail-silent (Byzantine) faults. They run at the application level and can be added on top of any TT system (possibly as a middleware component) without requiring modifications at the system level. The information on detected faults is accumulated using a penalty/reward algorithm to handle transient faults. After a fault is detected, the likelihood of node isolation can be adapted to different system configurations, including configurations where functions with different criticality levels are integrated. All protocols are formally verified using model checking. Using actual automotive and aerospace parameters, we also experimentally demonstrate the transient fault handling capabilities of the protocols. Marco Serafini, Péter Bokor, Neeraj Suri, Jonny Vinter, Astrit Ademaj, Wolfgang Brandstätter, Fulvio Tagliabo, Jens Koch |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2011 | On the design of perturbation-resilient atomic commit protocols for mobile transactionsabstractDistributed mobile transactions utilize commit protocols to achieve atomicity and consistent decisions. This is challenging, as mobile environments are typically characterized by frequent perturbations such as network disconnections and node failures. On one hand environmental constraints on mobile participants and wireless links may increase the resource blocking time of fixed participants. On the other hand frequent node and link failures complicate the design of atomic commit protocols by increasing both the transaction abort rate and resource blocking time. Hence, the deployment of classical commit protocols (such as two-phase commit) does not reasonably extend to distributed infrastructure-based mobile environments driving the need for perturbation-resilient commit protocols. In this article, we comprehensively consider and classify the perturbations of the wireless infrastructure-based mobile environment according to their impact on the outcome of commit protocols and on the resource blocking times. For each identified perturbation class a commit solution is provided. Consolidating these subsolutions, we develop a family of fault-tolerant atomic commit protocols that are tunable to meet the desired perturbation needs and provide minimized resource blocking times and optimized transaction commit rates. The framework is also evaluated using simulations and an actual testbed deployment. Brahim Ayari, Abdelmajid Khelil, Neeraj Suri |
ACM Trans. Comput. Syst. | 3 |
| 2010 | Scrooge: Reducing the costs of fast Byzantine replication in presence of unresponsive replicasabstractByzantine-Fault-Tolerant (BFT) state machine replication is an appealing technique to tolerate arbitrary failures. However, Byzantine agreement incurs a fundamental trade-off between being fast (i.e. optimal latency) and achieving optimal resilience (i.e. 2f + b + 1 replicas, where f is the bound on failures and b the bound on Byzantine failures). Achieving fast Byzantine replication despite f failures requires at least f + b - 2 additional replicas. In this paper we show, perhaps surprisingly, that fast Byzantine agreement despite f failures is practically attainable using only b - 1 additional replicas, which is independent of the number of crashes tolerated. This makes our approach particularly appealing for systems that must tolerate many crashes (large f) and few Byzantine faults (small b). The core principle of our approach is to have replicas agree on a quorum of responsive replicas before agreeing on requests. This is key to circumventing the resilience lower bound of fast Byzantine agreement. Marco Serafini, Péter Bokor, Dan Dobre, Matthias Majuntke, Neeraj Suri |
DSN | 5 |
| 2010 | LEHP: Localized energy hole profiling in Wireless Sensor NetworksabstractWireless Sensor Networks (WSN) display nonuniform energy usage distribution. This is mainly induced by the sink centric traffic or by non-uniform distribution of sensing activities and manifests as energy holes throughout the WSN. Holes can threaten the availability of the WSN by network partitioning and sensing voids. They are hard to predict, and consequently, proper function of the network requires systematic maintenance. Unfortunately, existing approaches do not systematically profile holes and focus only on very specific type of holes. In this work we present new distributed energy profiling algorithms for generalized types of energy holes. The algorithms search for boundary nodes and use them as a reference to calculate the energy needs of nodes within the hole. These, when aggregated, create angular and radial energy profiles. Extensive simulations show that the algorithms, when used for WSN maintenance, significantly help to extend the lifetime of the network. Piotr Szczytowski, Abdelmajid Khelil, Neeraj Suri |
ISCC | 3 |
| 2010 | ParTAC: A Partition-Tolerant Atomic Commit Protocol for MANETsabstractThe support of distributed atomic transactions in Mobile Ad-hoc Networks (MANET) is a key requirement for many mobile application scenarios. Atomicity is a fundamental property that ensures that all nodes decide a consistent outcome. As MANETs are characterized by frequent perturbations due to network partitioning and the fragility of nodes, providing atomicity is challenging. Existing protocols that ensure strict atomicity in MANETs are either bound to specific mobility pattern or based on building blocks such as consensus or group membership, not allowing arbitrary partitions or requiring exact knowledge about the members of a partition. These assumptions limit the deployment of these protocols to very restricted MANET scenarios, and may lead to poor commit rate, high message overhead or blocking related to intolerably long Commit/Abort decision times. In this paper, we present the first Partition-Tolerant Atomic Commit protocol (ParTAC) for MANETs which does not rely on consensus or group partition membership. As a consequence, ParTAC supports a significantly wider range of mobility patterns and partitioning scenarios than existing protocols. To reduce Commit/Abort decision times and prevent the protocol from blocking, ParTAC follows a best-effort strategy by defining a lifetime for every transaction after which the transaction is aborted. Further, we introduce a new coordination strategy based on a flexible preselection of multiple coordinators among the participating nodes. Thus, the failure of a single coordinator can be tolerated in the presence of network partitioning. Moreover, transactions can be aborted by any coordinator based on lifetime expiration. ParTAC is evaluated using simulations to demonstrate the performance of the protocol in terms of commit rate, message efficiency and Commit/Abort decision time. Brahim Ayari, Abdelmajid Khelil, Neeraj Suri |
Mobile Data Management | 3 |
| 2010 | Data-Based Agreement for Inter-vehicle CoordinationabstractData-based agreement is increasingly used to implement traceable coordination across mobile entities such as ad-hoc networked (autonomous) vehicles. In our work, we focus on data-based agreement using database transactions where mobile entities agree on a set of coordinated tasks that need to be performed by them in an atomic way. Atomicity means that all transaction participants agree on a set of tasks which will be performed by them or no one of them is performing any task. The data about the agreed tasks and their corresponding stakeholders are kept in local databases as a proof for the obtained agreement. This proof might be needed by users and regularities/authorities involved depending on the application scenario. In this demo, we demonstrate our effort to provide for partition-aware atomic commit protocols for transactional data-based agreement. Brahim Ayari, Abdelmajid Khelil, Kamel Saffar, Neeraj Suri |
Mobile Data Management | 4 |
| 2010 | Eventually linearizable shared objectsabstractLinearizability is the strongest known consistency property of shared objects. In asynchronous message passing systems, Linearizability can be achieved with ◊S and a majority of correct processes. In this paper we introduce the notion of Eventual Linearizability, the strongest known consistency property that can be attained with ◊S and any number of crashes. We show that linearizable shared object implementations can be augmented to support weak operations, which need to be linearized only eventually. Unlike strong operations that require to be always linearized, weak operations terminate in worst case runs. However, there is a tradeoff between ensuring termination of weak and strong operations when processes have only access to ◊S. If weak operations terminate in the worst case, then we show that strong operations terminate only in the absence of concurrent weak operations. Finally, we show that an implementation based on P exists that guarantees termination of all operations. Marco Serafini, Dan Dobre, Matthias Majuntke, Péter Bokor, Neeraj Suri |
PODC | 5 |
| 2010 | INDEXYS, a Logical Step beyond GENESYS
Andreas Eckel, Paul Milbredt, Zaid Al-Ars, Stefan Schneele, Bart Vermeulen, György Csertán, Christoph Scheerer, Neeraj Suri, Abdelmajid Khelil, Gerhard Fohler |
SAFECOMP | 8 |
| 2010 | Profiling the operational behavior of OS device drivers
Constantin Sârbu, Andréas Johansson, Neeraj Suri, Nachiappan Nagappan |
Empir. Softw. Eng. | 3 |
| 2010 | A software integration approach for designing and assessing dependable embedded systems
Neeraj Suri, Arshad Jhumka, Martin Hiller, András Pataricza, Shariful Islam, Constantin Sârbu |
J. Syst. Softw. | 1 |
| 2010 | Using Underutilized CPU Resources to Enhance Its ReliabilityabstractSoft errors (or transient faults) are temporary faults that arise in a circuit due to a variety of internal noise and external sources such as cosmic particle hits. Though soft errors still occur infrequently, they are rapidly becoming a major impediment to processor reliability. This is due primarily to processor scaling characteristics. In the past, systems designed to tolerate such faults utilized costly customized solutions, entailing the use of replicated hardware components to detect and recover from microprocessor faults. As the feature size keeps shrinking and with the proliferation of multiprocessor on die in all segments of computer-based systems, the capability to detect and recover from faults is also desired for commodity hardware. For such systems, however, performance and power constitute the main drivers, so the traditional solutions prove inadequate and new approaches are required. We introduce two independent and complementary microarchitecture-level techniques: double execution and double decoding. Both exploit the typically low average processor resource utilization of modern processors to enhance processor reliability. double execution protects the out-of-order part of the CPU by executing each instruction twice. Double decoding uses a second, low-performance low-power instruction decoder to detect soft errors in the decoder logic. These simple-to-implement techniques are shown to improve the processor's reliability with relatively low performance, power, and hardware overheads. Finally, the resulting ¿excessive¿ reliability can even be traded back for performance by increasing clock rate and/or reducing voltage, thereby improving upon single execution approaches. Avi Timor, Avi Mendelson, Yitzhak Birk, Neeraj Suri |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2009 | Role-Based Symmetry Reduction of Fault-Tolerant Distributed Protocols with Language Support
Péter Bokor, Marco Serafini, Neeraj Suri, Helmut Veith |
ICFEM | 3 |
| 2009 | Abortable Fork-Linearizable Storage
Matthias Majuntke, Dan Dobre, Marco Serafini, Neeraj Suri |
OPODIS | 4 |
| 2009 | Efficient Robust Storage Using Secret Tokens
Dan Dobre, Matthias Majuntke, Marco Serafini, Neeraj Suri |
SSS | 4 |
| 2009 | Brief Announcement: Efficient Model Checking of Fault-Tolerant Distributed Protocols Using Symmetry Reduction
Péter Bokor, Marco Serafini, Neeraj Suri, Helmut Veith |
DISC | 3 |
| 2008 | INcreasing Security and Protection through Infrastructure REsilience: The INSPIRE Project
Salvatore D'Antonio, Luigi Romano, Abdelmajid Khelil, Neeraj Suri |
CRITIS | 4 |
| 2008 | Dependable Embedded Systems Special Day Panel: Issues and Challenges in Dependable Embedded SystemsabstractThe paper presents a panel discussion on the issues and challenges in dependable embedded system from both the academic and industrial perspectives. The panelists are Jacob Abraham from the University of Texas at Austin-USA, Stefan Poledna from TTTech-Austria, Avi Mendelson from Intel-Israel, and Subhasish Mitra from Stanford University-USA. Neeraj Suri, Christof Fetzer, Jacob A. Abraham, Stefan Poledna, Avi Mendelson, Subhasish Mitra |
DATE | 1 |
| 2008 | Message from the DCCS program chairabstractPresents the introductory welcome message from the conference proceedings. Neeraj Suri |
DSN | 1 |
| 2008 | Profiling the Operational Behavior of OS Device DriversabstractAs the complexity of modern Operating Systems (OS) increases, testing key OS components such as device drivers(DD) becomes increasingly complex given the multitude of possible DD interactions. If representative operational activity profiles of DDs within an OS could be obtained, these could significantly improve the understanding of the actual operational DD state space towards guiding the test efforts. Focusing on characterizing DD operational activities, this paper proposes a quantitative technique for profiling the runtime behavior of DDs using a set of occurrence and temporal metrics obtained via I/O traffic characterization. Such profiles are used to improve test adequacy against real-world workloads by enabling similarity quantification across them. The profiles also reveal execution hotspots in terms of DD functionalities activated in the field, thus allowing for dedicated test campaigns. A case study on actual Windows drivers substantiates our proposed approach. Constantin Sârbu, Andréas Johansson, Neeraj Suri, Nachiappan Nagappan |
ISSRE | 3 |
| 2008 | On the Time-Complexity of Robust and Amnesic Storage
Dan Dobre, Matthias Majuntke, Neeraj Suri |
OPODIS | 3 |
| 2008 | A comparative study of data transport protocols in wireless sensor networksabstractSeveral proposals describing transport layer protocols for sensor networks appear in the literature. As each proposal is typically evaluated in the context of carefully selected parameters and scenarios, the benefits can be subjective. Also, given the limited details available of different proposals, it is difficult for developers of sensor network applications to select from the range of alternative transport protocols. This paper develops a common basis for evaluation of varied proposals. We first classify and review the existing protocols and evaluate them by measuring their performance in terms of responsiveness and efficiency in a conformal simulation environment and for a wide range of operational conditions. Common sources of poor performance are identified. Based on this experience, a set of design principles for the designers of applications and future transport protocols is presented. Faisal Karim Shaikh, Abdelmajid Khelil, Neeraj Suri |
WOWMOM | 3 |
| 2007 | On the Selection of Error Model(s) for OS Robustness EvaluationabstractThe choice of error model used for robustness evaluation of operating systems (OSs) influences the evaluation run time, implementation complexity, as well as the evaluation precision. In order to find an "effective" error model for OS evaluation, this paper systematically compares the relative effectiveness of three prominent error models, namely bit-flips, data type errors and fuzzing errors using fault injection at the interface between device drivers OS. Bit-flips come with higher costs (time) than the other models, but allow for more detailed results. Fuzzing is cheaper to implement but is found to be less precise. A composite error model is presented where the low cost of fuzzing is combined with the higher level of details of bit-flips, resulting in high precision with moderate setup and execution costs. Andréas Johansson, Neeraj Suri, Brendan Murphy |
DSN | 2 |
| 2007 | A Tunable Add-On Diagnostic Protocol for Time-Triggered SystemsabstractWe present a tunable diagnostic protocol for generic time-triggered (TT) systems to detect crash and send/receive omission faults. Compared to existing diagnostic and membership protocols for TT systems, it does not rely on the single-fault assumption and tolerates malicious faults. It runs at the application level and can be added on top of any TT system (possibly as a middleware component) without requiring modifications at the system level. The information on detected faults is accumulated using a penalty/reward algorithm to handle transient faults. After a fault is detected, the likelihood of node isolation can be adapted to different system configurations, including those where functions with different criticality levels are integrated. Using actual automotive and aerospace parameters, we experimentally demonstrate the transient fault handling capabilities of the protocol. Marco Serafini, Neeraj Suri, Jonny Vinter, Astrit Ademaj, Wolfgang Brandstätter, Fulvio Tagliabo, Jens Koch |
DSN | 2 |
| 2007 | A Multi Variable Optimization Approach for the Design of Integrated Dependable Real-Time Embedded Systems
Shariful Islam, Neeraj Suri |
EUC | 2 |
| 2007 | On the Impact of Injection Triggers for OS Robustness EvaluationabstractTraditionally, in fault injection-based robustness evaluation of software (specifically for operating systems - OS's), faults or errors are injected at specific code locations. This paper studies the sensitivity and accuracy of the robustness evaluation results arising from varying the timing of injecting the faults into the OS. A strategy to guide the triggering of fault injection is proposed, based on the observation that the operational usage profile of a driver shows a high degree of regularity in the calls being made. The concept of call blocks (i.e., a distinct sequence of calls made to the driver) can be used to guide injections into different system states, corresponding to the driver operations carried out. A real-world case study compares the effectiveness of the proposed strategy to traditional location-based approaches, demonstrating that significant and useful insights can be gained by modulating the injection instants. Andréas Johansson, Neeraj Suri, Brendan Murphy |
ISSRE | 2 |
| 2007 | On Modeling the Reliability of Data Transport in Wireless Sensor NetworksabstractData transport is a core function for wireless sensor networks (WSNs) with different applications having varied requirements on the reliability and timeliness of data delivery. While node redundancy, inherent in WSNs, increases the fault tolerance, no guarantees on reliability levels can be assured. Furthermore, the frequent failures within WSNs impact the observed reliability over time and make it more challenging to achieve the desired reliability. Unfortunately, a framework for modeling reliability of data transport protocols in WSNs is currently missing. The existence of such a framework would simplify evaluation, comparison and also adaptation of these protocols. We formulate the problem of data transport in a WSN as a set of operations carried out on raw data. The operations aim at filtering the raw data to streamline its reliable transport towards the sink. Based on this formulation we systematically define a reliability framework. This paper argues for the usefulness of the reliability framework by classifying existing transport protocols and comparing their reliability Faisal Karim Shaikh, Abdelmajid Khelil, Neeraj Suri |
PDP | 3 |
| 2007 | On the Latency Efficiency of Message-Parsimonious Asynchronous Atomic BroadcastabstractWe address the problem of message-parsimonious asynchronous atomic broadcast when a subset t out of n parties may exhibit byzantine behavior. Message parsimony involves using only the optimal O(n) message exchanges per atomically delivered payload in the normal case. Message parsimony is desirable for Internet-like deployment environments in which message loss rates are non-negligible. Protocol PABC, the only previously-known message-parsimonious solution, suffered from two limitations vis-a-vis the solutions with O(n2) message complexity: more communication steps and the use of digital signatures. We present a protocol termed AMP that for the first time provides signature-free message parsimony while at the same time reducing the number of communication steps to the minimum necessary. In contrast to many previous atomic broadcast solutions, our protocol satisfies both safety and liveness in the asynchronous model. Dan Dobre, HariGovind V. Ramasamy, Neeraj Suri |
SRDS | 3 |
| 2007 | The Fail-Heterogeneous Architectural ModelabstractFault tolerant distributed protocols typically utilize a homogeneous fault model, either fail-crash or fail-Byzantine, where all processors are assumed to fail in the same manner. In practice, due to complexity and evolvability reasons, only a subset of the nodes can actually be designed to have a restricted, fail-crash failure mode, provided that they are free of design faults. Based on this consideration, we propose a fail-heterogeneous architectural model for distributed systems which considers two classes of nodes: (a) full-fledged execution nodes, which can be fail-Byzantine, and (b) lightweight, validated coordination nodes, which can only be fail-crash. To illustrate the model we introduce HeterTrust as a practical trustworthy service replication protocol. It has a low latency overhead, requires few execution nodes with diversified design, and prevents intruded servers from disclosing confidential data. We also discuss applications of the model to DoS attacks mitigation and to group membership. Marco Serafini, Neeraj Suri |
SRDS | 2 |
| 2007 | Online Diagnosis and Recovery: On the Choice and Impact of Tuning ParametersabstractA sequenced process of Fault Detection followed by the erroneous node's Isolation and system Reconfiguration (node exclusion or recovery), that is, the FDIR process, characterizes the sustained operations of a fault-tolerant system. For distributed systems utilizing message passing, a number of diagnostic (and associated FDIR) approaches, including our prior algorithms, exist in literature and practice. Invariably, the focus is on proving the completeness and correctness (all and only the faulty nodes are isolated) for the chosen fault model, without explicitly segregating permanent from transient faulty nodes. To capture diagnostic issues related to the persistence of errors (transient, intermittent, and permanent), we advocate the integration of count-and-threshold mechanisms into the FDIR framework. Targeting pragmatic system issues, we develop an adaptive online FDIR framework that handles a continuum of fault models and diagnostic protocols and comprehensively characterizes the role of various probabilistic parameters that, due to the count-and-threshold approach, influence the correctness and completeness of diagnosis and system reliability such as the fault detection frequency. The FDIR framework has been implemented on two prototypes for automotive and aerospace applications. The tuning of the protocol parameters at design time allows a significant improvement with respect to prior design choices. Marco Serafini, Andrea Bondavalli, Neeraj Suri |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2006 | One-step Consensus with Zero-DegradationabstractIn the asynchronous distributed system model, consensus is obtained in one communication step if all processes propose the same value. Assumingf \lt n/3, this is regardless of the failure detector output. A zero-degrading protocol reaches consensus in two communication steps in every stable run, i.e., when the failure detector makes no mistakes and its output does not change. We show that no leaderbased consensus protocol can be simultaneously one-step and zero-degrading. We propose two approaches to circumvent the impossibility result and present corresponding consensus protocols. Further, we present an atomic broadcast protocol that has a latency of 3d in every stable run and a latency of 2d in case of no collisions. Finally, we evaluate its performance in a cluster of workstations. Dan Dobre, Neeraj Suri |
DSN | 2 |
| 2006 | Dependability Driven Integration of Mixed Criticality SW ComponentsabstractMapping of software onto hardware elements under platform resource constraints is a crucial step in the design of embedded systems. As embedded systems are increasingly integrating both safety-critical and non-safety critical software functionalities onto a shared hardware platform, a dependability driven integration is desirable. Such an integration approach faces new challenges of mapping software components onto shared hardware resources while considering extra-functional (dependability, timing, power consumption, etc.) requirements of the system. Considering dependability and real-time as primary drivers, we present a systematic resource allocation approach for the consolidated mapping of safety critical and non-safety critical applications onto a distributed platform such that their operational delineation is maintained over integration. The objective of our allocation technique is to come up with a feasible solution satisfying multiple concurrent constraints. Ensuring criticality partitioning, avoiding error propagation and reducing interactions across components are addressed in our approach. In order to demonstrate the usefulness and effectiveness of the mapping, the developed approach is applied to an actual automotive system Shariful Islam, Robert Lindstrom, Neeraj Suri |
ISORC | 3 |
| 2006 | FT-PPTC: An Efficient and Fault-Tolerant Commit Protocol for Mobile EnvironmentsabstractTransactions are required not only for wired networks but also for the emerging wireless environments where mobile and fixed hosts participate side by side in the execution of the transaction. This heterogenous environment is characterized by constraints in mobile host capabilities, network connectivity and also an increasing number of possible failure modes. Classical atomic commit protocols used in wired networks are therefore not directly suitable for this heterogenous environment. Furthermore, the few commit protocols designed for mobile transactions either consider mobile hosts only as initiators though not as active participants, or show a high resource blocking time. We present the Fault-Tolerant Pre-Phase Transaction Commit (FT-PPTC) protocol for mobile environments. FT-PPTC decouples the commit of mobile participants from that of fixed participants. Consequently, the commit set can be reduced to a set of entities in the fixed network. Thus, the commit can easily be supported by any traditional atomic commit protocol, such as the established 2PC protocol. We integrate fault-tolerance as a key feature of FT-PPTC. Performance evaluations confirm the efficiency, scalability and low resource blocking time of our approach Brahim Ayari, Abdelmajid Khelil, Neeraj Suri |
SRDS | 3 |
| 2005 | Error Propagation Profiling of Operating SystemsabstractAn operating system (OS) constitutes a fundamental software (SW) component of a computing system. The robustness of its operations, or lack thereof, strongly influences the robustness of the entire system. Targeting enhancement of robustness at the OS level via use of add-on SW wrappers, this paper presents an error propagation profiling framework that assists in a) systematic identification and location of design and operational vulnerabilities, and b) quantification of their potential impact. Focusing on data (value) errors occurring in OS drivers, a set of measures is presented that aids a designer to locate such vulnerabilities, either on an OS service (system call) basis or a per driver basis. A case study and associated experimental process, using Windows CE .Net, is presented outlining the utility of our proposed approach. Andréas Johansson, Neeraj Suri |
DSN | 2 |
| 2005 | Designing Efficient Fail-Safe Multitolerant Systems
Arshad Jhumka, Neeraj Suri |
FORTE | 2 |
| 2004 | TTET: Event-Triggered Channels on a Time-Triggered BaseabstractThis paper develops solutions for efficient transfer of sporadic (event-triggered) data over a time-triggered communication channel. We present novel and efficient techniques for composite provisioning of the event triggered (ET) and time triggered (TT) paradigms to achieve both the predictability and flexibility inherent in TT and ET systems respectively. As a tangible demonstration of the techniques, present and assess a variation of the TTP/C protocol, where minor operational changes allow for its efficient handling of sporadic message traffic. Vilgot Claesson, Neeraj Suri |
ICECCS | 2 |
| 2004 | Panel Summary StatementsabstractNext generation space-based systems will necessitate onboard high performance computing, which is the key to enabling spacecraft autonomy, onboard sensor/science data processing, and multi-spacecraft interactive-cooperative robotics. For example, while conceptually straightforward, “Internet in the sky” communications will require multiple GOPS (giga-operation per second) to perform high-speed routing and protocoltranslation at real-time rates. The viability of high performance space computing stems from the advent of new enabling technologies. Such enabling technologies encompass wireless network communications, non-semiconductor-based (e.g., magnetic, carbon nanotube, ferroelectric and MEMS) components, and deep submicron semiconductors. In addition, it is anticipated that high performance space computing will leads to 1) more extensive use of COTS (commercially-of-the-shelf) products and standards, and 2) the emergence of larger, more complex embedded software running on multi-threaded, file-oriented operating systems. Accordingly, high performance space computing will bring space system dependability concepts and challenges into new, more sophisticated settings. Moreover, as space-based systems are rapidly becoming pervasive and crucial to our worldwide infrastructure, high performance space computing will further increase our reliance on space system operations. Hence, we are reaching the point where failures of space-based systems will have far-reaching and potentially catastrophic consequences in such areas as hazardous weather prediction, communications, finance, aircraft control, military operations, homeland security, and disaster relief and recovery. With the above motivation, this panel brings distinguished researchers and practitioners with wide ranging expertise in space systems and applications to a discussion. The thrust is to foster debating, exchanging, and integrating opinions and solutions for dependable high performance space computing. We particularly solicit different views from various perspectives on the following issues: • What are the critical dependability issues in high performance space computing? • What types of fault tolerance strategies and capabilities that are currently being developed for terrestrial high performance computing systems will be applicable to high performance computing in space? • Among the various means (e.g., V&V, fault tolerance) for dependable high performance space computing, which should we give higher priority? • How do we define fault models and conduct benchmarking to predict COTS performance degradation in the presence of faults? Raphael R. Some, Algirdas Avizienis, Jiri Gaisler, Hirokazu Ihara, Shubu Mukherjee, Neeraj Suri |
PRDC | 6 |
| 2004 | Why Have Progresses in Real-Time Fault Tolerant Computing Been Slow?
K. H. (Kane) Kim, Paul D. Ezhilchelvan, Jörg Kaiser, Louise E. Moser, Edgar Nett, Neeraj Suri |
SRDS | 6 |
| 2004 | Why Progress in (Composite) Fault Tolerant Real-Time Systems has been Slow (-er than Expected.. & What Can We Do About It?)abstractThe pervasiveness of computers in our current IT driven society (transportation, e-commerce, e-transactions, communication, process control), also implies our growing dependency on their "correct" functionality. In many a case, the real value of the systems and also our usage of these systems comes, in part, based on the dependency (real or perceived) we are consequently willing to put into the provisioning of the services i.e., the implicit or explicit assurance of trust we put for sustained delivery of desired services. Some systems are considered as safety-critical (flight/reactor control etc), though others are accorded varied degrees of criticality. Nevertheless, our expectancy extends to obtaining the proper services when the system is fault-free and especially when it encounters perturbations (design or operational), e.g., electromagnetic interference or a lightning strike for an aircraft. Consequently, it is important to qualitatively and quantitatively associate some measures of trust in the system's ability to "actually" deliver us the desired services in the presence of faults. This is often termed as "dependability" measures for a system with a plethora of fault-tolerance (FT) strategies to help achieve desired levels of dependability. As before, dependability entails the sustained delivery of services, be they service-critical or cost-critical, regardless of the perturbations encountered during their operation. Neeraj Suri |
SRDS | 1 |
| 2004 | EPIC: Profiling the Propagation and Effect of Data Errors in SoftwareabstractWe present an approach for analyzing the propagation and effect of data errors in modular software enabling the profiling of the vulnerabilities of software to find 1) the modules and signals most likely exposed to propagating errors and 2) the modules and signals which, when subjected to error, tend to cause more damage than others from a systems operation point-of-view. We discuss how to use the obtained profiles to identify where dependability structures and mechanisms will likely be the most effective, i.e., how to perform a cost-benefit analysis for dependability. A fault-injection-based method for estimation of the various measures is described and the software of a real embedded control system is profiled to show the type of results obtainable by the analysis framework. Martin Hiller, Arshad Jhumka, Neeraj Suri |
IEEE Trans. Computers | 3 |
| 2004 | An Efficient TDMA Start-Up and Restart Synchronization Approach for Distributed Embedded SystemsabstractA desired attribute in safety-critical embedded real-time systems is a system time and event synchronization capability on which predictable communication can be established. Focusing on bus-based communication protocols, we present a novel, efficient, and low-cost start-up and restart synchronization approach for TDMA environments. This approach utilizes information about a node's message length that forms a unique sequence to achieve synchronization such that communication overhead can be avoided. We present a fault-tolerant initial synchronization protocol with a bounded start-up time. The protocol avoids start-up collisions by deterministically postponing retries after a collision. We also present a resynchronization strategy that incorporates recovering nodes into synchronization. Vilgot Claesson, Henrik Lönn, Neeraj Suri |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2003 | The Event-Triggered and Time-Triggered Medium-Access MethodsabstractThe processes of accessing a shared communication media have been extensively researched in the dependability and real-time area. For embedded systems, the primary approaches have revolved around the event-triggered and the time-triggered paradigms. In this paper, our goal is to objectively and quantitatively outline the capabilities and limitations of each of these paradigms. The event-triggered approach is commonly perceived as providing high flexibility, while the time-triggered approach is expected to provide for a higher degree of predictable communication access to the media. We have quantified the spread of their differences, and provide a summary discussion about suggested best usage for each approach. The focus of our work is on the response times of the communication system, and also on the schedulability of the communication system in collaboration with tasks in the nodes. Vilgot Claesson, Cecilia Ekelin, Neeraj Suri |
ISORC | 3 |
| 2003 | Compositional Design of RT Systems: A Conceptual Basis for Specification of Linking InterfacesabstractComposition of a system is driven by the (a) identification and specification of basic components, and (b) specification of the interactions across the components, i.e., the communication linkages, that are needed to communicate value and temporal information across the components from which the aggregate system results. This paper addresses compositional design of distributed Real-Time (RT) systems focusing specifically on the role of specification of linking interfaces (LIFs) across components. Hermann Kopetz, Neeraj Suri |
ISORC | 2 |
| 2003 | A Framework for the Design and Validation of Efficient Fail-Safe Fault-Tolerant Programs
Arshad Jhumka, Neeraj Suri, Martin Hiller |
SCOPES | 2 |
| 2003 | The customizable fault/error model for dependable distributed systems
Chris J. Walter, Neeraj Suri |
Theor. Comput. Sci. | 2 |
| 2002 | On the Placement of Software Mechanisms for Detection of Data ErrorsabstractAn important aspect in the development of dependable software is to decide where to locate mechanisms for efficient error detection and recovery. We present a comparison between two methods for selecting locations for error detection mechanisms, in this case executable assertions (EAs), in black-box, modular software. Our results show that by placing EAs based on error propagation analysis one may reduce the memory and execution time requirements as compared to experience- and heuristic-based placement while maintaining the obtained detection coverage. Further, we show the sensitivity of the EA-provided coverage estimation on the choice of the underlying error model. Subsequently, we extend the analysis framework such that error-model effects are also addressed and introduce measures for classifying signals according to their effect on system output when errors are present. The extended framework facilitates profiling of software systems from varied dependability perspectives and is also less susceptible to the effects of having different error models for estimating detection coverage. Martin Hiller, Arshad Jhumka, Neeraj Suri |
DSN | 3 |
| 2002 | PROPANE: an environment for examining the propagation of errors in softwareabstractIn order to produce reliable software, it is important to have knowledge on how faults and errors may affect the software. In particular, designing efficient error detection mechanisms requires not only knowledge on which types of errors to detect but also the effect these errors may have on the software as well as how they propagate through the software. This paper presents the Propagation Analysis Environment (PROPANE) which is a tool for profiling and conducting fault injection experiments on software running on desktop computers. PROPANE supports the injection of both software faults (by mutation of source code) and data errors (by manipulating variable and memory contents). PROPANE supports various error types out-of-the-box and has support for user-defined error types. For logging, probes are provided for charting the values of variables and memory areas as well as for registering events during execution of the system under test. PROPANE has a flexible design making it useful for development of a wide range of software systems, e.g., embedded software, generic software components, or user-level desktop applications. We show examples of results obtained using PROPANE and how these can guide software developers to where software error detection and recovery could increase the reliability of the software system. Martin Hiller, Arshad Jhumka, Neeraj Suri |
ISSTA | 3 |
| 2002 | A Control Theory Approach for Analyzing the Effects of Data Errors in Safety-Critical Control SystemsabstractComputers are increasingly used for implementing control algorithms in safety-critical embedded applications, such as engine control, braking control and flight surface control. Addressing the consequent coupling of control performance with computer related errors, this paper develops a composite computer dependability/control theory methodology for analyzing the effects data errors have on control system dependability. The effect is measured as the resulting control error (defined as the difference between the desired value of a physical properly and its actual value). We use maximum bounds on this measure as the criterion for control system failure (i.e., if the control error exceeds a certain threshold, the system has failed). In this paper we a) present suitable models of computer faults for analysis of control level effects and related analysis methods, and b) apply traditional control theory analysis methods for understanding the effects of data errors on system dependability An automobile slip-control brake-system is used as an example showing the viability of our approach. Örjan Askerdal, Magnus Gäfvert, Martin Hiller, Neeraj Suri |
PRDC | 4 |
| 2001 | An Approach for Analysing the Propagation of Data Errors in SoftwareabstractWe present a novel approach for analysing the propagation of data errors in software. The concept of error permeability is introduced as a basic measure upon which we define a set of related measures. These measures guide us in the process of analysing the vulnerability of software to find the modules that are most likely exposed to propagating errors. Based on the analysis performed with error permeability and its related measures, we describe how to select suitable locations for error detection mechanisms (EDMs) and error recovery mechanisms (ERMs). A method for experimental estimation of error permeability, based on fault injection, is described and the software of a real embedded control system analysed to show the type of results obtainable by the analysis framework. The results show that the developed framework is very useful for analysing error propagation and software vulnerability and for deciding where to place EDMs and ERMs. Martin Hiller, Arshad Jhumka, Neeraj Suri |
DSN | 3 |
| 2001 | Modular Composition of Redundancy Management Protocols in Distributed Systems: An Outlook on Simplifying Protocol Level Formal Specification & VerificationabstractIn recent years, formal methods (FMs) have been extensively used for the verification and validation (V&V) of dependable distributed protocols. In our studies utilizing FMs for V&V, we have observed that a number of protocols providing for distributed and dependable services can often be formulated using a small set of basic functional primitives or their variations. Thus, from the formal viewpoint, the objective of this paper is to introduce techniques, utilizing concepts of category theory, that could effectively identify and reuse basic formal modules in order to simplify formal specification and verification for a spectrum of protocols. Purnendu Sinha, Neeraj Suri |
ICDCS | 2 |
| 2001 | Efficient TDMA Synchronization for Distributed Embedded SystemsabstractA desired attribute in safety critical embedded real-time systems is a system time/event synchronization capability on which predictable communication can be established. Focusing on bus-based communication protocols in TDMA environments, we present a novel, efficient, and low-cost synchronization approach with bounded start-up time. This approach utilizes information about each node's unique message lengths to achieve synchronization. The protocol avoids start-up collisions by postponing retries after a collision. We also present a re-synchronization strategy that incorporates recovering nodes into synchronization. Vilgot Claesson, Henrik Lönn, Neeraj Suri |
SRDS | 3 |
| 2001 | Assessing Inter-Modular Error Propagation in Distributed SoftwareabstractWith the functionality of most embedded systems based on software (SW), interactions amongst SW modules arise, resulting in error propagation across them. During SW development, it would be helpful to have a framework that clearly demonstrates the error propagation and containment capabilities of the different SW components. In this paper, we assess the impact of inter-modular error propagation. Adopting a white-box SW approach, we make the following contributions: (a) we study and characterize the error propagation process and derive a set of metrics that quantitatively represents the inter-modular SW interactions, (b) we use a real embedded target system used in an aircraft arrestment system to perform fault-injection experiments to obtain experimental values for the metrics proposed, (c) we show how the set of metrics can be used to obtain the required analytical framework for error propagation analysis. We find that the derived analytical framework establishes a very close correlation between the analytical and experimental values obtained. The intent is to use this framework to be able to systematically develop SW such that inter-modular error propagation is reduced by design. Arshad Jhumka, Martin Hiller, Neeraj Suri |
SRDS | 3 |
| 2000 | Designing High-Performance & Reliable Superscalar Architectures: The out of Order Reliable Superscalar (O3RS) ApproachabstractAs VLSI geometry continues to shrink and the level of integration increases, it is expected that the probability of faults, particularly transient faults, will increase in future microprocessors. So far, fault tolerance has chiefly been considered for special purpose or safety critical systems, but future technology will likely require integrating fault tolerance techniques into commercial systems. Such systems require low cost solutions that are transparent to the system operation and do not degrade overall performance. This paper introduces a new superscalar architecture, termed as 03RS that aims to incorporate such simple fault tolerance mechanisms as part of the basic architecture. Avi Mendelson, Neeraj Suri |
DSN | 2 |
| 2000 | Evaluating COTS Standards for Design of Dependable SystemsabstractThis experience report presents a study on the fault tolerance (FT) support capabilities of various COTS standards prior to their inclusion in design of dependable systems. A standalone analysis and relative comparison of the FT attributes for SCI, ATM, Futurebus+ and Fiber Channel is presented. Chris J. Walter, Neeraj Suri, T. Monaghan |
DSN | 2 |
| 1999 | On the Use of Formal Techniques for Analyzing Dependable Real-Time ProtocolsabstractThe effective design of composite dependable and real time protocols entails demonstrating their proof of correctness and, in practice, the efficient delivery of services. We focus on these aspects of correctness and efficiency, specifically considering the real time aspects where the need is to ensure satisfaction of stringent timing and operational constraints. We establish the use of mathematically rigorous techniques such as formal methods (FMs) in not only providing for their traditional usage in establishing correctness checks, but also for their capability of assessing and analyzing timing requirements in dependable real time protocols. We present our perspectives in utilizing FMs in developing exact case analyses of fault tolerant and real time protocols. We discuss the insights obtained and flaws identified in the hand analysis over the process of formally analyzing and verifying the correctness of an existing fault tolerant real time scheduling protocol. Purnendu Sinha, Neeraj Suri |
RTSS | 2 |
| 1999 | Editorial: Special Section on Dependable Real-Time SystemsabstractVER time, the discipline of computing systems has evolved and expanded in a multifaceted manner to cover many distinct and (sometimes) disparate streams. These encompass themes in system architecture, performance modeling, real-time, and fault tolerance, among several others. Many of these areas (fortunately or unfortunately?) have evolved separately and become established as discrete fields, although the overall design of computers and applications still, and inherently, involves their synergistic association. To open a topic of debate: This synergism is often lacking or, if present, not particularly acknowledged as coming from the pertinent area and, at times, reinvented. Among the offshoots of the broad discipline of computing systems, fault tolerance and real-time, in particular, are increasingly being viewed as areas with potential for synergy. The concept of fault tolerance is expanding in scope to address the broader area of dependability: reliability, performability, availability, safety, the consideration of temporal properties in defining “correct” services towards the goal of reliably delivering desired services at specified times. Similarly, the real-time community is increasingly incorporating notions of dependability within its definitions of provision of timely services. Realistically, the entire range of current system designs is reaching a complexity level that necessitates joint consideration of fault tolerance and real-time features, regardless of which gets primary emphasis. Examples abound: communication networks, transaction processing, and flight, industrial, medical, or automotive control all clearly require consideration of both dependability and timeliness. It is of interest to note that the two communities tackle largely overlapping sets of problems which can be broadly stated as: “assured delivery of correctly executed services” and constraints, i.e., “resource management, guarantees, and efficiency.” Interestingly, the fields of fault tolerance and real-time address these problems with different perspectives and aim at optimizing different sets of objectives within the same problems. For example, the fault tolerance emphasis involves reliability of the mechanisms providing for delivery of correctly executed services; the real-time community puts the emphasis on guaranteeing delivery of services in a timely manner. Similarly, resource management issues are addressed as redundancy management issues, versus issues of efficient use of uniprocessor/multiprocessors for task allocation from the realtime systems perspective. It is actually quite a long list of issues that get addressed differently by the fault-tolerant and real-time viewpoints. Moreover, the two fields rely on very similar fundamentals and techniques from engineering, mathematics, and theory. Indeed, at times this fosters a very healthy development of ideas and solutions from varied perspectives; often enough, it involves reinventing the wheel. Both the disciplines of fault tolerance and real-time have matured, with their individual seminal contributions, development of terminology, and conceptual basis. Perhaps this maturity now fosters a better chance in understanding complementary issues which have, over time, been developed with different perspectives, emphasis, and presentation. However, as they do tackle the common problem of “assured delivery of correctly executed services,” it is the intent of this special issue to address and encourage the integration of these two (perceived?) diverse areas toward a cohesive formulation of dependable real-time service and systems. In many ways, the notion and popularity of the term “Quality of Service (QoS)” already represents the juxtaposition of many fault-tolerant and real-time services. This special issue on dependable real-time systems endeavors to develop the area of systems with composite attributes of real-time and dependability within a common framework. The thrust is on understanding system models, identifying issues and concepts used across both communities for design and analysis, and elucidating the context and relevance of concepts synergistically across the areas of dependability and real time. The flavor of this special issue is both theoretical and experimental in scope to expand the underlying principles of the areas of both dependability and real-time. We paraphrase some topics from the call for papers which we believe depict some of the areas where there is significant work needed to link the fields of dependability and real-time. These include (in part): Neeraj Suri, Krithi Ramamritham |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1998 | A Framework for Dependability Driven Software IntegrationabstractThe integration of system and SW functions for efficiency, performance and especially dependability is of interest from a research and system design perspective. We propose a framework for directing the process of integration of SW functions, with the objective of designing and maintaining desired dependability attributes of the system over the integration process. Rules of composition for integrated functions, and measures to quantify the goodness of dependable system integration are also addressed. Neeraj Suri, Sunundu Ghosh, Thomas J. Marlowe |
ICDCS | 1 |
| 1997 | Cache based fault recovery for distributed systemsabstractNo cache based techniques for roll-forward fault recovery exist at present. A split-cache approach is proposed that provides efficient support for checkpointing and roll-forward fault recovery in distributed systems. This approach obviates the use of discrete stable storage or explicit synchronization among the processors. Stability of the checkpoint intervals is used as a driver for real time operations. Avi Mendelson, Neeraj Suri |
ICECCS | 2 |
| 1997 | Formally Verified On-Line DiagnosisabstractA reconfigurable fault tolerant system achieves the attributes of dependability of operations through fault detection, fault isolation and reconfiguration, typically referred to as the FDIR paradigm. Fault diagnosis is a key component of this approach, requiring an accurate determination of the health and state of the system. An imprecise state assessment can lead to catastrophic failure due to an optimistic diagnosis, or conversely, result in underutilization of resources because of a pessimistic diagnosis. Differing from classical testing and other off-line diagnostic approaches, we develop procedures for maximal utilization of the system state information to provide for continual, on-line diagnosis and reconfiguration capabilities as an integral part of the system operations. Our diagnosis approach, unlike existing techniques, does not require administered testing to gather syndrome information but is based on monitoring the system message traffic among redundant system functions. We present comprehensive on-line diagnosis algorithms capable of handling a continuum of faults of varying severity at the node and link level. Not only are the proposed algorithms on-line in nature, but are themselves tolerant to faults in the diagnostic process. Formal analysis is presented for all proposed algorithms. These proofs offer both insight into the algorithm operations and facilitate a rigorous formal verification of the developed algorithms. Chris J. Walter, Patrick Lincoln, Neeraj Suri |
IEEE Trans. Software Eng. | 3 |
| 1994 | Synchronization issues in real-time systemsabstractReal-time systems must accomplish executive and application tasks within specified timing constraints. In distributed real-time systems, the mechanisms that ensure fair access to shared resources, achieve consistent deadlines, meet timing or precedence constraints, and avoid deadlocks all utilize the notion of a common system-wide time base. A synchronization primitive is essential in meeting the demands of real-time critical computing. This paper provides a tutorial on the terminology, issues, and techniques essential to synchronization in real-time systems.> Neeraj Suri, Michelle M. Hugue, Chris J. Walter |
Proc. IEEE | 1 |