VLDB 2026 Research / reviewers in the wild / expert
Naghmeh Ramezani Ivaki
dblp:123/7759 · also Naghmeh Ivaki
· DBLP profile ↗
24ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-8376-6711ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 3 first-author · 10 since 2021Security and privacy · 11 · 2 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IoT security assessment: A systematic literature reviewabstractThe Internet of Things (IoT) has rapidly expanded across multiple sectors, exposing significant opportunities but also raising important concerns. This rapid growth has raised concerns about the security of IoT devices and the protection of the large volumes of data they collect, transmit, store, and process. Numerous large-scale attacks on IoT systems underscore the need for security measures, as well as comprehensive security assessments and benchmarking methods to verify and validate these systems. We conduct a Systematic Literature Review (SLR) to analyze previous studies, methodologies, and tools used to assess and benchmark the security of IoT systems, and to identify critical challenges and gaps in the existing literature. As a result, we highlight that, due to their complexity, IoT systems lack a comprehensive security framework that covers all layers and their security concerns. Despite awareness of known vulnerabilities, there is a lack of best practices, tools, and techniques to prevent, detect, and mitigate threats effectively. The absence of standardized security benchmarks complicates the evaluation and comparison of the solutions. There is also limited alignment with emerging standards such as ISO/IEC 27402 and SESIP. Finally, it is noteworthy that IoT gateway security remains unexplored despite its critical role in IoT ecosystems. CCS Concepts: • Computer systems organization → Embedded systems ; Redundancy ; Robotics; • Networks → Network reliability. Thaer Slaibi, Naghmeh Ramezani Ivaki, Marco Vieira |
J. Syst. Softw. | 2 |
| 2026 | Safety Assessment of UAV Operations in U-Space: A Comprehensive Study on Key Safety MetricsabstractUnmanned Aircraft Systems Traffic Management (UTM) and its European version, U-Space, are regulatory frameworks designed to ensure safe, efficient, and secure integration of Unmanned Aerial Vehicles (UAVs) into urban airspace by providing services such as monitoring, conflict resolution, and traffic management. To ensure the safety of UAVs' operations, a comprehensive safety assessment framework is crucial. To build such a framework, it is necessary to identify appropriate safety metrics and develop an approach to measure them, enabling the measurement and management of associated safety risks. In this work, we identify and analyze two categories of safety metrics: collision metrics and surveillance performance metrics. We present an approach grounded in U-space regulatory framework concepts to design and conduct a comprehensive experimental study investigating the impact of several factors that can affect UAV safety, including GPS and IMU failures of varying duration at different UAV speeds, update intervals, traffic densities, and weather conditions, through quantitative assessment of the identified safety metrics. The results reveal key insights into: 1) Identifying the metrics most affected by variations in factors in the presence of GPS or IMU failures, 2) Determination of metrics most correlated to safety risk level under varying conditions, 3) Establishment of risk thresholds for selected metrics under erroneous or varying conditions, contributing to the identification of reliable risk indicators, and 4) Evaluation of the performance and limitations of preventive mechanisms, such as the fail-safe system, under erroneous behavior of GPS and IMU and across different operational and environmental conditions. Omid Asghari, Naghmeh Ramezani Ivaki, Henrique Madeira |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | bBench: A Comprehensive Performance Benchmark for Blockchain ApplicationsabstractThe performance assessment of blockchain applications holds significant challenges due to their decentralized architecture, immutable smart contracts, distributed ledgers, and operational costs such as gas fees. Existing blockchain benchmarks often either fail to fully capture blockchain-specific behaviors or offer limited configurability and metric reporting. In this paper, we present a new and comprehensive benchmark designed explicitly for blockchain applications, named bBench. Building on established principles from traditional benchmarking and by specializing them in the blockchain context and supported by customized blockchain tools (i.e., Hyperledger Caliper, web3.eth, and node-os-utils), bBench characterizes blockchain application performance in four dimensions: network performance, resource utilization, storage usage, and operational cost. We demonstrate the effectiveness of our benchmark through a case study involving 12 smart contract applications with varying performance demands, some of which hold known vulnerabilities. The results show the benchmark’s ability to quantify performance deviations across different applications, as well as those caused by the activation of specific vulnerabilities. Fernando Richter Vidal, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
ISSRE | 2 |
| 2025 | Analyzing the impact of elusive faults on blockchain reliabilityabstractBlockchain has recently become very popular due to its use in cryptocurrencies and potential application in various domains (e.g., retail, healthcare, and insurance). The smart contract is a key part of blockchain systems and specifies an agreement between transaction participants. Nowadays, smart contracts are being deployed to carry residual faults, including severe vulnerabilities that lead to different types of failures at runtime. Fault detection tools can be used to detect faults that may then be removed from the code before deployment. However, in the case of smart contracts, the common opinion is that tools are immature and ineffective. In this work, we carry out a fault injection campaign to empirically analyze the runtime impact that realistic faults present in smart contracts may have on the reliability of blockchain systems. We pay particular attention to the faults that elude popular smart contract verification tools and show if and in which ways the faults lead the blockchain system to fail at runtime. We map the observations to the fault detection capabilities of three state-of-the-art fault detection tools, namely Mythril, Slither, and Securify. The results show that the tools individually have poor detection capabilities (e.g., Securify with 6.4% accuracy and Mythril with 60% accuracy) or tend to generate false alerts (i.e., only 1.74% of Slither's alerts are correct). The results also show several elusive faults responsible for severe blockchain failures, such as A_MCV, which impacts the integrity of the ledger, and I_MVMSV, which causes gas depletion, just to name a few. Fernando Richter Vidal, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
Blockchain Res. Appl. | 2 |
| 2025 | A systematic review on smart contracts security design patternsabstractAbstract Smart contracts have accelerated the adoption of blockchain technology across various domains by enabling coded agreements between transaction participants. However, increased software defects and vulnerabilities in smart contracts, driven by developer inexperience with languages like Solidity and a lack of effective detection tools, pose significant risks. Given the high value of assets managed on blockchain (e.g., cryptocurrencies), these vulnerabilities can lead to severe consequences. Researchers and practitioners have proposed numerous smart contract design patterns to mitigate certain faults or vulnerabilities. Despite these efforts, it remains unclear which types of defects these patterns target and how effectively they address the wide range of existing smart contract security vulnerabilities. In this paper, we review the state of the art in smart contract design patterns, categorizing them and analyzing their effectiveness in mitigating known security vulnerabilities. Our findings reveal that only five patterns directly aim to prevent security vulnerabilities, collectively addressing just 6 out of 94 security issues identified by OpenSCV (a state-of-the-art vulnerability taxonomy), highlighting the need for further research on smart contract security design patterns. Sadaf Azimi, Ali Golzari, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
Empir. Softw. Eng. | 3 |
| 2024 | A Comprehensive Study on Drones Resilience in the Presence of Inertial Measurement Unit FaultsabstractUnmanned aerial vehicles (UAVs) have gained immense popularity for their versatility and diverse applications. However, this increased usage has raised concerns about the safety and security of UAVs, emphasizing the critical role of their Inertial Measurement Units (IMUs) in ensuring accurate orientation and position data. IMU faults, including both Accelerometer faults and Gyrometer faults, can lead to severe consequences, such as mission failures, collisions, or loss of control. This study addresses the need to enhance UAV resilience in urban airspace by exploring the impact of various IMU faults. A comprehensive fault model is introduced in this paper, covering a range of faults from hardware malfunctions to external attacks. Through extensive fault injection experiments in a simulated environment, the study assesses the effects of different fault types and durations on mission outcomes, providing valuable insights for developing resilient and fault-tolerant UAV systems. Evaluation metrics, including inner and outer bubble violations, missions completed, flight duration, and distance traveled, offer a comprehensive understanding of IMU fault impacts in dynamic operational scenarios. Results reveal that longer injection durations, particularly at 30 seconds, increase bubble violations and significantly reduce mission completion rates. Accelerometer faults, such as “Accelerometer Freeze” and “Accelerometer Random” exhibit reduced mission completion rates of 42.5% and 5%, respectively. Gyrometer faults, especially “Gyrometer Minimum” and “Gyrometer Random” lead to the lowest mission completion rates (2.5%). Additionally, IMU faults (where the fault affects both the Accelerometer and Gyrometer), notably “IMU Minimum”, “IMU Freeze”, and “IMU Random” result in complete mission failures, highlighting the importance of understanding specific fault characteristics. These insights can contribute to developing fault tolerance mechanisms and resilient UAV systems in complex and dynamic environments. Anamta Khan, Naghmeh Ramezani Ivaki, Henrique Madeira |
DSN | 2 |
| 2024 | OpenSCV: an open hierarchical taxonomy for smart contract vulnerabilitiesabstractAbstract Smart contracts are nowadays at the core of most blockchain systems. Like all computer programs, smart contracts are subject to the presence of residual faults, including severe security vulnerabilities. However, the key distinction lies in how these vulnerabilities are addressed. In smart contracts, when a vulnerability is identified, the affected contract must be terminated within the blockchain, as due to the immutable nature of blockchains, it is impossible to patch a contract once deployed. In this context, research efforts have been focused on proactively preventing the deployment of smart contracts containing vulnerabilities, mainly through the development of vulnerability detection tools. Along with these efforts, several heterogeneous vulnerability classification schemes appeared (e.g., most notably DASP and SWC). At the time of writing, these are mostly outdated initiatives, even though new smart contract vulnerabilities are consistently uncovered. In this paper, we propose OpenSCV, a new and Open hierarchical taxonomy for Smart Contract vulnerabilities, which is open to community contributions and matches the current state of the practice while being prepared to handle future modifications and evolution. The taxonomy was built based on the analysis of the existing research on vulnerability classification, community-maintained classification schemes, and research on smart contract vulnerability detection. We show how OpenSCV covers the announced detection ability of the current vulnerability detection tools and highlight its usefulness in smart contract vulnerability research. To validate OpenSCV, we performed an expert-based analysis wherein we invited multiple experts engaged in smart contract security research to participate in a questionnaire. The feedback from these experts indicated that the categories in OpenSCV are representative, clear, easily understandable, comprehensive, and highly useful. Regarding the vulnerabilities, the experts confirmed that they are easily understandable. Fernando Richter Vidal, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
Empir. Softw. Eng. | 2 |
| 2024 | Vulnerability detection techniques for smart contracts: A systematic literature review
Fernando Richter Vidal, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
J. Syst. Softw. | 2 |
| 2023 | Lead Time Analysis for UAVs' Failure Prediction in U-spaceabstractIn recent years, UAVs have been increasingly used in urban environments due to agility in movement, simplicity in mechanics, low price, and ability to access locations that are difficult or impossible to reach by humans. A significant number of drones are expected to fly in the urban sky shortly. The profitable nature of commercial UAVs/drone applications in urban space will imply a high density of drones; therefore, avoiding mid-air collisions will be critical for the safe operation of the UAVs. In Europe, U-space services are being created to guarantee the safe operations of UAVs in urban Very Low Level (VLL) airspace. To avoid collisions, U-space considers a separation minima (i.e., the minimum safe distance between UAVs) surrounding each UAV. Thus, violating the separation minima, which might be caused by abnormal conditions (e.g., bad weather conditions), failure conditions (e.g., GPS failure in UAVs), or unreliable behavior of the system (e.g., inaccurate GPS positioning data or erratic position estimation by flight controller), could potentially result in conflicts that require immediate mitigation measures to avoid mid-air collisions. Failure prediction is a promising method for preventing separation minima violations in U-space services. However, in order to have effective failure prediction, the lead time, which is the time between the activation of a fault and its manifestation in a system as a failure, must account for both the prediction step and the subsequent mitigation actions. This paper aims to evaluate the lead time in UAV systems in the presence of positioning-related issues (as being critical for the safe operation of UAVs) from a U-space perspective. We used fault injection to inject 18 different types of faults (or emulating failures) in 28 different UAV missions. The results show that the lead time for 17 types of faults injected is at least 14 seconds (in some cases, no failure occurred). Thus, U-space has at least 14 seconds to predict and mitigate such faults. In the case of GPS failure (i.e., GPS signal is entirely missing), lead time is about 5 seconds, requiring faster strategies for failure prediction and mitigation plans. Omid Asghari, Naghmeh Ramezani Ivaki, Henrique Madeira |
PRDC | 2 |
| 2023 | A Machine Learning driven Fault Tolerance Mechanism for UAVs' Flight ControllerabstractUnmanned Aerial Vehicles (UAVs) are susceptible to various hazards (e.g., software or hardware failures, communication failures, or security attacks) that may hinder mission completion or compromise safety by violating the separation minima (i.e., the minimum distance that must be maintained between UAVs in order to ensure safe and efficient operations). To address this issue, this paper proposes a new machine learning-based fault-tolerant mechanism for UAV flight controllers that tolerates GPS-related faults. These faults are of paramount importance (i.e., accidental faults and/or security attacks that eventually cause failures in the GPS function/data), as accurate positioning and tracking are essential to assure safe operation in UAVs. The proposed machine learning models were built using 884,410 data records from 1,985 flight logs collected from the PX4 public repository. The trained models are used to predict the expected position of the UAV during a mission, and separation minima are used as a threshold to detect the GPS hazards by comparing it with the distance between two consecutive position values. When a hazard is detected (i.e., the distance is higher than separation minima), the predicted values by machine learning models are fed into the flight controller’s position estimator (i.e., an Extended Kalman Filter (EKF)). To evaluate the effectiveness of this approach, validation experiments were conducted on several realistically defined missions while being exposed to different types of failure conditions (e.g., GPS signal loss or GPS Spoofing), both with and without using the proposed fault-tolerant mechanism. The results show a remarkable reduction in safety violations (the number of separation minima violations was reduced from 94 to 1). Additionally, the proposed mechanism demonstrated a notable improvement in the distance traveled by UAV and the duration of the flight mission in failure conditions, showing its ability to mitigate faults effectively. These findings support the effectiveness of the proposed fault tolerance mechanism in enhancing UAV safety in the presence of issues caused by GPS. Anamta Khan, João R. Campos, Naghmeh Ramezani Ivaki, Henrique Madeira |
PRDC | 3 |
| 2023 | Trustworthiness models to categorize and prioritize code for security improvement
Nadia Patricia Da Silva Medeiros, Naghmeh Ramezani Ivaki, Pedro Costa 0002, Marco Vieira |
J. Syst. Softw. | 2 |
| 2022 | Are UAVs' Flight Controller Software Reliable?abstractUnmanned Ariel Vehicles (UAVs) are recently being studied and worked upon to make them safe and secure for the upcoming expected growth of UAVs in civilian airspace. These efforts resulted in services such as Unmanned Aircraft System Traffic Management (UTM) or U-space in Europe, providing services to regularize and organize (pre-flight), monitor/track (during the flight) drones in civilian airspace while avoiding collisions. The primary source of information for tracking drones during flight is GPS positioning data, which is used and filtered (after being fused with the other sensors' data) by the flight controller software to estimate the vehicle position, velocity, and orientation. Extended Kalman Filter (EKF), which is used in most open-source flight controllers such as PX4, is responsible for doing this estimation. This makes EKF a critical component of the whole system. This paper aims to study the reliability of flight controllers and their core component, namely EKF, in the presence of GPS-related failures. To do so, we injected faults (i.e., we emulated failures indeed) on GPS raw data ranging from small noises to complete failure (missing GPS signals) and GPS spoofing to study their impact on EKF estimation and on the system as a whole. We observed that for small faults (e.g., Fixed Small Noise or Freeze Values), EKF is efficient and can tolerate/compensate the faults, whereas there is a gap in the filter for handling bigger anomalies (e.g., Invalid Values or Random Values) in the GPS data. Our research also clearly demonstrates that GPS faults lasting 30 seconds or more have a noticeable effect, which represents a clear vulnerability since GPS can be subject of cyber attacks such as spoofing. The quantification of the impact of GPS-related failures in the PX4 is an essential step to measure and improve the reliability of UAVs' flight controller software. Anamta Khan, Naghmeh Ramezani Ivaki, Henrique Madeira |
PRDC | 2 |
| 2022 | Assessment of the Impact of U-space Faulty Conditions on Drones Conflict Rate
Anamta Khan, Carlos A. Chuquitarco Jiménez, Morcillo-Pallarés Pablo, Naghmeh Ramezani Ivaki, Juan Vicente Balbastre-Tejedor, Henrique Madeira |
SAFECOMP | 4 |
| 2021 | An Empirical Evaluation of the Effectiveness of Smart Contract Verification ToolsabstractBlockchain has become popular due to its use in cryptocurrencies and potential to support different business-critical services (e.g., financial services, retail). The smart contract is at the center of blockchain systems and is a coded specification of an agreement between interacting partners in a transaction. Like other software artifacts, smart contracts are prone to carry residual faults. As many contracts are being used to handle financial transactions, huge losses may occur if a vulnerability is exploited. Also, a faulty contract cannot be corrected once it has been deployed on the blockchain, it can only be terminated and a new one must be deployed, which aggravates the cost of deploying contracts with faults and marks the reputation of the provider. Smart contract verification tools have been emerging, but limited knowledge is available regarding their real effectiveness. In this paper, we define a smart contract defect classification scheme based on the Orthogonal Defect Classification and apply it to a contract dataset, which has been extracted from multiple sources and holds different types of defects. We use the dataset to evaluate three state of the art verification tools regarding their fault detection performance. Results show the relatively low effectiveness of the tools and their complementarity. Bruno Dia, Naghmeh Ramezani Ivaki, Nuno Laranjeiro |
PRDC | 2 |
| 2018 | An Approach for Trustworthiness Benchmarking Using Software MetricsabstractTrustworthiness is a paramount concern for users and customers in the selection of a software solution, specially in the context of complex and dynamic environments, such as Cloud and IoT. However, assessing and benchmarking trustworthiness (worthiness of software for being trusted) is a challenging task, mainly due to the variety of application scenarios (e.g., businesscritical, safety-critical), the large number of determinative quality attributes (e.g., security, performance), and last, but foremost, due to the subjective notion of trust and trustworthiness. In this paper, we present trustworthiness as a measurable notion in relative terms based on security attributes and propose an approach for the assessment and benchmarking of software. The main goal is to build a trustworthiness assessment model based on software metrics (e.g., Cyclomatic Complexity, CountLine, CBO) that can be used as indicators of software security. To demonstrate the proposed approach, we assessed and ranked several files and functions of the Mozilla Firefox project based on their trustworthiness score and conducted a survey among several software security experts in order to validate the obtained rank. Results show that our approach is able to provide a sound ranking of the benchmarked software. Nadia Patricia Da Silva Medeiros, Naghmeh Ramezani Ivaki, Pedro Costa 0002, Marco Vieira |
PRDC | 2 |
| 2018 | Effects of GPS Spoofing on Unmanned Aerial VehiclesabstractUnmanned Aerial Vehicles (UAVs) are no longer exclusively military and scientific solutions. These vehicles have been growing in popularity among hobbyist and also as industrial solutions for specific activities. The flying characteristics and the absence of a crew on board of these devices allow them to perform a wide variety of activities, which can be inaccessible to humans or may threat their life. Despite the advantages, they also bring up major concerns regarding security breaches in the flight controller software, which may lead to security (e.g., vehicle hijacking by attackers), safety (e.g., crashing the vehicle into a planned area or building), or privacy (e.g., eavesdropping or stealing video footage) problems. GPS spoofing is one the main threat of UAVs. The predictability and knowledge of GPS signal properties, create conditions to attackers to assume control of the UAV and use it for their own objectives. In this paper the GPS spoofing effect on UAV is analyzed through a series of tests, under a simulation environment. The results are shown as deviation from the original trajectory and attack success, and analyzed over time and by attack type. Daniel Mendes, Naghmeh Ramezani Ivaki, Henrique Madeira |
PRDC | 2 |
| 2018 | A survey on reliable distributed communication
Naghmeh Ramezani Ivaki, Nuno Laranjeiro, Filipe Araújo |
J. Syst. Softw. | 1 |
| 2017 | Software Metrics as Indicators of Security VulnerabilitiesabstractDetecting software security vulnerabilities and distinguishing vulnerable from non-vulnerable code is anything but simple. Most of the time, vulnerabilities remain undisclosed until they are exposed, for instance, by an attack during the software operational phase. Software metrics are widely-used indicators of software quality, but the question is whether they can be used to distinguish vulnerable software units from the non-vulnerable ones during development. In this paper, we perform an exploratory study on software metrics, their interdependency, and their relation with security vulnerabilities. We aim at understanding: i) the correlation between software architectural characteristics, represented in the form of software metrics, and the number of vulnerabilities; and ii) which are the most informative and discriminative metrics that allow identifying vulnerable units of code. To achieve these goals, we use, respectively, correlation coefficients and heuristic search techniques. Our analysis is carried out on a dataset that includes software metrics and reported security vulnerabilities, exposed by security attacks, for all functions, classes, and files of five widely used projects. Results show: i) a strong correlation between several project-level metrics and the number of vulnerabilities, ii) the possibility of using a group of metrics, at both file and function levels, to distinguish vulnerable and non-vulnerable code with a high level of accuracy. Nadia Patricia Da Silva Medeiros, Naghmeh Ramezani Ivaki, Pedro Costa 0002, Marco Vieira |
ISSRE | 2 |
| 2016 | Towards designing reliable messaging patternsabstractReliable communication is nowadays pervasively supported by TCP, which is poorly adapted for message-based communications, because it offers a streaming channel with no mechanisms to encapsulate messages. Moreover, TCP does not tolerate connection crashes. Thus, whenever reliable message-based communication is needed, developers either use heavy-weight middleware, like Java Message Service (JMS), or develop their own custom error-prone solutions for recovering from crashes. In this paper, we introduce two TCP-based design patterns that address these limitations, and facilitate the development of light-weight and reliable message-based applications. Our design solutions are modular, in the sense that they build on top of each other. Naghmeh Ramezani Ivaki, Nuno Laranjeiro, Filipe Araújo |
NCA | 1 |
| 2016 | The 2016 IEEE Services Emerging Technology Track on Dependable and Secure Services (DSS 2016)abstractThis emerging technology track focuses on key topics regarding dependability and security of software and services. Service-based systems are being used in business, safety, and mission-critical environments to achieve operational goals and possess special characteristics that bring in difficult challenges to the research and industry communities. Among these challenges, dependability and security have been widely identified as critical aspects that need to be addressed, especially when considering that many services are also nowadays being deployed on the web, used over unreliable networks, and potentially exposed to security threats. The goal of the Emerging Technology Track On Dependable and Secure Services is to bring together researchers and practitioners to present original research and industrial practice regarding techniques to improve the dependability and security of services. Services hold special characteristics, in particular their typically complex nature, high heterogeneity, and fast-changing dynamics. In such scenarios, infrastructure interdependencies, failure and recovery modeling and analysis, accidental threats and attack modeling and evaluation, testing approaches, testbeds, benchmarks, interoperability in presence of dependability and security guarantees, as well as techniques and tools to assess the impact of accidental and malicious threats, metrics for assessing dependability and security are among the crucial aspects to be addressed. Nuno Laranjeiro, Naghmeh Ramezani Ivaki, Marco Vieira |
SERVICES | 2 |
| 2014 | Fault-Tolerant bi-directional communications in web-based applicationsabstractThe Hypertext Transfer Protocol (HTTP) and the Transmission Control Protocol (TCP) are the most popular protocols used in the development of web-based applications. Despite their popularity, the use of these protocols brings two limitations to applications and systems that require reliable interactive real-time communications: 1) HTTP forces applications to work in a request-response paradigm, even if a reply is not necessary, not allowing the server to send anything to a client without the client explicitly requesting it; 2) TCP provides no recovery options for network outages, thus forcing developers to write their own error-prone, complex, and ad hoc solutions. In this paper we introduce a solution that offers both bi-directional and reliable communication to web-based applications, even in presence of connection failures. To make this possible, we combine the idea behind WebSockets and a Session-Based Fault-Tolerant design pattern. Naghmeh Ramezani Ivaki, Filipe Araújo |
ICPADS | 1 |
| 2014 | Session-based fault-tolerant design patternsabstractDespite offering reliability against dropped and reordered packets, the widely adopted Transmission Control Protocol (TCP) provides nearly no recovery options for longterm network outages. When the network fails, developers must rollback the application to some coherent state on their own, using error-prone solutions. Overcoming this limitation is, therefore, a deeply investigated and challenging problem. Existing solutions range from transport-layer to application-layer protocols, including additions to TCP, usually transparent to the application. None of these solutions is perfect, because they all impact TCP's simplicity, performance or ubiquity, if not all. To avoid these shortcomings, we contain TCP connection crashes inside a single session layer exposed as a sockets interface. Based on this interface, we create a blocking and a non-blocking fault-tolerant design pattern. We explore the blocking design in an open source File Transfer Protocol (FTP) server and perform a thorough evaluation of performance, complexity and overhead of both designs. Our results show that using one of the patterns to tolerate TCP connection crashes, in new or existing applications, involves a very limited effort and negligible penalties. Naghmeh Ramezani Ivaki, Filipe Araújo, Fernando J. Barros |
ICPADS | 1 |
| 2014 | Design of Multi-threaded Fault-Tolerant Connection-Oriented CommunicationabstractFault-tolerance is vital for dependable distributed applications that can deliver service, even in the presence of faults. Over the last few decades, above all protocols proposed to offer reliability and fault-tolerance, TCP grew to become one of the cornerstones of the Internet. However, despite emulating reliable communication in distributed environments, TCP does not handle connection failures when the connectivity is lost for some time, even if both endpoints are still running. When this occurs, developers must rollback the peers to some coherent state, many times with error-prone, ad hoc, or custom application-level solutions. In this paper, we refine the Acceptor-Connector design pattern to tackle the TCP unreliability problem. The pattern decouples the failure-related processing from the connection and service processing, efficiently handling different connections and their possible crashes concurrently, thereby yielding more reusable, extensible, and efficient distributed communication. The solution we propose incorporates proven multi-threaded solutions and a buffering scheme that discards the need for an application-layer acknowledgment scheme. This simplifies the development of reliable connection-oriented applications using the ubiquitous TCP protocol. Naghmeh Ramezani Ivaki, Filipe Araújo, Fernando J. Barros |
PRDC | 1 |
| 2012 | A Middleware for Exactly-Once Semantics in Request-Response InteractionsabstractAlthough the need for the exactly-once request-response interaction pattern is ubiquitous in distributed systems, making it work in practice is anything but simple. Ensuring the at-most-once part of the invocation is relatively easy. Unfortunately, the same is not true for the at-least-once guarantee, which depends on the recovery from crashes of the client, the server and the network. This is what makes the exactly-once interaction so difficult in practice: client and server must log their actions into stable storage, and they must be able to restart the network connections. In this paper, we present a middleware that implements the exactly-once request-response pattern, in presence of network and endpoints crashes. The main contribution of our work is to release the programmer from the complex tasks of recovering from message losses and network crashes. Naghmeh Ramezani Ivaki, Filipe Araújo, Raul Barbosa |
PRDC | 1 |