EDBT 2026 Demo / reviewers in the wild / expert
Yair Amir
dblp:15/5788
· DBLP profile ↗
45ranked-venue papers
30as first author
2since 2021 · last 2024
0009-0002-4358-7270ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 14 first-authorSecurity and privacy · 19 · 12 first-author · 2 since 2021Computer networks · 5 · 5 first-authorTheory of computation · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Tolerating Compound Threats in Critical Infrastructure Control SystemsabstractCompound threats, in which cyberattacks are targeted in the aftermath of a natural hazard, pose an important emerging threat for critical infrastructure. In this paper, we analyze the system design implications of compound threats for power grid SCADA systems for the first time. We introduce a novel compound threat model and develop a tool for analyzing resilience under this threat model. Using our tool, we compare the resilience of existing fault- and intrusion-tolerant SCADA system architectures in case studies based on two power utilities: Hawaiian Electric (HECO) and Florida Power & Light (FPL). We show that no existing system architecture adequately addresses compound threats, but that it is possible to improve resilience to such threats by explicitly considering natural hazards in the system design and by employing a new out-of-band reconfiguration mechanism for intrusion-tolerant systems. However, an important outcome of our work is that compound threats remain a challenging problem, with no complete solution. Sahiti Bommareddy, Maher Khan, Huzaifah Nadeem, Benjamin Gilby, Imes Chiu, John W. van de Lindt, Omar Nofal, Mathaios Panteli, Linton Wells, Yair Amir, Amy Babay |
SRDS | 10 |
| 2022 | Real-Time Byzantine Resilience for Power Grid SubstationsabstractIn the world of increasing cyber threats, a compromised protective relay can put power grid resilience at risk by irreparably damaging costly power assets or by causing significant disruptions. We present the first architecture and protocols for the substation that ensure correct protective relay operation in the face of successful relay intrusions and network attacks while meeting the required latency constraint of a quarter power cycle (4.167ms). Our architecture supports other rigid requirements, including continuous availability over a long system lifetime and seamless substation integration. We evaluate our implementation in a range of fault-free and faulty operation conditions, and provide deployment tradeoffs. Sahiti Bommareddy, Daniel Qian, Christopher A. Bonebrake, Paul M. Skare, Yair Amir |
SRDS | 5 |
| 2019 | Deploying Intrusion-Tolerant SCADA for the Power GridabstractWhile there has been considerable research on making power grid Supervisory Control and Data Acquisition (SCADA) systems resilient to attacks, the problem of transitioning these technologies into deployed SCADA systems remains largely unaddressed. We describe our experience and lessons learned in deploying an intrusion-tolerant SCADA system in two realistic environments: a red team experiment in 2017 and a power plant test deployment in 2018. These experiences resulted in technical lessons related to developing an intrusion-tolerant system with a real deployable application, preparing a system for deployment in a hostile environment, and supporting protocol assumptions in that hostile environment. We also discuss some meta-lessons regarding the cultural aspects of transitioning academic research into practice in the power industry. Amy Babay, John L. Schultz, Thomas Tantillo, Samuel Beckley, Eamon Jordan, Kevin Ruddell, Kevin Jordan, Yair Amir |
DSN | 8 |
| 2018 | Network-Attack-Resilient Intrusion-Tolerant SCADA for the Power GridabstractAs key components of the power grid infrastructure, Supervisory Control and Data Acquisition (SCADA) systems are likely to be targeted by nation-state-level attackers willing to invest considerable resources to disrupt the power grid. We present Spire, the first intrusion-tolerant SCADA system that is resilient to both system-level compromises and sophisticated network-level attacks and compromises. We develop a novel architecture that distributes the SCADA system management across three or more active sites to ensure continuous availability in the presence of simultaneous intrusions and network attacks. A wide-area deployment of Spire, using two control centers and two data centers spanning 250 miles, delivered nearly 99.999% of all SCADA updates initiated over a 30-hour period within 100ms. This demonstrates that Spire can meet the latency requirements of SCADA for the power grid. Amy Babay, Thomas Tantillo, Trevor Aron, Marco Platania, Yair Amir |
DSN | 5 |
| 2018 | Toward an Intrusion-Tolerant Power Grid: Challenges and OpportunitiesabstractWhile cyberattacks pose a relatively new challenge for power grid control systems, commercial cloud systems have needed to address similar threats for many years. However, technology and approaches developed for cloud systems do not necessarily transfer directly to the power grid, due to important differences between the two domains. We discuss our experience adapting intrusion-tolerant cloud technologies to the power domain and describe the challenges we have encountered and potential directions for overcoming those obstacles. Amy Babay, John L. Schultz, Thomas Tantillo, Yair Amir |
ICDCS | 4 |
| 2017 | Structured Overlay Networks for a New Generation of Internet ServicesabstractThe dramatic success and scaling of the Internet was made possible by the core principle of keeping it simple in the middle and smart at the edge (or the end-to-end principle). However, new applications bring new demands, and for many emerging applications, the Internet paradigm presents limitations. For applications in this new generation of Internet services, structured overlay networks offer a powerful framework for deploying specialized protocols that can provide new capabilities beyond what the Internet natively supports by leveraging global state and in-network processing. The structured overlay concept includes three principles: A resilient network architecture, a flexible overlay node software architecture that exploits global state and unlimited programmability, and flow-based processing. We demonstrate the effectiveness of structured overlay networks in supporting today's demanding applications and propose forward-looking ideas for leveraging the framework to develop protocols that push the boundaries of what is possible in terms of performance and resilience. Amy Babay, Claudiu Danilov 0001, John Lane, Michal Miskin-Amir, Daniel Obenshain, John L. Schultz, Jonathan Robert Stanton, Thomas Tantillo, Yair Amir |
ICDCS | 9 |
| 2017 | Timely, Reliable, and Cost-Effective Internet Transport Service Using Dissemination GraphsabstractEmerging applications such as remote manipulation and remote robotic surgery require communication that is both timely and reliable, but the Internet natively supports only communication that is either completely reliable with no timeliness guarantees (e.g. TCP) or timely with best-effort reliability (e.g. UDP). We present an overlay transport service that can provide highly reliable communication while meeting stringent timeliness guarantees (e.g. 130ms round-trip latency across the US) over the Internet. To enable routing schemes that can support the necessary timeliness and reliability, we introduce dissemination graphs, providing a unified framework for specifying routing schemes ranging from a single path, to multiple disjoint paths, to arbitrary graphs. We conduct an extensive analysis of real-world network data, finding that a routing approach using two disjoint paths performs well in most cases, and that cases where two disjoint paths do not perform well typically involve problems around a source or destination. Based on this analysis, we develop a timely dissemination-graph-based routing method that can add targeted redundancy in problematic areas of the network. This approach can cover over 99% of the performance gap between a traditional single-path approach and an optimal (but prohibitively expensive) scheme, while two dynamic disjoint paths cover about 70% of this gap, and two static disjoint paths cover about 45%. This performance improvement is obtained at a cost increase of about 2% over two disjoint paths. Amy Babay, Emily Wagner, Michael Dinitz, Yair Amir |
ICDCS | 4 |
| 2016 | Fast Total Ordering for Modern Data CentersabstractThe performance profile of local area networks has changed over the last decade, but many practical group communication and ordered messaging tools rely on core ideas invented over a decade ago. We present the Accelerated Ring protocol, a novel ordering protocol that improves on the performance of standard token-based protocols by allowing processes to pass the token before they have finished multicasting. This performance improvement is obtained while maintaining the correctness and other beneficial properties of token-based protocols. On 1-gigabit networks, a single-threaded daemon-based implementation of the protocol reaches network saturation, and can reduce latency by 45% compared to a standard token-based protocol while simultaneously increasing throughput by 30%. On 10-gigabit networks, the implementation reaches throughputs of 6 Gbps, and can reduce latency by 30-35% while simultaneously increasing throughput by 25-40%. A production implementation of the Accelerated Ring protocol has been adopted as the default ordering protocol for data center environments in Spread, a widely-used open-source group communication system. Amy Babay, Yair Amir |
ICDCS | 2 |
| 2016 | Practical Intrusion-Tolerant NetworksabstractAs the Internet becomes an important part of the infrastructure our society depends on, it is crucial to construct networks that are able to work even when part of the network is compromised. This paper presents the first practical intrusion-tolerant network service, targeting high-value applications such as monitoring and control of global clouds and management of critical infrastructure for the power grid. We use an overlay approach to leverage the existing IP infrastructure while providing the required resiliency and timeliness. Our solution overcomes malicious attacks and compromises in both the underlying network infrastructure and in the overlay itself. We deploy and evaluate the intrusion-tolerant overlay implementation on a global cloud spanning East Asia, North America, and Europe, and make it publicly available. Daniel Obenshain, Thomas Tantillo, Amy Babay, John L. Schultz, Andrew Newell, Md. Endadul Hoque, Yair Amir, Cristina Nita-Rotaru |
ICDCS | 7 |
| 2015 | Fast Total Ordering for Modern Data CentersabstractData center applications rely on messaging services that guarantee reliable, ordered message delivery for a wide range of distributed coordination tasks. Totally ordered multicast, which (informally) guarantees that all processes receive messages in exactly the same order, is particularly useful for maintaining consistent distributed state in systems as diverse as financial systems, distributed storage systems, cloud management, and big data analytics platforms. Amy Babay, Yair Amir |
ICDCS | 2 |
| 2015 | Increasing Network Resiliency by Optimally Assigning Diverse Variants to Routing NodesabstractNetworks with homogeneous routing nodes are constantly at risk as any vulnerability found against a node could be used to compromise all nodes. Introducing diversity among nodes can be used to address this problem. With few variants, the choice of assignment of variants to nodes is critical to the overall network resiliency. We present the Diversity Assignment Problem (DAP), the assignment of variants to nodes in a network, and we show how to compute the optimal solution in medium-size networks. We also present a greedy approximation to DAP that scales well to large networks. Our solution shows that a high level of overall network resiliency can be obtained even from variants that are weak on their own. We provide a variation of our problem that matches the specific communication requirements of applications run over the network (e.g., Paxos and BFT). Also, we analyze the loss in resiliency when optimally assigning variants based on inaccurate information about compromises. Andrew Newell, Daniel Obenshain, Thomas Tantillo, Cristina Nita-Rotaru, Yair Amir |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2014 | Towards a Practical Survivable Intrusion Tolerant Replication SystemabstractThe increasing number of cyber attacks against critical infrastructures, which typically require large state and long system lifetimes, necessitates the design of systems that are able to work correctly even if part of them is compromised. We present the first practical survivable intrusion tolerant replication system, which defends across space and time using compiler-based diversity and proactive recovery, respectively. Our system supports large-state applications, and utilizes the Prime BFT protocol (providing performance guarantees under attack) with a compiler-based diversification engine. We devise a novel theoretical model that computes how resilient the system is over its lifetime based on the rejuvenation rate and the number of replicas. This model shows that we can achieve a confidence in the system of 95% over 30 years even when we transfer a state of 1 terabyte after each rejuvenation. Marco Platania, Daniel Obenshain, Thomas Tantillo, Ricky Sharma, Yair Amir |
SRDS | 5 |
| 2014 | Authenticated Adversarial Routing
Yair Amir, Paul Bunn, Rafail Ostrovsky |
J. Cryptol. | 1 |
| 2013 | Increasing network resiliency by optimally assigning diverse variants to routing nodesabstractNetworks with homogeneous routing nodes are constantly at risk as any vulnerability found against a node could be used to compromise all nodes. Introducing diversity among nodes can be used to address this problem. With few variants, the choice of assignment of variants to nodes is critical to the overall network resiliency. We present the Diversity Assignment Problem (DAP), the assignment of variants to nodes in a network, and we show how to compute the optimal solution in medium-size networks. We also present a greedy approximation to DAP that scales well to large networks. Our solution shows that a high level of overall network resiliency can be obtained even from variants that are weak on their own. For real-world systems that grow incrementally over time, we provide an online version of our solution. Lastly, we provide a variation of our solution that is tunable for specific applications (e.g., BFT). Andrew Newell, Daniel Obenshain, Thomas Tantillo, Cristina Nita-Rotaru, Yair Amir |
DSN | 5 |
| 2011 | Prime: Byzantine Replication under AttackabstractExisting Byzantine-resilient replication protocols satisfy two standard correctness criteria, safety and liveness, even in the presence of Byzantine faults. The runtime performance of these protocols is most commonly assessed in the absence of processor faults and is usually good in that case. However, faulty processors can significantly degrade the performance of some protocols, limiting their practical utility in adversarial environments. This paper demonstrates the extent of performance degradation possible in some existing protocols that do satisfy liveness and that do perform well absent Byzantine faults. We propose a new performance-oriented correctness criterion that requires a consistent level of performance, even with Byzantine faults. We present a new Byzantine fault-tolerant replication protocol that meets the new correctness criterion and evaluate its performance in fault-free executions and when under attack. Yair Amir, Brian A. Coan, Jonathan Kirsch, John Lane |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2010 | A Robust Push-To-Talk Service for Wireless Mesh NetworksabstractPush-to-Talk (PTT) is a useful capability for rapidly deployable wireless mesh networks used by first responders. PTT allows several users to speak with each other while using a single, half-duplex, communication channel, such that only one user speaks at a time while all other users listen. Furthermore, enabling regular PSTN phone users (e.g., cell phones) to seamlessly participate in the wireless mesh PTT session is key to supporting the heterogeneous environment commonly found in such settings. This paper presents the architecture and protocol of a robust distributed PTT service for wireless mesh networks. The architecture supports any 802.11 client with SIP-based VoIP software and enables the participation of regular phones. Collectively, the mesh nodes provide the illusion of a single third party call controller, enabling clients to participate via any reachable mesh node. Each PTT group instantiates its own logical floor control manager that is highly available and resilient to mesh connectivity changes such as node crashes and recoveries and network partitions and merges. Experimental results on a fully deployed mesh network consisting of 14 mesh nodes and tens of emulated clients demonstrate the scalability and robustness of the system. Yair Amir, Raluca Musaloiu-Elefteri, Nilo Rivera |
SECON | 1 |
| 2010 | Steward: Scaling Byzantine Fault-Tolerant Replication to Wide Area NetworksabstractThis paper presents the first hierarchical byzantine fault-tolerant replication architecture suitable to systems that span multiple wide-area sites. The architecture confines the effects of any malicious replica to its local site, reduces message complexity of wide-area communication, and allows read-only queries to be performed locally within a site for the price of additional standard hardware. We present proofs that our algorithm provides safety and liveness properties. A prototype implementation is evaluated over several network topologies and is compared with a flat byzantine fault-tolerant approach. The experimental results show considerable improvement over flat byzantine replication algorithms, bringing the performance of byzantine replication closer to existing benign fault-tolerant replication techniques over wide area networks. Yair Amir, Claudiu Danilov 0001, Danny Dolev, Jonathan Kirsch, John Lane, Cristina Nita-Rotaru, Josh Olsen, David Zage |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2010 | The SMesh wireless mesh networkabstractWireless mesh networks extend the connectivity range of mobile devices by using multiple access points, some of them connected to the Internet, to create a mesh topology and forward packets over multiple wireless hops. However, the quality of service provided by the mesh is impaired by the delays and disconnections caused by handoffs, as clients move within the area covered by multiple access points. We present the architecture and protocols of SMesh, the first transparent wireless mesh system that offers seamless, fast handoff, supporting real-time applications such as interactive VoIP. The handoff and routing logic is done solely by the access points, and therefore connectivity is attainable by any 802.11 device. In SMesh, the entire mesh network is seen by the mobile clients as a single, omnipresent access point, giving the mobile clients the illusion that they are stationary. We use multicast for access points coordination and, during handoff transitions, we use more than one access point to handle the moving client. SMesh provides a hybrid routing protocol that optimizes routes over wireless and wired links in a multihomed environment. Experimental results on a fully deployed mesh network demonstrate the effectiveness of the SMesh architecture and its intra-domain and inter-domain handoff protocols. Yair Amir, Claudiu Danilov 0001, Raluca Musaloiu-Elefteri, Nilo Rivera |
ACM Trans. Comput. Syst. | 1 |
| 2009 | Authenticated Adversarial Routing
Yair Amir, Paul Bunn, Rafail Ostrovsky |
TCC | 1 |
| 2008 | Byzantine replication under attackabstractExisting Byzantine-resilient replication protocols satisfy two standard correctness criteria, safety and liveness, in the presence of Byzantine faults. In practice, however, faulty processors can, in some protocols, significantly degrade performance by causing the system to make progress at an extremely slow rate. While ldquocorrectrdquo in the traditional sense, systems vulnerable to such performance degradation are of limited practical use in adversarial environments. This paper argues that techniques for mitigating such performance attacks are needed to bridge this ldquopracticality gaprdquo for intrusion-tolerant replication systems. We propose a new performance-oriented correctness criterion, and we show how failure to meet this criterion can lead to performance degradation. We present a new Byzantine replication protocol that achieves the criterion and evaluate its performance in fault-free configurations and when under attack. Yair Amir, Brian A. Coan, Jonathan Kirsch, John Lane |
DSN | 1 |
| 2007 | Customizable Fault Tolerance for Wide-Area ReplicationabstractConstructing logical machines out of collections of physical machines is a well-known technique for improving the robustness and fault tolerance of distributed systems. We present a new, scalable replication architecture, built upon logical machines specifically designed to perform well in wide-area systems spanning multiple sites. The physical machines in each site implement a logical machine by running a local state machine replication protocol, and a wide-area replication protocol runs among the logical machines. Implementing logical machines via the state machine approach affords free substitution of the fault tolerance method used in each site and in the wide-area replication protocol, allowing one to balance performance and fault tolerance based on perceived risk. We present a new byzantine fault-tolerant protocol that establishes a reliable virtual communication link between logical machines. Our communication protocol is efficient (a necessity in wide-area environments), avoiding the need for redundant message sending during normal-case operation and allowing a logical machine to consume approximately the same wide-area bandwidth as a single physical machine. This dramatically improves the wide-area performance of our system compared to existing logical machine based approaches. We implemented a prototype system and compare its performance and fault tolerance to existing solutions. Yair Amir, Brian A. Coan, Jonathan Kirsch, John Lane |
SRDS | 1 |
| 2007 | An Inter-domain Routing Protocol for Multi-homed Wireless Mesh NetworksabstractThis paper presents an architecture and a hybrid routing protocol for multi-homed wireless mesh networks that provide uninterrupted connectivity and fast handoff. Our approach integrates wireless and wired connectivity, using multicast groups to coordinate decisions and seamlessly transfer connections between several Internet gateways as mobile clients move between access points. The protocol optimizes the use of the wireless medium by short-cutting wireless hops through wired connections, paying a very low overhead during handoffs. The paper demonstrates that inter-domain handoffs occur instantaneously, with virtually no loss or delay, for both TCP and UDP connections. Yair Amir, Claudiu Danilov 0001, Raluca Musaloiu-Elefteri, Nilo Rivera |
WOWMOM | 1 |
| 2006 | Scaling Byzantine Fault-Tolerant Replication toWide Area NetworksabstractThis paper presents the first hierarchical Byzantine fault-tolerant replication architecture suitable to systems that span multiple wide area sites. The architecture confines the effects of any malicious replica to its local site, reduces message complexity of wide area communication, and allows read-only queries to be performed locally within a site for the price of additional hardware. A prototype implementation is evaluated over several network topologies and is compared with a flat Byzantine fault-tolerant approach Yair Amir, Claudiu Danilov 0001, Jonathan Kirsch, John Lane, Danny Dolev, Cristina Nita-Rotaru, Josh Olsen, David Zage |
DSN | 1 |
| 2006 | Fast handoff for seamless wireless mesh networksabstractAbstract — This paper presents the architecture and protocols of SMesh, a completely transparent wireless mesh system that offers seamless, fast handoff, supporting VoIP and other real-time application traffic for any unmodified 802.11 device. In SMesh, the entire mesh network is seen by the mobile clients as a single, omnipresent access point. Fast handoff is achieved by ensuring that each client is served by at least one access point at any time. Mobile clients are handled by a single access point during stable connectivity times. During handoff transitions, SMesh uses more than one access point to handle the moving client. Access points continuously monitor the connectivity quality of any client in their range and efficiently share this information with other access points in the vicinity of that client to coordinate which of them should serve the client. Experimental results on a fully deployed mesh network consisting of 14 access points demonstrate the effectiveness of the SMesh architecture and its handoff protocol. I. Yair Amir, Claudiu Danilov 0001, Michael Hilsdale, Raluca Musaloiu-Elefteri, Nilo Rivera |
MobiSys | 1 |
| 2006 | An Overlay Architecture for High-Quality VoIP StreamsabstractThe cost savings and novel features associated with voice over IP (VoIP) are driving its adoption by service providers. Unfortunately, the Internet's best effort service model provides no quality of service guarantees. Because low latency and jitter are the key requirements for supporting high-quality interactive conversations, VoIP applications use UDP to transfer data, thereby subjecting themselves to quality degradations caused by packet loss and network failures. In this paper, we describe an architecture to improve the performance of such VoIP applications. Two protocols are used for localized packet loss recovery and rapid rerouting in the event of network failures. The protocols are deployed on the nodes of an application-level overlay network and require no changes to the underlying infrastructure. Experimental results indicate that the architecture and protocols can be combined to yield voice quality on par with the public switched telephone network Yair Amir, Claudiu Danilov 0001, Stuart Goose, David Hedqvist, Andreas Terzis |
IEEE Trans. Multim. | 1 |
| 2005 | 1-800-OVERLAYS: using overlay networks to improve VoIP qualityabstractThe cost savings and novel features associated with Voice over IP (VoIP) are driving its adoption by service providers. Such a transition however can successfully happen only if the quality and reliability offered is comparable to the existing PSTN. Unfortunately, the Internet's best effort service model provides no inherent quality of service guarantees. Because low latency and jitter is the key requirement for supporting high quality interactive conversations, VoIP applications use UDP to transfer data, thereby subjecting themselves to performance degradations caused by packet loss and network failures.In this paper we describe two algorithms to improve the performance of such VoIP applications. These mechanisms are used for localized packet loss recovery and rapid rerouting in the event of network failures. The algorithms are deployed on the routers of an application-level overlay network and require no changes to the underlying infrastructure. Initial experimental results indicate that these two approaches can be composed to yield voice quality on par with the PSTN. Yair Amir, Claudiu Danilov 0001, Stuart Goose, David Hedqvist, Andreas Terzis |
NOSSDAV | 1 |
| 2005 | Secure Spread: An Integrated Architecture for Secure Group CommunicationabstractGroup communication systems are high-availability distributed systems providing reliable and ordered message delivery, as well as a membership service, to group-oriented applications. Many such systems are built using a distributed client-server architecture where a relatively small set of servers provide service to numerous clients. In this work, we show how group communication systems can be enhanced with security services without sacrificing robustness and performance. More specifically, we propose several integrated security architectures for distributed client-server group communication systems. In an integrated architecture, security services are implemented in servers, in contrast to a layered architecture, where the same services are implemented in clients. We discuss performance and accompanying trust issues of each proposed architecture and present experimental results that demonstrate the superior scalability of an integrated architecture. Yair Amir, Cristina Nita-Rotaru, Jonathan Robert Stanton, Gene Tsudik |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2005 | A cost-benefit flow control for reliable multicast and unicast in overlay networksabstractWhen many parties share network resources on an overlay network, mechanisms must exist to allocate the resources and protect the network from overload. Compared to large physical networks such as the Internet, in overlay networks the dimensions of the task are smaller, so new and possibly more effective techniques can be used. In this work we take a fresh look at the problem of flow control in multisender multigroup reliable multicast and unicast and explore a cost-benefit approach that works in conjunction with Internet standard protocols such as TCP. In contrast to existing window-based flow control schemes, we avoid end-to-end per sender or per group feedback by looking only at the state of the virtual links between participating nodes. This produces control traffic proportional only to the number of overlay network links and independent of the number of groups, senders, or receivers. We show the effectiveness of the resulting protocol through simulations and validate the simulations with live Internet experiments. We demonstrate near-optimal utilization of network resources, fair sharing of individual congested links, and quick adaptation to network changes. Yair Amir, Baruch Awerbuch, Claudiu Danilov 0001, Jonathan Robert Stanton |
IEEE/ACM Trans. Netw. | 1 |
| 2004 | On the performance of group key agreement protocolsabstractGroup key agreement is a fundamental building block for secure peer group communication systems. Several group key management techniques were proposed in the last decade, all assuming the existence of an underlying group communication infrastructure to provide reliable and ordered message delivery as well as group membership information. Despite analysis, implementation, and deployment of some of these techniques, the actual costs associated with group key management have been poorly understood so far. This resulted in an undesirable tendency: on the one hand, adopting suboptimal security for reliable group communication, while, on the other hand, constructing excessively costly group key management protocols.This paper presents a thorough performance evaluation of five notable distributed key management techniques (for collaborative peer groups) integrated with a reliable group communication system. An in-depth comparison and analysis of the five techniques is presented based on experimental results obtained in actual local- and wide-area networks. The extensive performance measurement experiments conducted for all methods offer insights into their scalability and practicality. Furthermore, our analysis of the experimental results highlights several observations that are not obvious from the theoretical analysis. Yair Amir, Yongdae Kim, Cristina Nita-Rotaru, Gene Tsudik |
ACM Trans. Inf. Syst. Secur. | 1 |
| 2004 | Secure Group Communication Using Robust Contributory Key AgreementabstractContributory group key agreement protocols generate group keys based on contributions of all group members. Particularly appropriate for relatively small collaborative peer groups, these protocols are resilient to many types of attacks. Unlike most group key distribution protocols, contributory group key agreement protocols offer strong security properties such as key independence and perfect forward secrecy. We present the first robust contributory key agreement protocol resilient to any sequence of group changes. The protocol, based on the Group Diffie-Hellman contributory key agreement, uses the services of a group communication system supporting virtual synchrony semantics. We prove that it provides both virtual synchrony and the security properties of Group Diffie-Hellman, in the presence of any sequence of (potentially cascading) node failures, recoveries, network partitions, and heals. We implemented a secure group communication service, Secure Spread, based on our robust key agreement protocol and Spread group communication system. To illustrate its practicality, we compare the costs of establishing a secure group with the proposed protocol and a protocol based on centralized group key management, adapted to offer equivalent security properties. Yair Amir, Yongdae Kim, Cristina Nita-Rotaru, John L. Schultz, Jonathan Robert Stanton, Gene Tsudik |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2003 | N-Way Fail-Over Infrastructure for Reliable Servers and RoutersabstractMaintaining the availability of critical servers and routers is an important concern for many organizations. At the lowest level, IP addresses represent the global namespace by which services are accessible on the Internet. We introduce Wackamole, a completely distributed software solution based on a provably correct algorithm that negotiates the assignment of IP addresses among the currently available servers upon detection of faults. This reallocation ensures that at any given time any public IP address of the server cluster is covered exactly once, as long as at least one physical server survives the network fault. The same technique is extended to support highly available routers. The paper presents the design considerations, algorithm specification and correctness proof, discusses the practical usage for server clusters and for routers, and evaluates the performance of the system. 1 Yair Amir, Ryan Caudy, Ashima Munjal, Theo Schlossnagle, Ciprian Tutu |
DSN | 1 |
| 2003 | Reliable Communication in Overlay NetworksabstractReliable point-to-point communication is usually achieved in overlay networks by applying TCP on the end nodes of a connection. This paper presents a hop-by-hop reliability approach that considerably reduces the latency and jitter of reliable connections. Our approach is feasible and beneficial in overlay networks that do not have the scalability and interoperability requirements of the global Internet. The effects of the hop-by-hop reliability approach are quantified in simulation as well as in practice using a newly developed overlay network system that is fair with the ex-ternal traffic on the Internet. The experimental results show that the overhead associated with overlay network process-ing at the application level does not play an important fac-tor compared with the considerable gain of the approach. 1 Yair Amir, Claudiu Danilov 0001 |
DSN | 1 |
| 2002 | On the Performance of Group Key Agreement ProtocolsabstractGroup key agreement (GKA) is a fundamental building block for securing peer group communication systems (GCS). Several group key agreement protocols were proposed in the past, all assuming an underlying group communication infrastructure. This paper presents a performance evaluation of 5 notable GKA protocols integrated with a reliable group communication system (Spread). They are: centralized group key distribution (CKD), Burmester-Desmedt (BD), Steer et al. (STR), group Diffie-Hellman GDH) and tree-based group Diffie-Hellman (TGDH).. We present concrete results obtained in experiments on local- and wide-area networks. Our analysis of these results offers insights into their relative scalability and practicality. Yair Amir, Yongdae Kim, Cristina Nita-Rotaru, Gene Tsudik |
ICDCS | 1 |
| 2002 | From Total Order to Database ReplicationabstractThis paper presents in detail an efficient and provably correct algorithm for database replication over partitionable networks. Our algorithm avoids the need for end-to-end acknowledgments for each action while supporting network partitions and merges and allowing dynamic instantiation of new replicas. One round of end-to-end acknowledgments is required only upon a membership change event such as a network partition. New actions may be introduced to the system at any point, not only while in a primary component. We show how performance can be further improved for applications that allow relaxation of consistency requirements. We provide experimental results that demonstrate the efficiency of our approach. Yair Amir, Ciprian Tutu |
ICDCS | 1 |
| 2001 | Exploring Robustness in Group Key AgreementabstractSecure group communication is crucial for building distributed applications that work in dynamic environments and communicate over unsecured networks (e.g. the Internet). Key agreement is a critical part of providing security services for group communication systems. Most of the current contributory key agreement protocols are not designed to tolerate failures and membership changes during execution. In particular, nested or cascaded group membership events (such as partitions) are not accommodated. We present the first robust contributory key agreement protocols, resilient to any sequence of events while preserving the group communication membership and ordering guarantees. Yair Amir, Cristina Nita-Rotaru, John L. Schultz, Jonathan Robert Stanton, Yongdae Kim, Gene Tsudik |
ICDCS | 1 |
| 2000 | A Low Latency, Loss Tolerant Architecture and Protocol for Wide Area Group CommunicationabstractGroup communication systems are proven tools upon which to build fault-tolerant systems. As the demands for fault-tolerance increase and more applications require reliable distributed computing over wide area networks, wide area group communication systems are becoming very useful. However, building a wide area group communication system is a challenge. This paper presents the design of the transport protocols of the spread wide area group communication system. We focus on two aspects of the system. First, the value of using overlay networks for application level group communication services. Second, the requirements and design of effective low latency link protocols used to construct wide area group communication. We support our claims with the results of live experiments conducted over the Internet. Yair Amir, Claudiu Danilov 0001, Jonathan Robert Stanton |
DSN | 1 |
| 2000 | Secure Group Communication in Asynchronous Networks with Failures: Integration and ExperimentsabstractThe increasing popularity and diversity of collaborative applications prompts a need for highly secure and reliable communication platforms for dynamic peer groups. Security mechanisms for such groups tend to be both expensive and complex and their integration with reliable group communication services presents a formidable challenge, This paper discusses some important integration issues, reports on our implementation experience and provides experimental results. Our approach utilizes distributed group key management developed by the Cliques project. We enhance it to handle processor and network faults (under a fail-stop or crash-and-recover model) and asynchronous membership events (such as joins, leaves, merges and network partitions). Our approach leverages the strong properties provided by the Spread group communication system, such as message ordering, clean failure semantics and a membership service. The result of this work is a secure group communications layer and an API that provide the application programmer with both standard group communication services and flexible security services. Jonathan Robert Stanton, Yair Amir, Damian Hasse, Giuseppe Ateniese, Yongdae Kim, Cristina Nita-Rotaru, Theo Schlossnagle, John L. Schultz, Gene Tsudik |
ICDCS | 2 |
| 2000 | A Cost-Benefit framework for online management of a metacomputing systemabstractManaging a large collection of networked machines, with a series of incoming jobs, requires that the jobs be assigned to machines wisely. A new approach to this problem is presented, inspired by economic principles: the Cost–Benefit framework. This framework simplifies complex assignment and admission control decisions, and performs well in practice. We demonstrate this framework in the context of an Internet-wide market for computational services and verify its utility for a classic network of workstations. Yair Amir, Baruch Awerbuch, R. Sean Borgstrom |
Decis. Support Syst. | 1 |
| 2000 | An Opportunity Cost Approach for Job Assignment in a Scalable Computing ClusterabstractA new method is presented for job assignment to and reassignment between machines in a computing cluster. Our method is based on a theoretical framework that has been experimentally tested and shown to be useful in practice. This "opportunity cost" method converts the usage of several heterogeneous resources in a machine to a single homogeneous "cost." Assignment and reassignment are then performed based on that cost. This is in contrast to traditional, ad hoc methods for job assignment and reassignment. These treated each resource as an independent entity with its own constraints, as there was no clean way to balance one resource against another. Our method has been tested by simulations, as well as real executions, and was found to perform well. Yair Amir, Baruch Awerbuch, Amnon Barak, R. Sean Borgstrom, Arie Keren |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1998 | Seamlessly Selecting the Best Copy from Internet-Wide Replicated Web Servers
Yair Amir, Alec Peterson, David Shaw |
DISC | 1 |
| 1998 | Optimal Availability Quorum Systems: Theory and Practice
Yair Amir, Avishai Wool |
Inf. Process. Lett. | 1 |
| 1996 | Evaluating Quorum Systems Over the Internet (Abstract)abstractNo abstract available. Yair Amir, Avishai Wool |
PODC | 1 |
| 1995 | The Totem Single-Ring Ordering and Membership ProtocolabstractFault-tolerant distributed systems are becoming more important, but in existing systems, maintaining the consistency of replicated data is quite expensive. The Totem single-ring protocol supports consistent concurrent operations by placing a total order on broadcast messages. This total order is derived from the sequence number in a token that circulates around a logical ring imposed on a set of processors in a broadcast domain. The protocol handles reconfiguration of the system when processors fail and restart or when the network partitions and remerges. Extended virtual synchrony ensures that processors deliver messages and configuration changes to the application in a consistent, systemwide total order. An effective flow control mechanism enables the Totem single-ring protocol to achieve message-ordering rates significantly higher than the best prior total-ordering protocols. Yair Amir, Louise E. Moser, P. M. Melliar-Smith, Deborah A. Agarwal, P. Ciarfella |
ACM Trans. Comput. Syst. | 1 |
| 1994 | Extended Virtual SynchronyabstractWe formulate a model of extended virtual synchrony that defines a group communication transport service for multicast and broadcast communication in a distributed system. The model extends the virtual synchrony model of the Isis system to support continued operation in all components of a partitioned network. The significance of extended virtual synchrony is that, during network partitioning and remerging and during process failure and recovery, it maintains a consistent relationship between the delivery of messages and the delivery of configuration changes across all processes in the system and provides well-defined self-delivery and failure atomicity properties. We describe an algorithm that implements extended virtual synchrony and construct a filter that reduces extended virtual synchrony to virtual synchrony.> Louise E. Moser, Yair Amir, P. M. Melliar-Smith, Deborah A. Agarwal |
ICDCS | 2 |
| 1993 | Fast Message Ordering and Membership Using a Logical Token-Passing RingabstractThe Totem protocol supports consistent concurrent operations by placing a total order on broadcast messages. This total order is achieved by including a sequence number in a token circulated around a logical ring that is imposed on a set of processors in a broadcast domain. A membership algorithm handles reconfiguration, including restarting of a failed processor and remerging of a partitioned network. Effective flow-control allows the protocol to achieve message ordering rates two to three times higher than the best prior protocols. The single-ring total ordering protocol of Totem provides fault-tolerant agreed and safe delivery of messages within a broadcast domain.> Yair Amir, Louise E. Moser, P. M. Melliar-Smith, Deborah A. Agarwal, P. Ciarfella |
ICDCS | 1 |