Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

E. N. Elnozahy

dblp:e/ENElnozahy · also Elmootazbellah (Mootaz) Elnozahy · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-authorComputer networks · 2Security and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Distributed systems · 49% Energy-efficient computing · 20% Embedded and real-time systems · 13%
Computer networks
2 papers
Internet architecture and protocols · 50% Routing and switching · 50%
Network and information security
2 papers
Network security · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.152004
Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002
Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998
Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004
Distributed systems › fault tolerance
rollback recovery
0.142004
Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004
Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998
On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995
Energy-efficient computing
power management
0.122004
The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004
Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002
Distributed systems › fault tolerance
checkpointing
0.122004
Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004
Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992
Distributed systems › fault tolerance › checkpointing
checkpoint placement
0.012004
The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.012004
The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004
Embedded and real-time systems
real-time scheduling
0.012004
The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004
Network security › attack resilience › attack mitigation
denial-of-service defense
0.012002
Hop integrity in computer networks · IEEE/ACM Trans. Netw. 2002
Embedded and real-time systems › energy-efficient embedded systems
energy-efficient real-time systems
0.012002
Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002
Hardware reliability and fault tolerance › redundancy
modular redundancy
0.012002
Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002
Routing and switching › routing
secure routing
0.012000
Hop Integrity in Computer Networks · ICNP 2000
Network security › protocol security
network protocol security
0.012000
Hop Integrity in Computer Networks · ICNP 2000
Performance modeling and evaluation
loop detection
0.011999
Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999
Performance modeling and evaluation
workload characterization
0.011999
Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999
Hardware reliability and fault tolerance
error recovery
0.012004
The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004
Distributed systems
distributed coordination and fault tolerance
0.011995
On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995
Storage systems
logging
0.011995
On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995
Distributed systems › fault tolerance
message logging
0.011992
Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992
Distributed systems › fault tolerance › checkpointing
uncoordinated checkpointing
0.011992
Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992
Routing and switching
inter-domain routing
0.012000
Hop Integrity in Computer Networks · ICNP 2000
Performance modeling and evaluation › workload characterization
address trace
0.011999
Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999
Debugging and program repair › record and replay
deterministic replay
0.011998
Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998
Distributed systems
distributed coordination
0.011992
Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992

Methods — techniques the papers use, named apart from their topics

weak integrity protocol · 0.1strong integrity protocol · 0.1secret exchange protocol · 0.1technology roadmap · 0.0slack exploitation · 0.0performance projection · 0.0checkpointing · 0.0user-level thread package · 0.0instruction counter · 0.0energy efficiency analysis · 0.0trace compression · 0.0loop detection · 0.0performance evaluation · 0.0
YearPublicationVenuePosition
2007 Five Years with the High Productivity Computing Systems Program A Perspective
abstract
Summary form only given. For the past five years, I had the very enviable task of leading IBM's effort in DARPA's High Productivity Computing Systems (HPCS) program. IBM competed successfully with other contestants in and survived two down-selects, producing along the way ground-breaking research for peta-scale systems aimed at changing the status quo in high end computing. The HPCS program is unique in that it states productivity as a broader definition of the system value than just performance. Commercial viability is another goal, meant to add realism and produce usable systems at the end of the program with productivity and performance goals that well exceed the projected improvements using today's technology. This unprecedented mix adds interesting and challenging constraints on the research program, and the traditional ways of approaching the problem do not apply. This talk will give an overview of the challenges of running projects of this kind, and gives a forward looking statement about the future of the program and its projected impact on the industry and the academic communities.
E. N. Elnozahy
IPDPS1
2006 Four Years with the High Productivity Computing Systems Program - A Perspective
abstract
For the past four years, IBM has participated in DARPA's high productivity computing systems (HPCS) program, competing with other contestants in ground-breaking research for peta-scale systems aimed at changing the status quo in high end computing. The HPCS program is unique in that it states productivity as a broader definition of the system value than just performance. Commercial viability is another goal, meant to add realism and produce usable systems at the end of the program with productivity and performance goals that well exceed the projected improvements using today's technology. This unprecedented mix adds interesting and challenging constraints on the research program, and the traditional ways of approaching the problem do not apply. This talk gives an overview of the program as conducted in IBM, including a description of many technologies that were investigated and considered. The talk addresses also the challenges of running projects of this kind, and gives a forward looking statement about the future of the program and its projected impact on the industry and the academic communities
E. N. Elnozahy
ICPP1
2004 Analysis of an Energy Efficient Optimistic TMR Scheme
Dakai Zhu 0001, Rami G. Melhem, Daniel Mossé, E. N. Elnozahy
ICPADS4
2004 The Interplay of Power Management and Fault Recovery in Real-Time Systems
abstract
We describe how to exploit the scheduling slack in a real-time system to reduce energy consumption and achieve fault tolerance at the same time. During failure-free operation, a task takes checkpoints to enable recovery from failure. Additionally, the system exploits the slack to conserve energy by reducing the processor speed. If a task fails, it will restart from a saved checkpoint and execute at maximum speed to guarantee that the deadlines are met. We show that the number of checkpoints and their placements interact in subtle ways with the power management policy. We study two checkpoint placement policies for aperiodic tasks and analytically derive the optimal number of checkpoints to conserve energy under each. This optimal number allows the CPU speed to be slowed down to the level that yields minimum energy consumption, while still guaranteeing recoverability of tasks under each checkpointing policy. The results show that traditional periodic checkpointing is not the best policy for the combined purpose of conserving energy and guaranteeing recovery. Instead, better energy savings are possible through a nonuniform distribution of checkpoints that takes into account the energy consumption and reliability factors. Depending on the amount of slack and the checkpointing overhead, energy can be reduced by up to 68 percent under nonuniform checkpointing. We also demonstrate the applicability of these checkpoint placement policies to periodic tasks.
Rami G. Melhem, Daniel Mossé, E. N. Elnozahy
IEEE Trans. Computers3
2004 Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery
abstract
Over the past two decades, rollback-recovery via checkpoint-restart has been used with reasonable success for long-running applications, such as scientific workloads that take from few hours to few months to complete. Currently, several commercial systems and publicly available libraries exist to support various flavors of checkpointing. Programmers typically use these systems if they are satisfactory or otherwise embed checkpointing support themselves within the application. In this paper, we project the performance and functionality of checkpointing algorithms and systems as we know them today into the future. We start by surveying the current technology roadmap and particularly how Peta-Flop capable systems may be plausibly constructed in the next few years. We consider how rollback-recovery as practiced today will fare when systems may have to be constructed out of thousands of nodes. Our projections predict that, unlike current practice, the effect of rollback-recovery may play a more prominent role in how systems may be configured to reach the desired performance level. System planners may have to devote additional resources to enable rollback-recovery and the current practice of using "cheap commodity" systems to form large-scale clusters may face serious obstacles. We suggest new avenues for research to react to these trends.
E. N. Elnozahy, James S. Plank
IEEE Trans. Dependable Secur. Comput.1
2002 Key Trees and the Security of Interval Multicast
abstract
A key tree is a distributed data structure of security keys that can be used by a group of users. In this paper we describe how any user in the group can use the different keys in the key tree to securely multicast data to different subgroups within the group. The cost of securely multicasting data to a subgroup whose users are "consecutive" is O(log n) encryptions, where n is the total number of users in the group. The cost of securely multicasting data to an arbitrary subgroup is O(n/2) encryptions. However this cost can be reduced to one encryption by introducing an additional key tree to the group.
Mohamed G. Gouda, Chin-Tser Huang, E. N. Elnozahy
ICDCS3
2002 Energy-Efficient Duplex and TMR Real-Time Systems
abstract
Duplex and triple modular redundancy (TMR) systems are used when a high-level of reliability is desired. Real-time systems for autonomous critical missions need such degrees of reliability, but energy consumption becomes a dominant concern when these systems are built from high-performance processors that consume a large budget of electrical power for operation and cooling. Examples where energy consumption and real time are of paramount importance include reliable computers onboard mobile vehicles, such as the Mars Rover, satellites, and other autonomous vehicles. At first inspection, a duplex system uses about two thirds of the components that a TMR system does, leading one to conclude that duplex systems are more energy-efficient. This paper shows that this is not always the case. We present an analysis of the energy efficiency of duplex and TMR systems when used to tolerate transient failures. With no power management deployed, the analysis supports the intuitive impression about the relative superiority of duplex systems in energy consumption. The analysis shows, however that the gap in energy consumption between the two types of systems diminishes with proper power management. We introduce the concept of an optimistic TMR system that offers the same reliability and performance as the traditional one, but at a fraction of the energy consumption budget. Optimistic TMR systems are competitive with respect to energy consumption when compared with a power-aware duplex system, can even exceed it in some situations, and have the added bonus of providing tolerance to permanent faults.
E. N. Elnozahy, Rami G. Melhem, Daniel Mossé
RTSS1
2002 Hop integrity in computer networks
abstract
A computer network is said to provide hop integrity if, when any router, p, in the network receives a message, m, supposedly from an adjacent router, q, then p can check that m was indeed sent by q, was not modified after it was sent and was not a replay of an old message sent from q to p. We describe three protocols that can be added to the routers in a computer network so that the network can provide hop integrity, and thus overcome most denial-of-service attacks. These three protocols are a secret exchange protocol, a weak integrity protocol and a strong integrity protocol. All three protocols are stateless, require small overhead and do not constrain the network protocol in the routers in any way.
Mohamed G. Gouda, E. N. Elnozahy, Chin-Tser Huang, Tommy M. McGuire
IEEE/ACM Trans. Netw.2
2000 Hop Integrity in Computer Networks
abstract
A computer network is said to provide hop integrity iff when any router p in the network receives a message m supposedly from an adjacent router q, then p can check that m was indeed sent by q, was not modified after it was sent, and was not a replay of an old message sent from q to p. We describe three protocols that can be added to the routers in a computer network so that the network can provide hop integrity. These three protocols are a secret exchange protocol, a weak integrity protocol, and a strong integrity protocol. All three protocols are stateless, require small overhead, and do not constrain the network protocol in the routers in any way.
Mohamed G. Gouda, E. N. Elnozahy, Chin-Tser Huang, Tommy M. McGuire
ICNP2
1999 Address Trace Compression Through Loop Detection and Reduction
abstract
No abstract available.
E. N. Elnozahy
SIGMETRICS1
1998 Support for Software Interrupts in Log-Based Rollback-Recovery
abstract
The piecewise deterministic execution model is a fundamental assumption in many log-based rollback-recovery protocols. Process execution in this model consists of intervals, each starting with the receipt of a message at an application-defined execution point. Execution within each interval is deterministic and messages are the only source of nondeterminism that affects the computation. This simple model excludes the nondeterminism that results when asynchronous signals or interrupts occur at arbitrary execution points. As a result, a wide range of applications cannot use log-based rollback-recovery in practice. We present a solution that removes this restriction and allows applications to replay interrupts at the same execution points during recovery. The solution relies on using a software counter to compute the number of instructions between the asynchronous signals during normal operation. Should a failure occur, the instruction counts are used to force the replay of these signals at the same execution points. The execution of the application thus can be replayed to recreate the prefailure state while accommodating nondeterminism due to asynchronous signals. We then use the deterministic replay of interrupts to solve another problem, namely tracking nondeterminism due to interleaved shared memory access in multithreaded applications on a single processor. We use the instruction counter solution to implement a user-level thread package in which thread scheduling decisions can be replayed if a failure occurs. By repeating the scheduling decisions during an execution replay, threads access the shared memory in the same order and the execution to be reconstructed. This technique allows multithreaded applications to use log-based rollback-recovery with low overhead, which was not previously possible. We carried out two prototype implementations that have shown the overhead is no more than a 6 percent slowdown in application execution on the DEC Alpha, and from 6 percent to 18 percent on the Inter Pentium. Thus, restrictions of the piecewise deterministic execution model can be lifted at a reasonable cost.
J. Hamilton Slye, E. N. Elnozahy
IEEE Trans. Computers2
1995 On the Relevance of Communication Costs of Rollback-Recovery Protocols
abstract
Communication overhead has been traditionally the primary metric for evaluating rollback-recovery protocols.This paper reexamines the prominence of this metric in light of the recent increases in processor and network speeds.We introduce a new recovery algorithm for a family of rolibackrecovery protocols based on logging.The new algorithm incurs a higher communication overhead during recovery than previous algorithms, but it requires less access to stable storage and imposes no restrictions on the execution of live processes.Experimental results show that the new algorithm performs better than one that is optimized for low communication overhead.These results suggest that in modern environments, latency in accessing stable storage and intrusion of a particular algorithm on the execution of live processes are more important than the number of messages exchanged during recovery.
E. N. Elnozahy
PODC1
1992 The Performance of Consistent Checkpointing
abstract
Consistent checkpointing provides transparent fault tolerance for long-running distributed applications. Performance measurements of an implementation of consistent checkpointing are described. The measurements show that consistent checkpointing performs remarkably well. Eight computation-intensive distributed applications were executed on a network of 16 diskless Sun-3/60 workstations, and the performance without checkpointing was compared to the performance with consistent checkpoints taken at two-minute intervals. For six of the eight applications, the running time increased by less than 1% as a result of the checkpointing. The highest overhead measured was 5.8%. Incremental checkpointing and copy-on write checkpointing were the most effective techniques in lowering the running time overhead. It is argued that these measurements show that consistent checkpointing is an efficient way to provide fault tolerance for long-running distributed applications.>
E. N. Elnozahy, David B. Johnson 0001, Willy Zwaenepoel
SRDS1
1992 Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit
abstract
Manetho is a new transparent rollback-recovery protocol for long-running distributed computations. It uses a novel combination of antecedence graph maintenance, uncoordinated checkpointing, and sender-based message logging. Manetho simultaneously achieves the advantages of pessimistic message logging, namely limited rollback and, fast output commit, and the advantage of optimistic message logging, namely low failure-free overhead. These advantages come at the expense of a complex recovery scheme.>
E. N. Elnozahy, Willy Zwaenepoel
IEEE Trans. Computers1
1991 A comparison of two approaches to build reliable distributed file servers
abstract
Several existing distributed file systems provide reliability by server replication. An alternative approach is to use dual-ported disks accessible to a server and a backup. The two approaches are compared by examining an example of each. Deceit is a replicated file server that emphasizes flexibility. HA-NFS is an example of the second approach that emphasizes efficiency and simplicity. The two file servers run on the same hardware and implement SUN's NFS protocol. The comparison shows that replicated servers are more flexible and tolerant of a wider variety of faults. On the other hand, the dual-ported disks approach is more efficient and simpler to implement. When tolerating single failure, dual-ported disks also give somewhat better availability.>
Anupam Bhide, E. N. Elnozahy, Stephen P. Morgan, A. Siegel
ICDCS2