EDBT 2026 Demo / reviewers in the wild / expert
E. N. Elnozahy
dblp:e/ENElnozahy · also Elmootazbellah (Mootaz) Elnozahy
· DBLP profile ↗
15ranked-venue papers
8as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-authorComputer networks · 2Security and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Distributed systems · 49% Energy-efficient computing · 20% Embedded and real-time systems · 13% | |
| Computer networks
2 papers |
Internet architecture and protocols · 50% Routing and switching · 50% | |
| Network and information security
2 papers |
Network security · 100% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.1 | 5 | 2004 | Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002 Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998 Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004 |
Distributed systems › fault tolerance
rollback recovery |
0.1 | 4 | 2004 | Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004 Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998 On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995 |
Energy-efficient computing
power management |
0.1 | 2 | 2004 | The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004 Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002 |
Distributed systems › fault tolerance
checkpointing |
0.1 | 2 | 2004 | Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-Recovery · IEEE Trans. Dependable Secur. Comput. 2004 Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992 |
Distributed systems › fault tolerance › checkpointing
checkpoint placement |
0.0 | 1 | 2004 | The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.0 | 1 | 2004 | The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004 |
Embedded and real-time systems
real-time scheduling |
0.0 | 1 | 2004 | The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004 |
Network security › attack resilience › attack mitigation
denial-of-service defense |
0.0 | 1 | 2002 | Hop integrity in computer networks · IEEE/ACM Trans. Netw. 2002 |
Embedded and real-time systems › energy-efficient embedded systems
energy-efficient real-time systems |
0.0 | 1 | 2002 | Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002 |
Hardware reliability and fault tolerance › redundancy
modular redundancy |
0.0 | 1 | 2002 | Energy-Efficient Duplex and TMR Real-Time Systems · RTSS 2002 |
Routing and switching › routing
secure routing |
0.0 | 1 | 2000 | Hop Integrity in Computer Networks · ICNP 2000 |
Network security › protocol security
network protocol security |
0.0 | 1 | 2000 | Hop Integrity in Computer Networks · ICNP 2000 |
Performance modeling and evaluation
loop detection |
0.0 | 1 | 1999 | Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1999 | Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999 |
Hardware reliability and fault tolerance
error recovery |
0.0 | 1 | 2004 | The Interplay of Power Management and Fault Recovery in Real-Time Systems · IEEE Trans. Computers 2004 |
Distributed systems
distributed coordination and fault tolerance |
0.0 | 1 | 1995 | On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995 |
Storage systems
logging |
0.0 | 1 | 1995 | On the Relevance of Communication Costs of Rollback-Recovery Protocols · PODC 1995 |
Distributed systems › fault tolerance
message logging |
0.0 | 1 | 1992 | Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992 |
Distributed systems › fault tolerance › checkpointing
uncoordinated checkpointing |
0.0 | 1 | 1992 | Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992 |
Routing and switching
inter-domain routing |
0.0 | 1 | 2000 | Hop Integrity in Computer Networks · ICNP 2000 |
Performance modeling and evaluation › workload characterization
address trace |
0.0 | 1 | 1999 | Address Trace Compression Through Loop Detection and Reduction · SIGMETRICS 1999 |
Debugging and program repair › record and replay
deterministic replay |
0.0 | 1 | 1998 | Support for Software Interrupts in Log-Based Rollback-Recovery · IEEE Trans. Computers 1998 |
Distributed systems
distributed coordination |
0.0 | 1 | 1992 | Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output Commit · IEEE Trans. Computers 1992 |
Methods — techniques the papers use, named apart from their topics
weak integrity protocol · 0.1strong integrity protocol · 0.1secret exchange protocol · 0.1technology roadmap · 0.0slack exploitation · 0.0performance projection · 0.0checkpointing · 0.0user-level thread package · 0.0instruction counter · 0.0energy efficiency analysis · 0.0trace compression · 0.0loop detection · 0.0performance evaluation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Five Years with the High Productivity Computing Systems Program A PerspectiveabstractSummary form only given. For the past five years, I had the very enviable task of leading IBM's effort in DARPA's High Productivity Computing Systems (HPCS) program. IBM competed successfully with other contestants in and survived two down-selects, producing along the way ground-breaking research for peta-scale systems aimed at changing the status quo in high end computing. The HPCS program is unique in that it states productivity as a broader definition of the system value than just performance. Commercial viability is another goal, meant to add realism and produce usable systems at the end of the program with productivity and performance goals that well exceed the projected improvements using today's technology. This unprecedented mix adds interesting and challenging constraints on the research program, and the traditional ways of approaching the problem do not apply. This talk will give an overview of the challenges of running projects of this kind, and gives a forward looking statement about the future of the program and its projected impact on the industry and the academic communities. E. N. Elnozahy |
IPDPS | 1 |
| 2006 | Four Years with the High Productivity Computing Systems Program - A PerspectiveabstractFor the past four years, IBM has participated in DARPA's high productivity computing systems (HPCS) program, competing with other contestants in ground-breaking research for peta-scale systems aimed at changing the status quo in high end computing. The HPCS program is unique in that it states productivity as a broader definition of the system value than just performance. Commercial viability is another goal, meant to add realism and produce usable systems at the end of the program with productivity and performance goals that well exceed the projected improvements using today's technology. This unprecedented mix adds interesting and challenging constraints on the research program, and the traditional ways of approaching the problem do not apply. This talk gives an overview of the program as conducted in IBM, including a description of many technologies that were investigated and considered. The talk addresses also the challenges of running projects of this kind, and gives a forward looking statement about the future of the program and its projected impact on the industry and the academic communities E. N. Elnozahy |
ICPP | 1 |
| 2004 | Analysis of an Energy Efficient Optimistic TMR Scheme
Dakai Zhu 0001, Rami G. Melhem, Daniel Mossé, E. N. Elnozahy |
ICPADS | 4 |
| 2004 | The Interplay of Power Management and Fault Recovery in Real-Time SystemsabstractWe describe how to exploit the scheduling slack in a real-time system to reduce energy consumption and achieve fault tolerance at the same time. During failure-free operation, a task takes checkpoints to enable recovery from failure. Additionally, the system exploits the slack to conserve energy by reducing the processor speed. If a task fails, it will restart from a saved checkpoint and execute at maximum speed to guarantee that the deadlines are met. We show that the number of checkpoints and their placements interact in subtle ways with the power management policy. We study two checkpoint placement policies for aperiodic tasks and analytically derive the optimal number of checkpoints to conserve energy under each. This optimal number allows the CPU speed to be slowed down to the level that yields minimum energy consumption, while still guaranteeing recoverability of tasks under each checkpointing policy. The results show that traditional periodic checkpointing is not the best policy for the combined purpose of conserving energy and guaranteeing recovery. Instead, better energy savings are possible through a nonuniform distribution of checkpoints that takes into account the energy consumption and reliability factors. Depending on the amount of slack and the checkpointing overhead, energy can be reduced by up to 68 percent under nonuniform checkpointing. We also demonstrate the applicability of these checkpoint placement policies to periodic tasks. Rami G. Melhem, Daniel Mossé, E. N. Elnozahy |
IEEE Trans. Computers | 3 |
| 2004 | Checkpointing for Peta-Scale Systems: A Look into the Future of Practical Rollback-RecoveryabstractOver the past two decades, rollback-recovery via checkpoint-restart has been used with reasonable success for long-running applications, such as scientific workloads that take from few hours to few months to complete. Currently, several commercial systems and publicly available libraries exist to support various flavors of checkpointing. Programmers typically use these systems if they are satisfactory or otherwise embed checkpointing support themselves within the application. In this paper, we project the performance and functionality of checkpointing algorithms and systems as we know them today into the future. We start by surveying the current technology roadmap and particularly how Peta-Flop capable systems may be plausibly constructed in the next few years. We consider how rollback-recovery as practiced today will fare when systems may have to be constructed out of thousands of nodes. Our projections predict that, unlike current practice, the effect of rollback-recovery may play a more prominent role in how systems may be configured to reach the desired performance level. System planners may have to devote additional resources to enable rollback-recovery and the current practice of using "cheap commodity" systems to form large-scale clusters may face serious obstacles. We suggest new avenues for research to react to these trends. E. N. Elnozahy, James S. Plank |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2002 | Key Trees and the Security of Interval MulticastabstractA key tree is a distributed data structure of security keys that can be used by a group of users. In this paper we describe how any user in the group can use the different keys in the key tree to securely multicast data to different subgroups within the group. The cost of securely multicasting data to a subgroup whose users are "consecutive" is O(log n) encryptions, where n is the total number of users in the group. The cost of securely multicasting data to an arbitrary subgroup is O(n/2) encryptions. However this cost can be reduced to one encryption by introducing an additional key tree to the group. Mohamed G. Gouda, Chin-Tser Huang, E. N. Elnozahy |
ICDCS | 3 |
| 2002 | Energy-Efficient Duplex and TMR Real-Time SystemsabstractDuplex and triple modular redundancy (TMR) systems are used when a high-level of reliability is desired. Real-time systems for autonomous critical missions need such degrees of reliability, but energy consumption becomes a dominant concern when these systems are built from high-performance processors that consume a large budget of electrical power for operation and cooling. Examples where energy consumption and real time are of paramount importance include reliable computers onboard mobile vehicles, such as the Mars Rover, satellites, and other autonomous vehicles. At first inspection, a duplex system uses about two thirds of the components that a TMR system does, leading one to conclude that duplex systems are more energy-efficient. This paper shows that this is not always the case. We present an analysis of the energy efficiency of duplex and TMR systems when used to tolerate transient failures. With no power management deployed, the analysis supports the intuitive impression about the relative superiority of duplex systems in energy consumption. The analysis shows, however that the gap in energy consumption between the two types of systems diminishes with proper power management. We introduce the concept of an optimistic TMR system that offers the same reliability and performance as the traditional one, but at a fraction of the energy consumption budget. Optimistic TMR systems are competitive with respect to energy consumption when compared with a power-aware duplex system, can even exceed it in some situations, and have the added bonus of providing tolerance to permanent faults. E. N. Elnozahy, Rami G. Melhem, Daniel Mossé |
RTSS | 1 |
| 2002 | Hop integrity in computer networksabstractA computer network is said to provide hop integrity if, when any router, p, in the network receives a message, m, supposedly from an adjacent router, q, then p can check that m was indeed sent by q, was not modified after it was sent and was not a replay of an old message sent from q to p. We describe three protocols that can be added to the routers in a computer network so that the network can provide hop integrity, and thus overcome most denial-of-service attacks. These three protocols are a secret exchange protocol, a weak integrity protocol and a strong integrity protocol. All three protocols are stateless, require small overhead and do not constrain the network protocol in the routers in any way. Mohamed G. Gouda, E. N. Elnozahy, Chin-Tser Huang, Tommy M. McGuire |
IEEE/ACM Trans. Netw. | 2 |
| 2000 | Hop Integrity in Computer NetworksabstractA computer network is said to provide hop integrity iff when any router p in the network receives a message m supposedly from an adjacent router q, then p can check that m was indeed sent by q, was not modified after it was sent, and was not a replay of an old message sent from q to p. We describe three protocols that can be added to the routers in a computer network so that the network can provide hop integrity. These three protocols are a secret exchange protocol, a weak integrity protocol, and a strong integrity protocol. All three protocols are stateless, require small overhead, and do not constrain the network protocol in the routers in any way. Mohamed G. Gouda, E. N. Elnozahy, Chin-Tser Huang, Tommy M. McGuire |
ICNP | 2 |
| 1999 | Address Trace Compression Through Loop Detection and ReductionabstractNo abstract available. E. N. Elnozahy |
SIGMETRICS | 1 |
| 1998 | Support for Software Interrupts in Log-Based Rollback-RecoveryabstractThe piecewise deterministic execution model is a fundamental assumption in many log-based rollback-recovery protocols. Process execution in this model consists of intervals, each starting with the receipt of a message at an application-defined execution point. Execution within each interval is deterministic and messages are the only source of nondeterminism that affects the computation. This simple model excludes the nondeterminism that results when asynchronous signals or interrupts occur at arbitrary execution points. As a result, a wide range of applications cannot use log-based rollback-recovery in practice. We present a solution that removes this restriction and allows applications to replay interrupts at the same execution points during recovery. The solution relies on using a software counter to compute the number of instructions between the asynchronous signals during normal operation. Should a failure occur, the instruction counts are used to force the replay of these signals at the same execution points. The execution of the application thus can be replayed to recreate the prefailure state while accommodating nondeterminism due to asynchronous signals. We then use the deterministic replay of interrupts to solve another problem, namely tracking nondeterminism due to interleaved shared memory access in multithreaded applications on a single processor. We use the instruction counter solution to implement a user-level thread package in which thread scheduling decisions can be replayed if a failure occurs. By repeating the scheduling decisions during an execution replay, threads access the shared memory in the same order and the execution to be reconstructed. This technique allows multithreaded applications to use log-based rollback-recovery with low overhead, which was not previously possible. We carried out two prototype implementations that have shown the overhead is no more than a 6 percent slowdown in application execution on the DEC Alpha, and from 6 percent to 18 percent on the Inter Pentium. Thus, restrictions of the piecewise deterministic execution model can be lifted at a reasonable cost. J. Hamilton Slye, E. N. Elnozahy |
IEEE Trans. Computers | 2 |
| 1995 | On the Relevance of Communication Costs of Rollback-Recovery ProtocolsabstractCommunication overhead has been traditionally the primary metric for evaluating rollback-recovery protocols.This paper reexamines the prominence of this metric in light of the recent increases in processor and network speeds.We introduce a new recovery algorithm for a family of rolibackrecovery protocols based on logging.The new algorithm incurs a higher communication overhead during recovery than previous algorithms, but it requires less access to stable storage and imposes no restrictions on the execution of live processes.Experimental results show that the new algorithm performs better than one that is optimized for low communication overhead.These results suggest that in modern environments, latency in accessing stable storage and intrusion of a particular algorithm on the execution of live processes are more important than the number of messages exchanged during recovery. E. N. Elnozahy |
PODC | 1 |
| 1992 | The Performance of Consistent CheckpointingabstractConsistent checkpointing provides transparent fault tolerance for long-running distributed applications. Performance measurements of an implementation of consistent checkpointing are described. The measurements show that consistent checkpointing performs remarkably well. Eight computation-intensive distributed applications were executed on a network of 16 diskless Sun-3/60 workstations, and the performance without checkpointing was compared to the performance with consistent checkpoints taken at two-minute intervals. For six of the eight applications, the running time increased by less than 1% as a result of the checkpointing. The highest overhead measured was 5.8%. Incremental checkpointing and copy-on write checkpointing were the most effective techniques in lowering the running time overhead. It is argued that these measurements show that consistent checkpointing is an efficient way to provide fault tolerance for long-running distributed applications.> E. N. Elnozahy, David B. Johnson 0001, Willy Zwaenepoel |
SRDS | 1 |
| 1992 | Manetho: Transparent Rollback-Recovery with Low Overhead, Limited Rollback, and Fast Output CommitabstractManetho is a new transparent rollback-recovery protocol for long-running distributed computations. It uses a novel combination of antecedence graph maintenance, uncoordinated checkpointing, and sender-based message logging. Manetho simultaneously achieves the advantages of pessimistic message logging, namely limited rollback and, fast output commit, and the advantage of optimistic message logging, namely low failure-free overhead. These advantages come at the expense of a complex recovery scheme.> E. N. Elnozahy, Willy Zwaenepoel |
IEEE Trans. Computers | 1 |
| 1991 | A comparison of two approaches to build reliable distributed file serversabstractSeveral existing distributed file systems provide reliability by server replication. An alternative approach is to use dual-ported disks accessible to a server and a backup. The two approaches are compared by examining an example of each. Deceit is a replicated file server that emphasizes flexibility. HA-NFS is an example of the second approach that emphasizes efficiency and simplicity. The two file servers run on the same hardware and implement SUN's NFS protocol. The comparison shows that replicated servers are more flexible and tolerant of a wider variety of faults. On the other hand, the dual-ported disks approach is more efficient and simpler to implement. When tolerating single failure, dual-ported disks also give somewhat better availability.> Anupam Bhide, E. N. Elnozahy, Stephen P. Morgan, A. Siegel |
ICDCS | 2 |