Robert H. B. Netzer

dblp:01/1090 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 9 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorTheory of computation · 3Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Distributed systems · 52% Parallel and multicore computing · 25% Electronic design automation · 11%
Software engineering, system software, and programming languages
4 papers
Concurrent programming · 48% Debugging and program repair · 34% Operating systems · 15%
Theoretical computer science
3 papers
Distributed computing theory · 92% Logic in computer science · 8%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.021999
Consistency Issues in Distributed Checkpoints · IEEE Trans. Software Eng. 1999
Finding Consistent Global Checkpoints in a Distributed Computation · IEEE Trans. Parallel Distributed Syst. 1997
Distributed systems › fault tolerance
checkpointing
0.031999
Consistency Issues in Distributed Checkpoints · IEEE Trans. Software Eng. 1999
Replaying Distributed Programs without Message Logging · HPDC 1997
Necessary and Sufficient Conditions for Consistent Global Snapshots · IEEE Trans. Parallel Distributed Syst. 1995
Concurrent programming › concurrency bug detection
data race detection
0.031991
Techniques for Debugging Parallel Programs with Flowback Analysis · ACM Trans. Program. Lang. Syst. 1991
Improving the Accuracy of Data Race Detection · PPoPP 1991
Detecting Data Races on Weak Memory Systems · ISCA 1991
Parallel and multicore computing › parallel computing
parallel program debugging
0.031993
Adaptive message logging for incremental replay of message-passing programs · SC 1993
Optimal Tracing and Replay for Debugging Message-Passing Parallel Programs · SC 1992
Techniques for Debugging Parallel Programs with Flowback Analysis · ACM Trans. Program. Lang. Syst. 1991
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay
0.021993
Adaptive message logging for incremental replay of message-passing programs · SC 1993
Optimal Tracing and Replay for Debugging Message-Passing Parallel Programs · SC 1992
Hardware reliability and fault tolerance › error recovery
checkpoint recovery
0.011997
Finding Consistent Global Checkpoints in a Distributed Computation · IEEE Trans. Parallel Distributed Syst. 1997
Electronic design automation › hardware verification and test
debugging
0.011997
Replaying Distributed Programs without Message Logging · HPDC 1997
Distributed computing theory
checkpointing
0.011997
Finding Consistent Global Checkpoints in a Distributed Computation · IEEE Trans. Parallel Distributed Syst. 1997
Debugging and program repair
fault localization
0.021991
Techniques for Debugging Parallel Programs with Flowback Analysis · ACM Trans. Program. Lang. Syst. 1991
Improving the Accuracy of Data Race Detection · PPoPP 1991
Distributed computing theory › distributed algorithms
distributed snapshots
0.011995
Necessary and Sufficient Conditions for Consistent Global Snapshots · IEEE Trans. Parallel Distributed Syst. 1995
Operating systems › fault tolerance
checkpoint and rollback
0.011994
Optimal Tracing and Incremental Reexecution for Debugging Long-Running Programs · PLDI 1994
Distributed systems › fault tolerance
message logging
0.011993
Adaptive message logging for incremental replay of message-passing programs · SC 1993
Concurrent programming
concurrency bugs
0.011991
Techniques for Debugging Parallel Programs with Flowback Analysis · ACM Trans. Program. Lang. Syst. 1991
Concurrent programming › concurrency bug detection › data race detection
dynamic race detection
0.011991
Detecting Data Races on Weak Memory Systems · ISCA 1991
Distributed computing theory
distributed algorithms
0.011997
Finding Consistent Global Checkpoints in a Distributed Computation · IEEE Trans. Parallel Distributed Syst. 1997
Logic in computer science
causality
0.011995
Necessary and Sufficient Conditions for Consistent Global Snapshots · IEEE Trans. Parallel Distributed Syst. 1995
Distributed computing theory › event ordering
happened-before relation
0.011995
Necessary and Sufficient Conditions for Consistent Global Snapshots · IEEE Trans. Parallel Distributed Syst. 1995
Program analysis
dynamic analysis
0.011991
Improving the Accuracy of Data Race Detection · PPoPP 1991

Methods — techniques the papers use, named apart from their topics

necessary and sufficient condition proof · 0.0proof · 0.0enumeration algorithm · 0.0zigzag path analysis · 0.0semantic analysis · 0.0program dependence graph · 0.0incremental tracing · 0.0two-level bitvectors · 0.0adaptive tracing · 0.0adaptive message logging · 0.0runtime tracing · 0.0message logging · 0.0race validation · 0.0race ordering · 0.0
YearPublicationVenuePosition
2003 Detecting Race Conditions in Parallel Programs that Use Semaphores
Philip N. Klein, Robert H. B. Netzer, Hsueh-I Lu
Algorithmica2
2001 Deadlock-Free Incremental Replay of Message-Passing Programs
Franco Zambonelli, Robert H. B. Netzer
J. Parallel Distributed Comput.2
2000 Communication-Based Prevention of Useless Checkpoints in Fistributed Computations
Jean-Michel Hélary, Achour Mostéfaoui, Robert H. B. Netzer, Michel Raynal
Distributed Comput.3
1999 Consistency Issues in Distributed Checkpoints
abstract
A global checkpoint is a set of local checkpoints, one per process. The traditional consistency criterion for global checkpoints states that a global checkpoint is consistent if it does not include messages received and not sent. The paper investigates other consistency criteria, transitlessness, and strong consistency. A global checkpoint is transitless if it does not exhibit messages sent and not received. Transitlessness can be seen as a dual of traditional consistency. Strong consistency is the addition of transitlessness to traditional consistency. The main result of the paper is a statement of the necessary and sufficient condition answering the following question: "given an arbitrary set of local checkpoints, can this set be extended to a global checkpoint that satisfies P" (where P is traditional consistency, transitlessness, or strong consistency). From a practical point of view, this condition, when applied to transitlessness, is particularly interesting as it helps characterize which messages do not need to be recorded by checkpointing protocols.
Jean-Michel Hélary, Robert H. B. Netzer, Michel Raynal
IEEE Trans. Software Eng.2
1997 Replaying Distributed Programs without Message Logging
abstract
Debugging long program runs can be difficult because of the delays required to repeatedly re-run the execution. Even a moderately long run of five minutes can incur aggravating delays. To address this problem, techniques exist that allow re-executing a distributed program from intermediate points by using combinations of checkpointing and message logging. In this paper we explore another idea: how to support replay without logging the contents of any message. When no messages are logged, the set of global states from which replay is possible is constrained, and it has been unknown how to compute this set without exhaustively searching the space of all global states, whose size is exponential in the number of processes. We present a simple and efficient hybrid on-the-fly/post-mortem algorithm for detecting the necessary and sufficient conditions under which parts of the execution can be replayed without message logs. A small amount of trace (two vectors) is recorded at each checkpoint and a fast post-mortem algorithm computes global states from which replay can begin. This algorithm is independent of the checkpointing technique used.
Robert H. B. Netzer, Yikang Xu
HPDC1
1997 Preventing Useless Checkpoints in Distributed Computations
abstract
A useless checkpoint is a local checkpoint that cannot be part of a consistent global checkpoint. The paper addresses the following important problem. Given a set of processes that take (basic) local checkpoints in an independent and unknown way, the problem is to design a communication induced checkpointing protocol that directs processes to take additional local (forced) checkpoints to ensure that no local checkpoint is useless. A general and efficient protocol answering this problem is proposed. It is shown that several existing protocols that solve the same problem are particular instances of it. The design of this general protocol is motivated by the use of communication induced checkpointing protocols in "consistent global checkpoint" based distributed applications. Detection of stable or unstable properties, rollback recovery and determination of distributed breakpoints are examples of such applications.
Jean-Michel Hélary, Achour Mostéfaoui, Robert H. B. Netzer, Michel Raynal
SRDS3
1997 Finding Consistent Global Checkpoints in a Distributed Computation
abstract
Consistent global checkpoints have many uses in distributed computations. A central question in applications that use consistent global checkpoints is to determine whether a consistent global checkpoint that includes a given set of local checkpoints can exist. Netzer and Xu (1995) presented the necessary and sufficient conditions under which such a consistent global checkpoint can exist, but they did not explore what checkpoints could be constructed. In this paper, we prove exactly which local checkpoints can be used for constructing such consistent global checkpoints. We illustrate the use of our results with a simple and elegant algorithm to enumerate all such consistent global checkpoints.
D. Manivannan 0001, Robert H. B. Netzer, Mukesh Singhal
IEEE Trans. Parallel Distributed Syst.2
1996 Race-Condition Detection in Parallel Computation with Semaphores (Extended Abstract)
Philip N. Klein, Hsueh-I Lu, Robert H. B. Netzer
ESA3
1995 Optimal tracing and replay for debugging message-passing parallel programs
Robert H. B. Netzer, Barton P. Miller
J. Supercomput.1
1995 Necessary and Sufficient Conditions for Consistent Global Snapshots
abstract
Consistent global snapshots are important in many distributed applications. We prove the exact conditions for an arbitrary checkpoint, or a set of checkpoints, to belong to a consistent global snapshot, a previously open problem. To describe the conditions, we introduce a generalization of Lamport's (1978) happened-before relation called a zigzag path.>
Robert H. B. Netzer
IEEE Trans. Parallel Distributed Syst.1
1994 Critical-Path-Based Message Logging for incremental Replay of Message-Passing Programs
abstract
Debugging long-running, nondeterministic message-passing parallel programs requires incremental replay, the ability to exactly replay selected parts of an execution. To support incremental replay, we must log enough messages and checkpoint processes often enough to allow any requested replay to complete quickly. We present an adaptive tracing strategy to keep the message-logging overhead down. We let the user specify a bound on the maximum time any replay request is allowed to take. Our algorithm tracks what each process's critical path will be during a replay and logs enough messages to ensure the critical path will never exceed the bound. Overhead is kept low by not logging messages that can be recomputed during a replay. Experiments indicate that we log about 0.1-5% of the messages while still providing a reasonable bound on any replay.>
Robert H. B. Netzer, Sairam Subramanian
ICDCS1
1994 Optimal Tracing and Incremental Reexecution for Debugging Long-Running Programs
abstract
Debugging requires execution replay. Locations of bugs are rarely known in advance, so an execution must be repeated over and over to track down bugs. A problem arises with repeated reexecution for long-running programs and programs that have complex interactions with their environment. Replaying long-running programs from the start incurs too much delay. Replaying programs that interact with their environment requires the difficult (and sometimes impossible) task of exactly reproducing this environment (such as the connections over a one-day period to an X server). We solve these problems by incremental checkpointing and replay. By periodically checkpointing parts of the execution''s state, it can be restarted from intermediate points, bounding the delay to replay any part of the execution and allowing parts of the execution to be skipped. We present adaptive tracing strategies that provide bounded-time incremental replay and that are nearly optimal. Our techniques track reads and writes to memory using space-efficient two-level bitvectors. Our implementation on a Sparc 10 traces less than 15 kilobytes/sec for CPU-intensive programs and for interactive programs the slowdown is low enough that tracing can be left on all the time.
Robert H. B. Netzer, Mark H. Weaver
PLDI1
1993 Adaptive message logging for incremental replay of message-passing programs
abstract
Article Free Access Share on Adaptive message logging for incremental replay of message-passing programs Authors: R. H. B. Netzer Dept. of Computer Science, Brown University, Box 1910, Providence, RI Dept. of Computer Science, Brown University, Box 1910, Providence, RIView Profile , J. Xu Dept. of Computer Science, Brown University, Box 1910, Providence, RI Dept. of Computer Science, Brown University, Box 1910, Providence, RIView Profile Authors Info & Claims Supercomputing '93: Proceedings of the 1993 ACM/IEEE conference on SupercomputingDecember 1993 Pages 840–849https://doi.org/10.1145/169627.169850Published:01 December 1993Publication History 16citation222DownloadsMetricsTotal Citations16Total Downloads222Last 12 Months9Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Robert H. B. Netzer
SC1
1993 Detecting Race Conditions in Parallel Programs that Use One Semaphore
Hsueh-I Lu, Philip N. Klein, Robert H. B. Netzer
WADS3
1992 Efficient Race Condition Detection for Shared-Memory Programs with Post/Wait Synchronization
Robert H. B. Netzer, Sanjoy Ghosh
ICPP (2)1
1992 Optimal Tracing and Replay for Debugging Message-Passing Parallel Programs
abstract
A techinque for tracing and replaying message-passing programs for debugging is presented. The technique is optimal in the common case and has good performance in the worst case. By making runtime tracing decisions, only a fraction of the total number of messages is traced, gaining two orders of magnitude reduction over traditional techniques which trace every message. Experiments indicate that only 1% of the messages often need to be traced. These traces are sufficient to provide replay, allowing an execution to be reproduced any number of times for debugging. This work is novel in that runtime decisions are used to detect and trace only those messages that introduce nondeterminacy. With the proposed strategy, large reductions in trace size allow long-running programs to be replayed that were previously unmanageable. In addition, the reduced tracing experiments alleviate tracing bottlenecks, allowing executions to be debugged with substantially lower execution-time overhead.>
Robert H. B. Netzer, Barton P. Miller
SC1
1991 Detecting Data Races on Weak Memory Systems
abstract
For shared-memory systems, the most commonly assumed programmer's model of memory is sequential consistency. The weaker models of weak ordering, release consistency with sequentially consistent synchronization operations, data-race-free-0, and data-race-free-1 provide higher performance by guaranteeing sequential consistency to only a restricted class of programs - mainly programs that do not exhibit data races. To allow programmers to use the intuition and algorithms already developed for sequentially consistent systems, it is important to determine when a program written for a weak system exhibits no data races. In this paper, we investigate the extension of dynamic data race detection techniques developed for sequentially consistent systems to weak systems. A potential problem is that in the presence of a data race, weak systems fail to guarantee sequential consistency and therefore dynamic techniques may not give meaningful results. However, we reason that in practice a weak system...
Sarita V. Adve, Mark D. Hill, Barton P. Miller, Robert H. B. Netzer
ISCA4
1991 Improving the Accuracy of Data Race Detection
abstract
For shared-memory parallel programs that use explicit synchronization, data race detection is an important part of debugging. A data race exists when concurrently executing sections of code access common shared variables. In programs intended to be data race free, they are sources of nondeterminism usually considered bugs. Previous methods for detecting data races in executions of parallel programs can determine when races occurred, but can report many data races that are artifacts of others and not direct manifestations of program bugs. Artifacts exist because some races can cause others and can also make false races appear real. Such artifacts can overwhelm the programmer with information irrelevant for debugging. This paper presents results showing how to identify nonartifact data races by validation and ordering. Data race validation attempts to determine which races involve events that either did execute concurrently or could have (called feasible data races). We show how each de...
Robert H. B. Netzer, Barton P. Miller
PPoPP1
1991 Techniques for Debugging Parallel Programs with Flowback Analysis
abstract
Flowback analysis is a powerful technique for debugging programs.It allows the programmer to examine dynamic dependence in a program's execution history without having to reexecute the program.The goal is to present to the programmer a graphical view of the dynamic program dependence.We are building a system, called PPD, that performs flowback analysis while keeping the execution time overhead low.We also extend the semantics of flowback analysis to parallel programs.This paper describes details of the graphs and algorithms needed to implement efficient flowback analysis for parallel programs.Execution-time overhead is kept low by recording only a small amount of trace during a program's execution.We use semantic analysis and a technique called incremental tracing to keep the time and space overhead low.As part of the semantic analysis, PPD uses a static program dependence graph structure that reduces the amount of work done at compile time and takes advantage of the dynamic information produced during execution time.Parallel programs have been accommodated in two ways.First, the flowback dependence can span process boundaries; that is, the most recent modification to a variable might be traced to a different process than that one that contains the current reference.The static dynamic program dependence graphs of the individual processes are tied together with synchronization and data dependence information to form complete graphs that represent the entire program.Second, our algorithms will detect potential data-race conditions in the access to shared variables.The programmer can be directed to the cause of the race condition.PPD is currently being implemented for the C programming language on a Sequent Symmetry shared-memory .
Jong-Deok Choi, Barton P. Miller, Robert H. B. Netzer
ACM Trans. Program. Lang. Syst.3
1990 On the Complexity of Event Ordering for Shared-Memory Parallel Program Executions
Robert H. B. Netzer, Barton P. Miller
ICPP (2)1