EDBT 2026 Demo / reviewers in the wild / expert
Chris Metcalf
dblp:01/6439 · also Christopher D. Metcalf
· DBLP profile ↗
3ranked-venue papers
0as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 67% Debugging and program repair · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis
dynamic analysis |
0.1 | 1 | 2005 | TraceBack: first fault diagnosis by reconstruction of distributed control flow · PLDI 2005 |
Debugging and program repair
fault localization |
0.1 | 1 | 2005 | TraceBack: first fault diagnosis by reconstruction of distributed control flow · PLDI 2005 |
Program analysis › dynamic analysis
runtime instrumentation |
0.1 | 1 | 2005 | TraceBack: first fault diagnosis by reconstruction of distributed control flow · PLDI 2005 |
Methods — techniques the papers use, named apart from their topics
control-flow instrumentation · 0.1binary program analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | TraceBack: first fault diagnosis by reconstruction of distributed control flowabstractFaults that occur in production systems are the most important faults to fix, but most production systems lack the debugging facilities present in development environments. TraceBack provides debugging information for production systems by providing execution history data about program problems (such as crashes, hangs, and exceptions). TraceBack supports features commonly found in production environments such as multiple threads, dynamically loaded modules, multiple source languages (e.g., Java applications running with JNI modules written in C++), and distributed execution across multiple computers. TraceBack supports first fault diagnosis-discovering what went wrong the first time a fault is encountered. The user can see how the program reached the fault state without having to re-run the computation; in effect enabling a limited form of a debugger in production code.TraceBack uses static, binary program analysis to inject low-overhead runtime instrumentation at control-flow block granularity. Post-facto reconstruction of the records written by the instrumentation code produces a source-statement trace for user diagnosis. The trace shows the dynamic instruction sequence leading up to the fault state, even when the program took exceptions or terminated abruptly (e.g., kill -9).We have implemented TraceBack on a variety of architectures and operating systems, and present examples from a variety of platforms. Performance overhead is variable, from 5% for Apache running SPECweb99, to 16%-25% for the Java SPECJbb benchmark, to 60% average for SPECint2000. We show examples of TraceBack's cross-language and cross-machine abilities, and report its use in diagnosing problems in production software. Andrew Ayers, Richard Schooler, Chris Metcalf, Anant Agarwal, Junghwan Rhee, Emmett Witchel |
PLDI | 3 |
| 1996 | NuMesh: An architecture optimized for scheduled communication
David Shoemaker, Frank Honoré, Chris Metcalf, Steve Ward |
J. Supercomput. | 3 |
| 1993 | The NuMesh: A Modular, Scalable Communications SubstrateabstractMany standardized hardware communication interfaces offer runtime flexibility and configurability at the cost of efficiency.An alternate approach is the use of a highly-efficien~minimal communication element with as much communication decision-making as possible done at compile time.NuMesh is a packaging and interconnect technology supporting high-bandwidth systolic communications on a 3D nearest-neighbor lattice; our goal is to combine Lego-like modularity with supercomputer performance.To date, the primary focus of the project has been the class of applications whose static communication patterns can be precompiled into independent and carefully choreographed finite state machines running on each node.Several extensions of the NuMesh to more general communication paradigms have been implemented, and the issues involved are under active exploration.This paper presents an overview of our approach, as well as an introduction to our current-generation prototype.We also discuss our software environment and simulation technology, and enumerate some of the applications and programming models we have developed to make full use of the capabdities of the NuMesh. Steve Ward, Karim Abdalla, Rajeev Dujari, Michael Fetterman, Frank Honoré, Ricardo Jenez, Philippe Laffont, Kenneth Mackenzie, Chris Metcalf, Milan Minsky, John Nguyen, John Pezaris, Gill A. Pratt, Russell Tessier |
International Conference on Supercomputing | 9 |