EDBT 2026 Demo / reviewers in the wild / expert
Bob Moench
dblp:48/8524 · also Robert Moench
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 60% Electronic design automation · 20% High-performance computing · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Debugging and program repair · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Debugging and program repair › software debugging
relative debugging |
0.3 | 2 | 2015 | Relative debugging for a highly parallel hybrid computer system · SC 2015 Data centric highly parallel debugging · HPDC 2010 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2015 | Relative debugging for a highly parallel hybrid computer system · SC 2015 |
GPUs and heterogeneous computing › heterogeneous computing systems
hybrid computer |
0.2 | 1 | 2015 | Relative debugging for a highly parallel hybrid computer system · SC 2015 |
Electronic design automation › hardware verification and test
debugging |
0.1 | 1 | 2012 | Scalable parallel debugging with statistical assertions · PPoPP 2012 |
Debugging and program repair › automated debugging
assertion-based debugging |
0.1 | 1 | 2010 | Data centric highly parallel debugging · HPDC 2010 |
Debugging and program repair › concurrent program debugging
parallel program debugging |
0.1 | 1 | 2010 | Data centric highly parallel debugging · HPDC 2010 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2015 | Relative debugging for a highly parallel hybrid computer system · SC 2015 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.0 | 1 | 2012 | Scalable parallel debugging with statistical assertions · PPoPP 2012 |
High-performance computing
scientific computing |
0.0 | 1 | 2012 | Scalable parallel debugging with statistical assertions · PPoPP 2012 |
Methods — techniques the papers use, named apart from their topics
data model · 0.4case study · 0.4statistical assertion · 0.3parallelization · 0.3hashing · 0.1assertions · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Relative debugging for a highly parallel hybrid computer systemabstractRelative debugging traces software errors by comparing two executions of a program concurrently - one code being a reference version and the other faulty. Relative debugging is particularly effective when code is migrated from one platform to another, and this is of significant interest for hybrid computer architectures containing CPUs accelerators or coprocessors. In this paper we extend relative debugging to support porting stencil computation on a hybrid computer. We describe a generic data model that allows programmers to examine the global state across different types of applications, including MPI/OpenMP, MPI/OpenACC, and UPC programs. We present case studies using a hybrid version of the `stellarator' particle simulation DELTA5D, on Titan at ORNL, and the UPC version of Shallow Water Equations on Crystal, an internal supercomputer of Cray. These case studies used up to 5,120 GPUs and 32,768 CPU cores to illustrate that the debugger is effective and practical. Luiz De Rose, Andrew Gontarek, Aaron Vose, Bob Moench, David Abramson 0001, Minh Ngoc Dinh, Chao Jin 0001 |
SC | 4 |
| 2015 | A data-centric framework for debugging highly parallel applicationsabstractSummary Contemporary parallel debuggers allow users to control more than one processing thread while supporting the same examination and visualisation operations of that of sequential debuggers. This approach restricts the use of parallel debuggers when it comes to large scale scientific applications run across hundreds of thousands compute cores. First, manually observing the runtime data to detect error becomes impractical because the data is too big. Second, performing expensive but useful debugging operations becomes infeasible as the computational codes become more complex, involving larger data structures, and as the machines become larger. This study explores the idea of a data‐centric debugging approach, which could be used to make parallel debuggers more powerful. It discusses the use ofad hocdebug‐time assertions that allow a user to reason about the state of a parallel computation. These assertions support the verification and validation of program state at runtime as a whole rather than focusing on that of only a single process state. Furthermore, the debugger's performance can be improved by exploiting the underlying parallel platform because the available compute cores can execute parallel debugging functions, while a program is idling at a breakpoint. We demonstrate the system with several case studies and evaluate the performance of the tool on a 20 000 cores Cray XE6. Copyright © 2013 John Wiley & Sons, Ltd. Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose |
Softw. Pract. Exp. | 5 |
| 2012 | A Scalable Parallel Debugging Library with Pluggable Communication ProtocolsabstractParallel debugging faces challenges in both scalability and efficiency. A number of advanced methods have been invented to improve the efficiency of parallel debugging. As the scale of system increases, these methods highly rely on a scalable communication protocol in order to be utilized in large-scale distributed environments. This paper describes a debugging middleware that provides fundamental debugging functions supporting multiple communication protocols. Its pluggable architecture allows users to select proper communication protocols as plug-ins for debugging on different platforms. It aims to be utilized by various advanced debugging technologies across different computing platforms. The performance of this debugging middleware is examined on a Cray XE Supercomputer with 21,760 CPU cores. Chao Jin 0001, David Abramson 0001, Minh Ngoc Dinh, Andrew Gontarek, Bob Moench, Luiz De Rose |
CCGRID | 5 |
| 2012 | Scalable parallel debugging with statistical assertionsabstractTraditional debuggers are of limited value for modern scientific codes that manipulate large complex data structures. This paper discusses a novel debug-time assertion, called a "Statistical Assertion", that allows a user to reason about large data structures, and the primitives are parallelised to provide an efficient solution. We present the design and implementation of statistical assertions, and illustrate the debugging technique with a molecular dynamics simulation. We evaluate the performance of the tool on a 12,000 cores Cray XE6. Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose |
PPoPP | 5 |
| 2011 | Assertion Based Parallel DebuggingabstractProgramming languages have advanced tremendously over the years, but program debuggers have hardly changed. Sequential debuggers do little more than allow a user to control the flow of a program and examine its state. Parallel ones support the same operations on multiple processes, which are adequate with a small number of processors, but become unwieldy and ineffective on very large machines. Typical scientific codes have enormous multi-dimensional data structures and it is impractical to expect a user to view the data using traditional display techniques. In this paper we discuss the use of debug-time assertions, and show that these can be used to debug parallel programs. The techniques reduce the debugging complexity because they reason about the state of large arrays without requiring the user to know the expected value of every element. Assertions can be expensive to evaluate, but their performance can be improved by running them in parallel. We demonstrate the system with a case study finding errors in a parallel version of the Shallow Water Equations, and evaluate the performance of the tool on a 4,096 cores Cray XE6. Minh Ngoc Dinh, David Abramson 0001, Donny Kurniawan, Chao Jin 0001, Bob Moench, Luiz De Rose |
CCGRID | 5 |
| 2010 | Data centric highly parallel debuggingabstractDebugging parallel programs is an order of magnitude more complex than sequential ones, and yet, most parallel debuggers provide little extra functionality than their sequential counterparts. This problem becomes more serious as computational codes become more complex, involving larger data structures, and as the machines become larger. Peta-scale machines consisting of millions of cores pose a significant challenge for existing techniques. We argue that debugging must become more data-centric, and believe that "assertions" provide a useful model. Assertions allow a user to declare their expectations about the program state as a whole rather than focusing on that of only a single process state. Previously, we have implemented a special type of assertion that supports debugging applications as they evolve or are ported to different platforms. They allow a user to compare the state of one program against another reference version. These 'relative debugging' assertions, whilst powerful, pose significant implementation challenges for large peta-scale machines. In this paper we discuss a hashing technique that provides a scalable solution for very large problems on very large machines. We illustrate the scheme on 65k cores of Kraken, a Cray XT5 at the University of Tennessee. David Abramson 0001, Minh Ngoc Dinh, Donny Kurniawan, Bob Moench, Luiz De Rose |
HPDC | 4 |