Bob Moench

dblp:48/8524 · also Robert Moench · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 60% Electronic design automation · 20% High-performance computing · 12%
Software engineering, system software, and programming languages
3 papers
Debugging and program repair · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair › software debugging
relative debugging
0.322015
Relative debugging for a highly parallel hybrid computer system · SC 2015
Data centric highly parallel debugging · HPDC 2010
GPUs and heterogeneous computing
GPU computing
0.212015
Relative debugging for a highly parallel hybrid computer system · SC 2015
GPUs and heterogeneous computing › heterogeneous computing systems
hybrid computer
0.212015
Relative debugging for a highly parallel hybrid computer system · SC 2015
Electronic design automation › hardware verification and test
debugging
0.112012
Scalable parallel debugging with statistical assertions · PPoPP 2012
Debugging and program repair › automated debugging
assertion-based debugging
0.112010
Data centric highly parallel debugging · HPDC 2010
Debugging and program repair › concurrent program debugging
parallel program debugging
0.112010
Data centric highly parallel debugging · HPDC 2010
Parallel and multicore computing
parallel programming models
0.112015
Relative debugging for a highly parallel hybrid computer system · SC 2015
High-performance computing › scientific computing systems
molecular dynamics simulation
0.012012
Scalable parallel debugging with statistical assertions · PPoPP 2012
High-performance computing
scientific computing
0.012012
Scalable parallel debugging with statistical assertions · PPoPP 2012

Methods — techniques the papers use, named apart from their topics

data model · 0.4case study · 0.4statistical assertion · 0.3parallelization · 0.3hashing · 0.1assertions · 0.1
YearPublicationVenuePosition
2015 Relative debugging for a highly parallel hybrid computer system
abstract
Relative debugging traces software errors by comparing two executions of a program concurrently - one code being a reference version and the other faulty. Relative debugging is particularly effective when code is migrated from one platform to another, and this is of significant interest for hybrid computer architectures containing CPUs accelerators or coprocessors. In this paper we extend relative debugging to support porting stencil computation on a hybrid computer. We describe a generic data model that allows programmers to examine the global state across different types of applications, including MPI/OpenMP, MPI/OpenACC, and UPC programs. We present case studies using a hybrid version of the `stellarator' particle simulation DELTA5D, on Titan at ORNL, and the UPC version of Shallow Water Equations on Crystal, an internal supercomputer of Cray. These case studies used up to 5,120 GPUs and 32,768 CPU cores to illustrate that the debugger is effective and practical.
Luiz De Rose, Andrew Gontarek, Aaron Vose, Bob Moench, David Abramson 0001, Minh Ngoc Dinh, Chao Jin 0001
SC4
2015 A data-centric framework for debugging highly parallel applications
abstract
Summary Contemporary parallel debuggers allow users to control more than one processing thread while supporting the same examination and visualisation operations of that of sequential debuggers. This approach restricts the use of parallel debuggers when it comes to large scale scientific applications run across hundreds of thousands compute cores. First, manually observing the runtime data to detect error becomes impractical because the data is too big. Second, performing expensive but useful debugging operations becomes infeasible as the computational codes become more complex, involving larger data structures, and as the machines become larger. This study explores the idea of a data‐centric debugging approach, which could be used to make parallel debuggers more powerful. It discusses the use ofad hocdebug‐time assertions that allow a user to reason about the state of a parallel computation. These assertions support the verification and validation of program state at runtime as a whole rather than focusing on that of only a single process state. Furthermore, the debugger's performance can be improved by exploiting the underlying parallel platform because the available compute cores can execute parallel debugging functions, while a program is idling at a breakpoint. We demonstrate the system with several case studies and evaluate the performance of the tool on a 20 000 cores Cray XE6. Copyright © 2013 John Wiley & Sons, Ltd.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
Softw. Pract. Exp.5
2012 A Scalable Parallel Debugging Library with Pluggable Communication Protocols
abstract
Parallel debugging faces challenges in both scalability and efficiency. A number of advanced methods have been invented to improve the efficiency of parallel debugging. As the scale of system increases, these methods highly rely on a scalable communication protocol in order to be utilized in large-scale distributed environments. This paper describes a debugging middleware that provides fundamental debugging functions supporting multiple communication protocols. Its pluggable architecture allows users to select proper communication protocols as plug-ins for debugging on different platforms. It aims to be utilized by various advanced debugging technologies across different computing platforms. The performance of this debugging middleware is examined on a Cray XE Supercomputer with 21,760 CPU cores.
Chao Jin 0001, David Abramson 0001, Minh Ngoc Dinh, Andrew Gontarek, Bob Moench, Luiz De Rose
CCGRID5
2012 Scalable parallel debugging with statistical assertions
abstract
Traditional debuggers are of limited value for modern scientific codes that manipulate large complex data structures. This paper discusses a novel debug-time assertion, called a "Statistical Assertion", that allows a user to reason about large data structures, and the primitives are parallelised to provide an efficient solution. We present the design and implementation of statistical assertions, and illustrate the debugging technique with a molecular dynamics simulation. We evaluate the performance of the tool on a 12,000 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Chao Jin 0001, Andrew Gontarek, Bob Moench, Luiz De Rose
PPoPP5
2011 Assertion Based Parallel Debugging
abstract
Programming languages have advanced tremendously over the years, but program debuggers have hardly changed. Sequential debuggers do little more than allow a user to control the flow of a program and examine its state. Parallel ones support the same operations on multiple processes, which are adequate with a small number of processors, but become unwieldy and ineffective on very large machines. Typical scientific codes have enormous multi-dimensional data structures and it is impractical to expect a user to view the data using traditional display techniques. In this paper we discuss the use of debug-time assertions, and show that these can be used to debug parallel programs. The techniques reduce the debugging complexity because they reason about the state of large arrays without requiring the user to know the expected value of every element. Assertions can be expensive to evaluate, but their performance can be improved by running them in parallel. We demonstrate the system with a case study finding errors in a parallel version of the Shallow Water Equations, and evaluate the performance of the tool on a 4,096 cores Cray XE6.
Minh Ngoc Dinh, David Abramson 0001, Donny Kurniawan, Chao Jin 0001, Bob Moench, Luiz De Rose
CCGRID5
2010 Data centric highly parallel debugging
abstract
Debugging parallel programs is an order of magnitude more complex than sequential ones, and yet, most parallel debuggers provide little extra functionality than their sequential counterparts. This problem becomes more serious as computational codes become more complex, involving larger data structures, and as the machines become larger. Peta-scale machines consisting of millions of cores pose a significant challenge for existing techniques. We argue that debugging must become more data-centric, and believe that "assertions" provide a useful model. Assertions allow a user to declare their expectations about the program state as a whole rather than focusing on that of only a single process state. Previously, we have implemented a special type of assertion that supports debugging applications as they evolve or are ported to different platforms. They allow a user to compare the state of one program against another reference version. These 'relative debugging' assertions, whilst powerful, pose significant implementation challenges for large peta-scale machines. In this paper we discuss a hashing technique that provides a scalable solution for very large problems on very large machines. We illustrate the scheme on 65k cores of Kraken, a Cray XT5 at the University of Tennessee.
David Abramson 0001, Minh Ngoc Dinh, Donny Kurniawan, Bob Moench, Luiz De Rose
HPDC4