VLDB 2026 Research / reviewers in the wild / expert
Paul Stodghill
dblp:67/4594
· DBLP profile ↗
17ranked-venue papers
0as first author
0since 2021 · last 2006
0000-0003-3875-8450ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Distributed systems · 33% High-performance computing · 24% Performance modeling and evaluation · 14% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 90% Program analysis · 10% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › fault tolerance › checkpointing
application-level checkpointing |
0.1 | 2 | 2004 | Implementation and Evaluation of a Scalable Application-Level Checkpoint-Recovery Scheme for MPI Programs · SC 2004 Automated application-level checkpointing of MPI programs · PPoPP 2003 |
Distributed systems
fault tolerance |
0.1 | 2 | 2004 | Implementation and Evaluation of a Scalable Application-Level Checkpoint-Recovery Scheme for MPI Programs · SC 2004 Automated application-level checkpointing of MPI programs · PPoPP 2003 |
Parallel and multicore computing › parallel programming models › message passing
MPI applications |
0.1 | 2 | 2006 | Mobile MPI programs in computational grids · PPoPP 2006 Implementation and Evaluation of a Scalable Application-Level Checkpoint-Recovery Scheme for MPI Programs · SC 2004 |
Performance modeling and evaluation
benchmarking |
0.1 | 2 | 2005 | Automatic measurement of memory hierarchy parameters · SIGMETRICS 2005 Is Search Really Necessary to Generate High-Performance BLAS? · Proc. IEEE 2005 |
Distributed systems
grid computing |
0.1 | 1 | 2006 | Mobile MPI programs in computational grids · PPoPP 2006 |
Cloud and datacenter computing
utility computing |
0.1 | 1 | 2006 | Mobile MPI programs in computational grids · PPoPP 2006 |
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2005 | Is Search Really Necessary to Generate High-Performance BLAS? · Proc. IEEE 2005 |
Memory systems
memory hierarchy |
0.1 | 1 | 2005 | Automatic measurement of memory hierarchy parameters · SIGMETRICS 2005 |
Performance modeling and evaluation › benchmarking
microbenchmarking |
0.1 | 1 | 2005 | Automatic measurement of memory hierarchy parameters · SIGMETRICS 2005 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2005 | Is Search Really Necessary to Generate High-Performance BLAS? · Proc. IEEE 2005 |
Distributed systems › fault tolerance
checkpointing |
0.0 | 1 | 2004 | Implementation and Evaluation of a Scalable Application-Level Checkpoint-Recovery Scheme for MPI Programs · SC 2004 |
High-performance computing
numerical linear algebra |
0.0 | 1 | 2003 | A comparison of empirical and model-driven optimization · PLDI 2003 |
Compilers and program optimization › sparse computation › sparse tensor compilation
sparse tensor code generation |
0.0 | 1 | 2000 | A Framework for Sparse Matrix Code Synthesis from High-level Specifications · SC 2000 |
High-performance computing › iterative methods
conjugate gradient |
0.0 | 1 | 2000 | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path · SC 2000 |
Processor architecture and microarchitecture › multithreading
fine-grain multithreading |
0.0 | 1 | 2000 | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path · SC 2000 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2000 | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path · SC 2000 |
Compilers and program optimization › code generation
parallel code generation |
0.0 | 1 | 1997 | Compiling Parallel Code for Sparse Matrix Applications · SC 1997 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.0 | 1 | 2005 | Automatic measurement of memory hierarchy parameters · SIGMETRICS 2005 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2000 | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary Path · SC 2000 |
High-performance computing
sparse linear algebra |
0.0 | 1 | 2000 | A Framework for Sparse Matrix Code Synthesis from High-level Specifications · SC 2000 |
Program analysis › data flow analysis
constant propagation |
0.0 | 1 | 1991 | Dependence Flow Graphs: An Algebraic Approach to Program Dependencies · POPL 1991 |
Program analysis
data flow analysis |
0.0 | 1 | 1991 | Dependence Flow Graphs: An Algebraic Approach to Program Dependencies · POPL 1991 |
Compilers and program optimization
intermediate representation |
0.0 | 1 | 1991 | Dependence Flow Graphs: An Algebraic Approach to Program Dependencies · POPL 1991 |
Methods — techniques the papers use, named apart from their topics
global search · 0.1analytical performance modeling · 0.1model-driven optimization · 0.1empirical optimization · 0.1application-level checkpointing · 0.1microbenchmarking · 0.1program transformation · 0.0preprocessor · 0.0precompiler instrumentation · 0.0checkpointing protocol · 0.0common enumeration identification · 0.0cartesian product embedding · 0.0relational algebra · 0.0static single assignment · 0.0abstract interpretation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2006 | Experimental evaluation of application-level checkpointing for OpenMP programsabstractIt is becoming important for long-running scientific applications to tolerate hardware faults. The most commonly used approach is checkpoint and restart (CPR) - the computation's state is saved periodically to disk. Upon failure the computation is restarted from the last saved state. The common CPR mechanism, called System-level Checkpointing (SLC), requires modifying the Operating System and the communication libraries to enable them to save the state of the entire parallel application. This approach is not portable since a checkpointer for one system rarely works on another. Application-level Checkpointing (ALC) is a portable alternative where the programmer manually modifies their program to enable CPR, a very labor-intensive task.We are investigating the use of compiler technology to instrument codes to embed the ability to tolerate faults into applications themselves, making them self-checkpointing and self-restarting on any platform. In [9] we described a general approach for checkpointing shared memory APIs at the application level. Since [9] applied to only a toy feature set common to most shared memory APIs, this paper shows the practicality of this approach by extending it to a specific popular shared memory API: OpenMP. We describe the challenges involved in providing automated ALC for OpenMP applications and experimentally validate this approach by showing detailed performance results for our implementation of this technique. Our experiments with the NAS OpenMP benchmarks [1] and the EPCC microbench-marks [21] show generally low overhead on three different architectures: Linux/IA64, Tru64/Alpha and Solaris/Sparc and highlight important lessons about the performance characteristics of this aproach. Greg Bronevetsky, Keshav Pingali, Paul Stodghill |
ICS | 3 |
| 2006 | A distributed system based on web services for computational science simulationsabstractIn this paper, we describe the ASP system, a testbed based on Web Services for coupled multi-physics simulations. The system is organized as a collection of geographically-distributed software components in which each component provides a Web Service and uses standard SOAP-based Web Service protocols to interact with other components. There are a number of advantages to organizing a system in this way, which we discuss. We have analyzed the performance of our system for a typical application and for a number of problem sizes, and have found that the overhead for using SOAPbased Web Services is small and tends to decrease as the problem size increases. Our results suggest that potential performance bottlenecks identified in the literature may not be major issues in practice, and that a standards-compliant implementation like ours can delivery excellent scalable performance even on coupled problems, provided Web Services are used judiciously. 1. Keshav Pingali, Paul Stodghill |
ICS | 2 |
| 2006 | Recent advances in checkpoint/recovery systemsabstractCheckpoint and recovery (CPR) systems have many uses in high-performance computing. Because of this, many developers have implemented it, by hand, into their applications. One of the uses of checkpointing is to help mitigate the effects of interruptions in computational service (both planned and unplanned) In fact, some supercomputing centers expect their users to use checkpointing as a matter of policy. And yet, few centers provide fully automatic checkpointing systems for their high-end production machines. The paper is a status report on our work on the family of C3systems for (almost) fully automatic checkpointing for scientific applications. To date, we have shown that our techniques can be used for checkpointing sequential, MPI and OpenMP applications written in C, Fortran, and several other languages. A novel aspect of our work is that we have not built a single checkpointing system, rather, we have developed a methodology and a set of techniques that have enabled us to develop a number of systems, each meeting different design goals and efficiency requirements Greg Bronevetsky, Rohit Fernandes, Daniel Marques, Keshav Pingali, Paul Stodghill |
IPDPS | 5 |
| 2006 | Mobile MPI programs in computational gridsabstractUtility computing is becoming a popular way of exploiting the potential of computational grids. In utility computing, users are provided with computational power in a transparent manner similar to the way in which electrical utilities supply power to their customers. To take full advantage of utility computing, an application needs to be mobile; that is, it needs to be able to migrate between heterogeneous computing platforms while it is executing. Further, it needs to be able to adapt to the computing resources at each site, such as the number of available physical processors. At present, there are few high-performance computing applications of this sort, and re-engineering legacy codes to be mobile can take enormous effort.In this paper, we describe theph$PC^3$ system, which converts C/MPI codes into mobile programs almost transparently. Because it is based on portable application-level checkpointing, it enables the state of running applications to be saved so that the application can be restarted on different architectures, operating systems and MPI implementations. Moreover, the number of processors on these machines can be different. To our knowledge, this is the first system to provide all these features. Experimental results show that the overhead introduced by the system is usually small. Rohit Fernandes, Keshav Pingali, Paul Stodghill |
PPoPP | 3 |
| 2005 | Think globally, search locallyabstractA key step in program optimization is the determination of optimal values for code optimization parameters such as cache tile sizes and loop unrolling factors. One approach, which is implemented in most compilers, is to use analytical models to determine these values. The other approach, used in library generators like ATLAS, is to perform a global empirical search over the space of parameter values.Neither approach is completely suitable for use in general-purpose compilers that must generate high quality code for large programs running on complex architectures. Model-driven optimization may incur a performance penalty of 10-20% even for a relatively simple code like matrix multiplication. On the other hand, global search is not tractable for optimizing large programs for complex architectures because the optimization space is too large.In this paper, we advocate a methodology for generating high-performance code without increasing search time dramatically. Our methodology has three components: (i) modeling, (ii) local search, and (iii) model refinement. We demonstrate this methodology by using it to eliminate the performance gap between code produced by a model-driven version of ATLAS described by us in prior work, and code produced by the original ATLAS system using global search. Kamen Yotov, Keshav Pingali, Paul Stodghill |
ICS | 3 |
| 2005 | Automatic measurement of memory hierarchy parametersabstractThe running time of many applications is dominated by the cost of memory operations. To optimize such applications for a given platform, it is necessary to have a detailed knowledge of the memory hierarchy parameters of that platform. In practice, this information is poorly documented if at all. Moreover, there is growing interest in self-tuning, autonomic software systems that can optimize themselves for different platforms; these systems must determine memory hierarchy parameters automatically without human intervention.One solution is to use micro-benchmarks to determine the parameters of the memory hierarchy. In this paper, we argue that existing micro-benchmarks are inadequate, and present novel micro-benchmarks for determining parameters of all levels of the memory hierarchy, including registers, all data caches and the translation look-aside buffer. We have implemented these micro-benchmarks in a tool called X-Ray that can be ported easily to new platforms. We present experimental results that show that X-Ray successfully determines memory hierarchy parameters on current platforms, and compare its accuracy with that of existing tools. Kamen Yotov, Keshav Pingali, Paul Stodghill |
SIGMETRICS | 3 |
| 2005 | Is Search Really Necessary to Generate High-Performance BLAS?abstractA key step in program optimization is the estimation of optimal values for parameters such as tile sizes and loop unrolling factors. Traditional compilers use simple analytical models to compute these values. In contrast, library generators like ATLAS use global search over the space of parameter values by generating programs with many different combinations of parameter values, and running them on the actual hardware to determine which values give the best performance. It is widely believed that traditional model-driven optimization cannot compete with search-based empirical optimization because tractable analytical models cannot capture all the complexities of modern high-performance architectures, but few quantitative comparisons have been done to date. To make such a comparison, we replaced the global search engine in ATLAS with a model-driven optimization engine and measured the relative performance of the code produced by the two systems on a variety of architectures. Since both systems use the same code generator, any differences in the performance of the code produced by the two systems can come only from differences in optimization parameter values. Our experiments show that model-driven optimization can be surprisingly effective and can generate code with performance comparable to that of code generated by ATLAS using global search. Kamen Yotov, Xiaoming Li 0004, Gang Ren 0002, María Jesús Garzarán, David A. Padua, Keshav Pingali, Paul Stodghill |
Proc. IEEE | 7 |
| 2004 | Implementation and Evaluation of a Scalable Application-Level Checkpoint-Recovery Scheme for MPI ProgramsabstractThe running times of many computational science applications are much longer than the mean-time-to-failure of current high-performance computing platforms. To run to completion, such applications must tolerate hardware failures. Checkpoint-and-restart (CPR) is the most commonly used scheme for accomplishing this - the state of the computation is saved periodically on stable storage, and when a hardware failure is detected, the computation is restarted from the most recently saved state. Most automatic CPR schemes in the literature can be classified as system-level checkpointing schemes because they take core-dump style snapshots of the computational state when all the processes are blocked at global barriers in the program. Unfortunately, a system that implements this style of checkpointing is tied to a particular platform; in addition, it cannot be used if there are no global barriers in the program. We are exploring an alternative called application-level, non-blocking checkpointing. In our approach, programs are transformed by a pre-processor so that they become self-checkpointing and self-restartable on any platform; there is also no assumption about the existence of global barriers in the code. In this paper, we describe our implementation of application-level, non-blocking checkpointing. We present experimental results on both a Windows cluster and a Compaq Alpha cluster, which show that the overheads introduced by our approach are small. Martin Schulz 0001, Greg Bronevetsky, Rohit Fernandes, Daniel Marques, Keshav Pingali, Paul Stodghill |
SC | 6 |
| 2003 | Collective operations in application-level fault-tolerant MPIabstractFault-tolerance is becoming a critical issue on high-performance platforms. Checkpointing techniques make programs fault-tolerant by saving their state periodically and restoring this state after failure. System-level checkpointing saves the state of the entire machine on stable storage, but this usually has too much overhead. In practice, programmers do manual checkpointing by writing code to (i) save the values of key program variables at critical points in the program, and (ii) restore the entire computational state from these values during recovery. However, this can be difficult to do in general MPI programs without global barriers.In an earlier paper, we presented a distributed checkpoint coordination protocol which handles MPI's point-to-point constructs, while dealing with the unique challenges of application-level checkpointing. The protocol is implemented by a thin software layer that sits between the application program and the MPI library, so it does not require any modifications to the MPI library. However, it did not handle collective communication, which is a very important part of MPI. In this paper, we extend the protocol to handle MPI's collective communication constructs. We also present experimental results that show that the overhead introduced by the protocol for collective operations is small. Greg Bronevetsky, Daniel Marques, Keshav Pingali, Paul Stodghill |
ICS | 4 |
| 2003 | A comparison of empirical and model-driven optimizationabstractEmpirical program optimizers estimate the values of key optimization parameters by generating different program versions and running them on the actual hardware to determine which values give the best performance. In contrast, conventional compilers use models of programs and machines to choose these parameters. It is widely believed that model-driven optimization does not compete with empirical optimization, but few quantitative comparisons have been done to date. To make such a comparison, we replaced the empirical optimization engine in ATLAS (a system for generating a dense numerical linear algebra library called the BLAS) with a model-driven optimization engine that used detailed models to estimate values for optimization parameters, and then measured the relative performance of the two systems on three different hardware platforms. Our experiments show that model-driven optimization can be surprisingly effective, and can generate code whose performance is comparable to that of code generated by empirical optimizers for the BLAS. Kamen Yotov, Xiaoming Li 0004, Gang Ren 0002, Michael Cibulskis, Gerald DeJong, María Jesús Garzarán, David A. Padua, Keshav Pingali, Paul Stodghill, Peng Wu 0001 |
PLDI | 9 |
| 2003 | Automated application-level checkpointing of MPI programsabstractThe running times of many computational science applications, such as protein-folding using ab initio methods, are much longer than the mean-time-to-failure of high-performance computing platforms. To run to completion, therefore, these applications must tolerate hardware failures.In this paper, we focus on the stopping failure model in which a faulty process hangs and stops responding to the rest of the system. We argue that tolerating such faults is best done by an approach called application-level coordinated non-blocking checkpointing, and that existing fault-tolerance protocols in the literature are not suitable for implementing this approach.We then present a suitable protocol, which is implemented by a co-ordination layer that sits between the application program and the MPI library. We show how this protocol can be used with a precompiler that instruments C/MPI programs to save application and MPI library state. An advantage of our approach is that it is independent of the MPI implementation. We present experimental results that argue that the overhead of using our system can be small. Greg Bronevetsky, Daniel Marques, Keshav Pingali, Paul Stodghill |
PPoPP | 4 |
| 2000 | Next-generation generic programming and its application to sparse matrix computationsabstractThe contributions of this paper are the following. Nikolay Mateev, Keshav Pingali, Paul Stodghill, Vladimir Kotlyar |
ICS | 3 |
| 2000 | A Framework for Sparse Matrix Code Synthesis from High-level SpecificationsabstractWe present compiler technology for synthesizing sparse matrix code from (i) dense matrix code, and (ii) a description of the index structure of a sparse matrix. Our approach is to embed statement instances into a Cartesian product of statement iteration and data spaces, and to produce efficient sparse code by identifying common enumerations for multiple references to sparse matrices. The approach works for imperfectly-nested codes with dependences, and produces sparse code competitive with hand-written library code for the Basic Linear Algebra Subroutines (BLAS). Nawaaz Ahmed, Nikolay Mateev, Keshav Pingali, Paul Stodghill |
SC | 4 |
| 2000 | Landing CG on EARTH: A Case Study of Fine-Grained Multithreading on an Evolutionary PathabstractWe report on our work in developing a fine-grained multithreaded solution for the communication-intensive Conjugate Gradient (CG) problem. In our recent work, we developed a simple yet efficient program for sparse matrix-vector multiply on a multi-threaded system. This paper presents an effective mechanism for the reduction-broadcast phase, which is integrated with the sparse MVM, resulting in a scalable implementation of the complete CG application. Three major observations from our experiments on the EARTH multithreaded testbed are: (1) The scalability of our CG implementation is impressive, e.g., absolute speedup is 90 on 120 processors for the NAS CG class B input. (2) Our dataflow-style reduction-broadcast network based on fine-grain multithreading is twice as fast as a serial reduction scheme on the same system. (3) By slowing down the network by a factor of 2, no notable degradation of overall CG performance was observed. Kevin B. Theobald, Gagan Agrawal, Rishi Kumar, Gerd Heber, Guang R. Gao, Paul Stodghill, Keshav Pingali |
SC | 6 |
| 1997 | A Relational Approach to the Compilation of Sparse Matrix Programs
Vladimir Kotlyar, Keshav Pingali, Paul Stodghill |
Euro-Par | 3 |
| 1997 | Compiling Parallel Code for Sparse Matrix ApplicationsabstractWe have developed a framework based on relational algebra for compiling efficient sparse matrix code from dense DO-ANY loops and a specification of the representation of the sparse matrix. In this paper, we show how this framework can be used to generate parallel code, and present experimental data that demonstrates that the code generated by our Bernoulli compiler achieves performance competitive with that of hand-written codes for important computational kernels. Vladimir Kotlyar, Keshav Pingali, Paul Stodghill |
SC | 3 |
| 1991 | Dependence Flow Graphs: An Algebraic Approach to Program DependenciesabstractThe topic of intermediate languages for optimizing and parallelizing compilers has received muchattention lately. In this paper, we argue that any good representation of a program must havetwo crucial properties: first, it must be a data structure that can be rapidly traversed to determine dependence information, and second this representation must be a program in its own right, with a parallel, local model of execution. In this paper, we illustrate the importance of these points by examining algorithms for a standard optimization --- global constant propagation. We discuss the problems in working with current representations. Then, we propose a novel representation called the dependence flow graph which has each of the properties mentioned above. Weshow that this representation leads to a simple algorithm, based on abstract interpretation, for solving the constant propagation problem. Our algorithm is simpler than, and as efficient as, the best known algorithms for this problem. An interesting feature of our representation is that it naturally incorporates the best aspects of many other representations, including continuation-passing style, data and program dependence graphs, static single assignment form and dataflow program graphs. Keshav Pingali, Micah D. Beck, Mayan Moudgill, Paul Stodghill |
POPL | 5 |