Elizabeth R. Jessup

dblp:j/ERJessup · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0002-7740-9985ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Debugging and program repair · 62% Program analysis · 31% Compilers and program optimization · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
fault localization
0.412019
Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019
Program analysis › static analysis
program slicing
0.412019
Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019
Debugging and program repair
root cause analysis
0.412019
Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019
Environmental and earth informatics
climate modeling
0.112019
Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019
Compilers and program optimization
domain-specific compilation
0.112009
Automating the generation of composed linear algebra kernels · SC 2009
Memory systems › memory bandwidth
memory bandwidth optimization
0.012009
Automating the generation of composed linear algebra kernels · SC 2009

Methods — techniques the papers use, named apart from their topics

runtime variable sampling · 0.8hybrid program slicing · 0.8directed graph · 0.8community partitioning · 0.8centrality ranking · 0.8
YearPublicationVenuePosition
2019 Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model
abstract
Large-scale simulation codes that model complicated science and engineering applications typically have huge and complex code bases. For such simulation codes, where bit-for-bit comparisons are too restrictive, finding the source of statistically significant discrepancies (e.g., from a previous version, alternative hardware or supporting software stack) in output is non-trivial at best. Although there are many tools for program comprehension through debugging or slicing, few (if any) scale to a model as large as the Community Earth System Model (CESM#8482;), which consists of more than 1.5 million lines of Fortran code. Currently for the CESM, we can easily determine whether a discrepancy exists in the output using a by now well-established statistical consistency testing tool. However, this tool provides no information as to the possible cause of the detected discrepancy, leaving developers in a seemingly impossible (and frustrating) situation. Therefore, our aim in this work is to provide the tools to enable developers to trace a problem detected through the CESM output to its source. To this end, our strategy is to reduce the search space for the root cause(s) to a tractable size via a series of techniques that include creating a directed graph of internal CESM variables, extracting a subgraph (using a form of hybrid program slicing), partitioning into communities, and ranking nodes by centrality. Runtime variable sampling then becomes feasible in this reduced search space. We demonstrate the utility of this process on multiple examples of CESM simulation output by illustrating how sampling can be performed as part of an efficient parallel iterative refinement procedure to locate error sources, including sensitivity to CPU instructions. By providing CESM developers with tools to identify and understand the reason for statistically distinct output, we have positively impacted the CESM software development cycle and, in particular, its focus on quality assurance.
Daniel Milroy, Allison H. Baker, Dorit Hammerling, Youngsung Kim, Elizabeth R. Jessup, Thomas Hauser
HPDC5
2015 Generating Efficient Tensor Contractions for GPUs
abstract
Many scientific and numerical applications, including quantum chemistry modeling and fluid dynamics simulation, require tensor product and tensor contraction evaluation. Tensor computations are characterized by arrays with numerous dimensions, inherent parallelism, moderate data reuse and many degrees of freedom in the order in which to perform the computation. The best-performing implementation is heavily dependent on the tensor dimensionality and the target architecture. In this paper, we map tensor computations to GPUs, starting with a high-level tensor input language and producing efficient CUDA code as output. Our approach is to combine tensor-specific mathematical transformations with a GPU decision algorithm, machine learning and auto tuning of a large parameter space. Generated code shows significant performance gains over sequential and Open MP parallel code, and a comparison with Open ACC shows the importance of auto tuning and other optimizations in our framework for achieving efficient results.
Thomas Nelson, Axel Rivera, Prasanna Balaprakash, Mary W. Hall, Paul D. Hovland, Elizabeth R. Jessup, Boyana Norris
ICPP6
2015 Reliable Generation of High-Performance Matrix Algebra
abstract
Scientific programmers often turn to vendor-tuned Basic Linear Algebra Subprograms (BLAS) to obtain portable high performance. However, many numerical algorithms require several BLAS calls in sequence, and those successive calls do not achieve optimal performance. The entire sequence needs to be optimized in concert. Instead of vendor-tuned BLAS, a programmer could start with source code in Fortran or C (e.g., based on the Netlib BLAS) and use a state-of-the-art optimizing compiler. However, our experiments show that optimizing compilers often attain only one-quarter of the performance of hand-optimized code. In this article, we present a domain-specific compiler for matrix kernels, the Build to Order BLAS (BTO), that reliably achieves high performance using a scalable search algorithm for choosing the best combination of loop fusion, array contraction, and multithreading for data parallelism. The BTO compiler generates code that is between 16% slower and 39% faster than hand-optimized code.
Thomas Nelson, Geoffrey Belter, Jeremy G. Siek, Elizabeth R. Jessup, Boyana Norris
ACM Trans. Math. Softw.4
2009 Automating the generation of composed linear algebra kernels
abstract
Memory bandwidth limits the performance of important kernels in many scientific applications. Such applications often use sequences of Basic Linear Algebra Subprograms (BLAS), and highly efficient implementations of those routines enable scientists to achieve high performance at little cost. However, tuning the BLAS in isolation misses opportunities for memory optimization that result from composing multiple subprograms. Because it is not practical to create a library of all BLAS combinations, we have developed a domain-specific compiler that generates them on demand. In this paper, we describe a novel algorithm for compiling linear algebra kernels and searching for the best combination of optimization choices. We also present a new hybrid analytic/empirical method for quickly evaluating the profitability of each optimization. We report experimental results showing speedups of up to 130% relative to the GotoBLAS on an AMD Opteron and up to 137% relative to MKL on an Intel Core 2.
Geoffrey Belter, Elizabeth R. Jessup, Ian Karlin, Jeremy G. Siek
SC2
2008 Build to order linear algebra kernels
abstract
The performance bottleneck for many scientific applications is the cost of memory access inside linear algebra kernels. Tuning such kernels for memory efficiency is a complex task that reduces the productivity of computational scientists. Software libraries such as the Basic Linear Algebra Subprograms (BLAS) ameliorate this problem by providing a standard interface for which computer scientists and hardware vendors have created highly-tuned implementations. Scientific applications often require a sequence of BLAS operations, which presents further opportunities for memory optimization. However, because BLAS are tuned in isolation they do not take advantage of these opportunities. This phenomenon motivated the recent addition to the BLAS of several routines that perform sequences of operations. Unfortunately, the exact sequence of operations needed in a given situation is highly application dependent, so many more routines are needed. In this paper we present preliminary work on a domain- specific compiler that generates implementations for arbitrary sequences of basic linear algebra operations and tunes them for memory efficiency. We report experimental results for dense kernels and show speedups of 25 % to 120 % relative to sequences of calls to GotoBLAS and vendor-tuned BLAS on Intel Xeon and IBM PowerPC platforms.
Jeremy G. Siek, Ian Karlin, Elizabeth R. Jessup
IPDPS3
1999 The PMESC Programming Library for Distributed-Memory MIMD Computers
Silvia A. Crivelli, Elizabeth R. Jessup
J. Parallel Distributed Comput.2
1995 The Cost of Eigenvalue Computation on Distributed-Memory MIMD Multiprocessors
Silvia A. Crivelli, Elizabeth R. Jessup
Parallel Comput.2