EDBT 2026 Demo / reviewers in the wild / expert
Elizabeth R. Jessup
dblp:j/ERJessup
· DBLP profile ↗
7ranked-venue papers
0as first author
0since 2021 · last 2019
0000-0002-7740-9985ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Debugging and program repair · 62% Program analysis · 31% Compilers and program optimization · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Debugging and program repair
fault localization |
0.4 | 1 | 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019 |
Program analysis › static analysis
program slicing |
0.4 | 1 | 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019 |
Debugging and program repair
root cause analysis |
0.4 | 1 | 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019 |
Environmental and earth informatics
climate modeling |
0.1 | 1 | 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate Model · HPDC 2019 |
Compilers and program optimization
domain-specific compilation |
0.1 | 1 | 2009 | Automating the generation of composed linear algebra kernels · SC 2009 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.0 | 1 | 2009 | Automating the generation of composed linear algebra kernels · SC 2009 |
Methods — techniques the papers use, named apart from their topics
runtime variable sampling · 0.8hybrid program slicing · 0.8directed graph · 0.8community partitioning · 0.8centrality ranking · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate ModelabstractLarge-scale simulation codes that model complicated science and engineering applications typically have huge and complex code bases. For such simulation codes, where bit-for-bit comparisons are too restrictive, finding the source of statistically significant discrepancies (e.g., from a previous version, alternative hardware or supporting software stack) in output is non-trivial at best. Although there are many tools for program comprehension through debugging or slicing, few (if any) scale to a model as large as the Community Earth System Model (CESM#8482;), which consists of more than 1.5 million lines of Fortran code. Currently for the CESM, we can easily determine whether a discrepancy exists in the output using a by now well-established statistical consistency testing tool. However, this tool provides no information as to the possible cause of the detected discrepancy, leaving developers in a seemingly impossible (and frustrating) situation. Therefore, our aim in this work is to provide the tools to enable developers to trace a problem detected through the CESM output to its source. To this end, our strategy is to reduce the search space for the root cause(s) to a tractable size via a series of techniques that include creating a directed graph of internal CESM variables, extracting a subgraph (using a form of hybrid program slicing), partitioning into communities, and ranking nodes by centrality. Runtime variable sampling then becomes feasible in this reduced search space. We demonstrate the utility of this process on multiple examples of CESM simulation output by illustrating how sampling can be performed as part of an efficient parallel iterative refinement procedure to locate error sources, including sensitivity to CPU instructions. By providing CESM developers with tools to identify and understand the reason for statistically distinct output, we have positively impacted the CESM software development cycle and, in particular, its focus on quality assurance. Daniel Milroy, Allison H. Baker, Dorit Hammerling, Youngsung Kim, Elizabeth R. Jessup, Thomas Hauser |
HPDC | 5 |
| 2015 | Generating Efficient Tensor Contractions for GPUsabstractMany scientific and numerical applications, including quantum chemistry modeling and fluid dynamics simulation, require tensor product and tensor contraction evaluation. Tensor computations are characterized by arrays with numerous dimensions, inherent parallelism, moderate data reuse and many degrees of freedom in the order in which to perform the computation. The best-performing implementation is heavily dependent on the tensor dimensionality and the target architecture. In this paper, we map tensor computations to GPUs, starting with a high-level tensor input language and producing efficient CUDA code as output. Our approach is to combine tensor-specific mathematical transformations with a GPU decision algorithm, machine learning and auto tuning of a large parameter space. Generated code shows significant performance gains over sequential and Open MP parallel code, and a comparison with Open ACC shows the importance of auto tuning and other optimizations in our framework for achieving efficient results. Thomas Nelson, Axel Rivera, Prasanna Balaprakash, Mary W. Hall, Paul D. Hovland, Elizabeth R. Jessup, Boyana Norris |
ICPP | 6 |
| 2015 | Reliable Generation of High-Performance Matrix AlgebraabstractScientific programmers often turn to vendor-tuned Basic Linear Algebra Subprograms (BLAS) to obtain portable high performance. However, many numerical algorithms require several BLAS calls in sequence, and those successive calls do not achieve optimal performance. The entire sequence needs to be optimized in concert. Instead of vendor-tuned BLAS, a programmer could start with source code in Fortran or C (e.g., based on the Netlib BLAS) and use a state-of-the-art optimizing compiler. However, our experiments show that optimizing compilers often attain only one-quarter of the performance of hand-optimized code. In this article, we present a domain-specific compiler for matrix kernels, the Build to Order BLAS (BTO), that reliably achieves high performance using a scalable search algorithm for choosing the best combination of loop fusion, array contraction, and multithreading for data parallelism. The BTO compiler generates code that is between 16% slower and 39% faster than hand-optimized code. Thomas Nelson, Geoffrey Belter, Jeremy G. Siek, Elizabeth R. Jessup, Boyana Norris |
ACM Trans. Math. Softw. | 4 |
| 2009 | Automating the generation of composed linear algebra kernelsabstractMemory bandwidth limits the performance of important kernels in many scientific applications. Such applications often use sequences of Basic Linear Algebra Subprograms (BLAS), and highly efficient implementations of those routines enable scientists to achieve high performance at little cost. However, tuning the BLAS in isolation misses opportunities for memory optimization that result from composing multiple subprograms. Because it is not practical to create a library of all BLAS combinations, we have developed a domain-specific compiler that generates them on demand. In this paper, we describe a novel algorithm for compiling linear algebra kernels and searching for the best combination of optimization choices. We also present a new hybrid analytic/empirical method for quickly evaluating the profitability of each optimization. We report experimental results showing speedups of up to 130% relative to the GotoBLAS on an AMD Opteron and up to 137% relative to MKL on an Intel Core 2. Geoffrey Belter, Elizabeth R. Jessup, Ian Karlin, Jeremy G. Siek |
SC | 2 |
| 2008 | Build to order linear algebra kernelsabstractThe performance bottleneck for many scientific applications is the cost of memory access inside linear algebra kernels. Tuning such kernels for memory efficiency is a complex task that reduces the productivity of computational scientists. Software libraries such as the Basic Linear Algebra Subprograms (BLAS) ameliorate this problem by providing a standard interface for which computer scientists and hardware vendors have created highly-tuned implementations. Scientific applications often require a sequence of BLAS operations, which presents further opportunities for memory optimization. However, because BLAS are tuned in isolation they do not take advantage of these opportunities. This phenomenon motivated the recent addition to the BLAS of several routines that perform sequences of operations. Unfortunately, the exact sequence of operations needed in a given situation is highly application dependent, so many more routines are needed. In this paper we present preliminary work on a domain- specific compiler that generates implementations for arbitrary sequences of basic linear algebra operations and tunes them for memory efficiency. We report experimental results for dense kernels and show speedups of 25 % to 120 % relative to sequences of calls to GotoBLAS and vendor-tuned BLAS on Intel Xeon and IBM PowerPC platforms. Jeremy G. Siek, Ian Karlin, Elizabeth R. Jessup |
IPDPS | 3 |
| 1999 | The PMESC Programming Library for Distributed-Memory MIMD Computers
Silvia A. Crivelli, Elizabeth R. Jessup |
J. Parallel Distributed Comput. | 2 |
| 1995 | The Cost of Eigenvalue Computation on Distributed-Memory MIMD Multiprocessors
Silvia A. Crivelli, Elizabeth R. Jessup |
Parallel Comput. | 2 |