EDBT 2026 Demo / reviewers in the wild / expert
R. Clint Whaley
dblp:56/3316 · also R. Clinton Whaley
· DBLP profile ↗
14ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 60% Memory systems · 23% Parallel and multicore computing · 10% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache
cache-aware algorithm design |
0.1 | 1 | 2010 | Scaling LAPACK panel operations using parallel cache assignment · PPoPP 2010 |
High-performance computing › numerical linear algebra
dense linear algebra |
0.1 | 1 | 2010 | Scaling LAPACK panel operations using parallel cache assignment · PPoPP 2010 |
High-performance computing › linear algebra library
BLAS |
0.1 | 2 | 2005 | Self-Adapting Linear Algebra Algorithms and Software · Proc. IEEE 2005 Automatically Tuned Linear Algebra Software · SC 1998 |
High-performance computing › performance optimization
auto-tuning |
0.0 | 1 | 1998 | Automatically Tuned Linear Algebra Software · SC 1998 |
Embedded and real-time systems › model-based design
code generation |
0.0 | 1 | 1998 | Automatically Tuned Linear Algebra Software · SC 1998 |
High-performance computing › performance optimization
numerical program optimization |
0.0 | 1 | 1998 | Automatically Tuned Linear Algebra Software · SC 1998 |
Parallel and multicore computing
MPI |
0.0 | 1 | 1996 | ScaLAPACK: A Portable Linear Algebra Library for Distributed Memory Computers - Design Issues and Performance · SC 1996 |
High-performance computing
numerical linear algebra |
0.0 | 1 | 1996 | ScaLAPACK: A Portable Linear Algebra Library for Distributed Memory Computers - Design Issues and Performance · SC 1996 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.0 | 1 | 1998 | Automatically Tuned Linear Algebra Software · SC 1998 |
Methods — techniques the papers use, named apart from their topics
cache assignment · 0.1block algorithms · 0.1search-based autotuning · 0.1empirical performance modeling · 0.1empirical search · 0.0automatic code generation · 0.0message passing · 0.0block-partitioned algorithms · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Effectively Exploiting Parallel Scale for All Problem Sizes in LU FactorizationabstractLU factorization is one of the most widely-used methods for solving linear equations, and thus its performance underlies a broad range of scientific computing. As architectural trends have replaced clock rate improvements with increases in parallel scale, library writers have responded by using tiled algorithms, where operand size is constrained in order to maximize parallelism, as seen in the well-known PLASMA library. This approach has two main drawbacks: (1) asymptotic performance is reduced due to limited operand size, (2) performance of small to medium sized problems is reduced due to unnecessary data motion in the parallel caches. In this paper we introduce a new approach where asymptotic performance is maximized by using special low-overhead kernel primitives that are auto-generated by the ATLAS framework, while unnecessary cache motion is minimized by using explicit cache management. We show that this technique can outperform all known libraries at all problem sizes on commodity parallel Intel and AMD platforms, with asymptotic LU performance of roughly 91% of hardware theoretical peak for a 12-core Intel Xeon, and 87% for a 32-core AMD Opteron. Md Rakib Hasan, R. Clint Whaley |
IPDPS | 2 |
| 2013 | Vectorization past dependent branches through speculationabstractModern architectures increasingly rely on SIMD vectorization to improve performance for floating point intensive scientific applications. However, existing compiler optimization techniques for automatic vectorization are inhibited by the presence of unknown control flow surrounding partially vectorizable computations. In this paper, we present a new approach, speculative vectorization, which speculates past dependent branches to aggressively vectorize computational paths that are expected to be taken frequently at runtime, while simply restarting the calculation using scalar instructions when the speculation fails. We have integrated our technique in an iterative optimizing compiler and have employed empirical tuning to select the profitable paths for speculation. When applied to optimize 9 floating-point benchmarks, our optimizing compiler has achieved up to 6.8X speedup for single precision and 3.4X for double precision kernels using AVX, while vectorizing some operations considered not vectorizable by prior techniques. Majedul Haque Sujon, R. Clint Whaley, Qing Yi |
PACT | 2 |
| 2013 | Scaling LAPACK panel operations using parallel cache assignmentabstractIn LAPACK many matrix operations are cast as block algorithms which iteratively process a panel using an unblocked algorithm and then update a remainder matrix using the high performance Level 3 BLAS. The Level 3 BLAS have excellent scaling, but panel processing tends to be bus bound, and thus scales with bus speed rather than the number of processors ( p ). Amdahl's law therefore ensures that as p grows, the panel computation will become the dominant cost of these LAPACK routines. Our contribution is a novel parallel cache assignment approach to the panel factorization which we show scales well with p . We apply this general approach to the QR, QL, RQ, LQ and LU panel factorizations. We show results for two commodity platforms: an 8-core Intel platform and a 32-core AMD platform. For both platforms and all twenty implementations (five factorizations each of which is available in 4 types), we present results that demonstrate that our approach yields significant speedup over the existing state of the art. Anthony M. Castaldo, R. Clint Whaley, Siju Samuel |
ACM Trans. Math. Softw. | 2 |
| 2011 | Achieving Scalable Parallelization for the Hessenberg FactorizationabstractMuch of dense linear algebra has been successfully blocked to concentrate the majority of its time in the Level 3 BLAS, which are not only efficient for serial computation, but also scale well for parallelism. For the Hessenberg factorization, which is a critical step in computing the eigenvalues and vectors, however, performance of the best known algorithm is still strongly limited by the memory speed, which does not tend to scale well at all. In this paper we present an adaptation of our Parallel Cache Assignment (PCA) technique to the Hessenberg factorization, and show that it achieves super linear speedup over the corresponding serial algorithm and a more than four-fold speedup over the best known algorithm for small and medium sized problems. Anthony M. Castaldo, R. Clint Whaley |
CLUSTER | 2 |
| 2010 | Scaling LAPACK panel operations using parallel cache assignmentabstractIn LAPACK many matrix operations are cast as block algorithms which iteratively process a panel using an unblocked algorithm and then update a remainder matrix using the high performance Level 3 BLAS. The Level~3 BLAS have excellent weak scaling, but panel processing tends to be bus bound, and thus scales with bus speed rather than the number of processors (p). Amdahl's law therefore ensures that as p grows, the panel computation will become the dominant cost of these LAPACK routines. Our contribution is a novel parallel cache assignment approach which we show scales well with p. We apply this general approach to the QR and LU panel factorizations on two commodity 8-core platforms with very different cache structures, and demonstrate superlinear panel factorization speedups on both machines. Other approaches to this problem demand complicated reformulations of the computational approach, new kernels to be tuned, new mathematics, an inflation of the high-order flop count, and do not perform as well. By demonstrating a straight-forward alternative that avoids all of these contortions and scales with p, we address a critical stumbling block for dense linear algebra in the age of massive parallelism. Anthony M. Castaldo, R. Clint Whaley |
PPoPP | 2 |
| 2009 | Minimizing startup costs for performance-critical threadingabstractUsing the well-known ATLAS and LAPACK dense linear algebra libraries, we demonstrate that the parallel management overhead (PMO) can grow with problem size on even statically scheduled parallel programs with minimal task interaction. Therefore, the widely held view that these thread management issues can be ignored in such computationally intensive libraries is wrong, and leads to substantial slowdown on today's machines. We survey several methods for reducing this overhead, the best of which we have not seen in the literature. Finally, we demonstrate that by applying these techniques at the kernel level, performance in applications such as LU and QR factorizations can be improved by almost 40% for small problems, and as much as 15% for large O(N3) computations. These techniques are completely general, and should yield significant speedup in almost any performance-critical operation.We then show that the lion's share of the remaining parallel inefficiency comes from bus contention, and, in the future work section, outline some promising avenues for further improvement. Anthony M. Castaldo, R. Clint Whaley |
IPDPS | 2 |
| 2008 | Achieving accurate and context-sensitive timing for code optimizationabstractAbstract Key computational kernels must run near their peak efficiency for most high‐performance computing (HPC) applications. Getting this level of efficiency has always required extensive tuning of the kernel on a particular platform of interest. The success or failure of an optimization is usually measured by invoking a timer. Understanding how to build reliable and context‐sensitive timers is one of the most neglected areas in HPC, and this results in a host of HPC software that looks good when reported in the papers, but delivers only a fraction of the reported performance when used by actual HPC applications. In this paper, we motivate the importance of timer design and then discuss the techniques and methodologies we have developed in order to accurately time HPC kernel routines for our well‐known empirical tuning framework, ATLAS. Copyright © 2008 John Wiley & Sons, Ltd. R. Clint Whaley, Anthony M. Castaldo |
Softw. Pract. Exp. | 1 |
| 2005 | Tuning High Performance Kernels through Empirical CompilationabstractThere are a few application areas, which remain almost untouched by the historical and continuing advancement of compilation research. For the extremes of optimization required for high performance computing on one end, and embedded systems at the opposite end of the spectrum, many critical routines are still hand-tuned, often directly in assembly. At the same time, architecture implementations are performing an increasing number of compiler-like transformations in hardware, making it harder to predict the performance impact of a given series of optimizations applied at the ISA level. These issues, together with the rate of hardware evolution dictated by Moore's Law, make it almost impossible to keep key kernels running at peak efficiency. Automated empirical systems, where direct timings are used to guide optimization, have provided the most successful response to these challenges. This paper describes our approach to performing empirical optimization, which utilizes a low-level iterative compilation framework specialized for optimizing high performance computing kernels. We present results showing that this approach can not only provide speedups over traditional optimizing compilers, but can improve overall performance when compared to the best hand-tuned kernels selected by the empirical search of our well-known ATLAS package. R. Clint Whaley, David B. Whalley |
ICPP | 1 |
| 2005 | Self-Adapting Linear Algebra Algorithms and SoftwareabstractOne of the main obstacles to the efficient solution of scientific problems is the problem of tuning software, both to the available architecture and to the user problem at hand. We describe approaches for obtaining tuned high-performance kernels and for automatically choosing suitable algorithms. Specifically, we describe the generation of dense and sparse Basic Linear Algebra Subprograms (BLAS) kernels, and the selection of linear solver algorithms. However, the ideas presented here extend beyond these areas, which can be considered proof of concept. Richard Carl Demmel, Jack J. Dongarra, Victor Eijkhout, Erika Fuentes, Antoine Petitet, Richard W. Vuduc, R. Clint Whaley, Katherine A. Yelick |
Proc. IEEE | 7 |
| 2005 | Minimizing development and maintenance costs in supporting persistently optimized BLASabstractThe Basic Linear Algebra Subprograms (BLAS) define one of the most heavily used performance-critical APIs in scientific computing today. It has long been understood that the most important of these routines, the dense Level 3 BLAS, may be written efficiently given a highly optimized general matrix multiply routine. In this paper, however, we show that an even larger set of operations can be efficiently maintained using a much simpler matrix multiply kernel. Indeed, this is how our own project, ATLAS (which provides one of the most widely used BLAS implementations in use today), supports a large variety of performance-critical routines. Copyright © 2004 John Wiley & Sons, Ltd. R. Clint Whaley, Antoine Petitet |
Softw. Pract. Exp. | 1 |
| 2001 | Automated empirical optimizations of software and the ATLAS project
R. Clint Whaley, Antoine Petitet, Jack J. Dongarra |
Parallel Comput. | 1 |
| 1998 | Automatically Tuned Linear Algebra SoftwareabstractThis paper describes an approach for the automatic generation and optimization of numerical software for processors with deep memory hierarchies and pipelined functional units. The production of such software for machines ranging from desktop workstations to embedded processors can be a tedious and time consuming process. The work described here can help in automating much of this process. We will concentrate our efforts on the widely used linear algebra kernels called the Basic Linear Algebra Subroutines (BLAS). In particular, the work presented here is for general matrix multiply, DGEMM. However much of the technology and approach developed here can be applied to the other Level 3 BLAS and the general strategy can have an impact on basic linear algebra operations in general and may be extended to other important kernel operations. R. Clint Whaley, Jack J. Dongarra |
SC | 1 |
| 1997 | Practical Experience in the Numerical Dangers of Heterogeneous ComputingabstractSpecial challenges exist in writing reliable numerical library software for heterogeneous computing environments. Although a lot of software for distributed-memory parallel computers has been written, porting this software to a network of workstations requires careful consideration. The symptoms of heterogeneous computing failures can range from erroneous results without warning to deadlock. Some of the problems are straightforward to solve, but for others the solutions are not so obvious, or incur an unacceptable overhead. Making software robust on heterogeneous systems often requires additional communication. We describe and illustrate the problems encountered during the development of ScaLAPACK and the NAG Numerical PVM Library. Where possible, we suggest ways to avoid potential pitfalls, or if that is not possible, we recommend that the software not be used on heterogeneous networks. L. Susan Blackford, Andrew J. Cleary, Antoine Petitet, R. Clint Whaley, James Demmel, Inderjit S. Dhillon, H. Ren, Ken Stanley, Jack J. Dongarra, Sven Hammarling |
ACM Trans. Math. Softw. | 4 |
| 1996 | ScaLAPACK: A Portable Linear Algebra Library for Distributed Memory Computers - Design Issues and PerformanceabstractThis paper outlines the content and performance of ScaLAPACK, a collection of mathematical software for linear algebra computations on distributed memory computers. The importance of developing standards for computational and message passing interfaces is discussed. We present the different components and building blocks of ScaLAPACK, and indicate the difficulties inherent in producing correct codes for networks of heterogeneous processors. Finally, this paper briefly describes future directions for the ScaLAPACK library and concludes by suggesting alternative approaches to mathematical libraries, explaining how ScaLAPACK could be integrated into efficient and user-friendly distributed systems. L. Susan Blackford, Andrew J. Cleary, James Demmel, Inderjit S. Dhillon, Jack J. Dongarra, Sven Hammarling, Greg Henry, Antoine Petitet, Ken Stanley, David W. Walker, R. Clint Whaley |
SC | 12 |