VLDB 2026 Research / reviewers in the wild / expert
Paul D. Hovland
dblp:79/3143
· DBLP profile ↗
37ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-0907-2567ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 9 · 5 since 2021Theory of computation · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thinking Fast and Correct: Automated Rewriting of Numerical Code through Compiler AugmentationabstractFloating-point numbers are finite-precision approximations to real numbers and are ubiquitous in computer applications in nearly every field. Selecting the right floating-point representation that balances performance and numerical accuracy is a difficult task – one that has become even more critical as hardware trends toward high-performance, low-precision operations. Although the common wisdom around changing floating-point precision implies that accuracy and performance are inversely correlated, more advanced techniques can often circumvent this tradeoff. Applying complex numerical optimizations to real-world code, however, is an arduous engineering task that requires expertise in numerical analysis and performance engineering, and application-specific numerical context. While there is a plethora of existing tools that partially automate this process, they are limited in the scope of optimization techniques or still require substantial human intervention. We present Poseidon, a modular and extensible framework that fully automates floating-point optimizations for real-world applications within a production compiler. Our key insight is that a small surrogate profile often reveals sufficient numerical context to drive effective rewrites. Poseidon operates as a two-phase compiler: the first compilation instruments the program to capture numerical context; the second compilation consumes profiled data, generates and evaluates candidate rewrites, and solves for optimal performance/accuracy tradeoffs. Poseidon’s interoperability with standard compiler analyses and optimizations grants it analysis and optimization advantages unavailable to existing source- and binary-level approaches. On multiple large-scale applications, Poseidon leads to outsized benefits in performance without substantially changing accuracy, and outsized accuracy benefits without diminishing performance. On a quaternion differentiator, Poseidon enables a 1.46× speedup with a relative error of 10−7. On DOE’s LULESH hydrodynamics application, Poseidon improves program accuracy to exactly match a 512-bit simulation run without substantially reducing performance. Siyuan Brant Qian, Vimarsh Sathia, Ivan R. Ivanov, Jan Hückelheim, Paul D. Hovland, William S. Moses |
CGO | 5 |
| 2025 | Verifying PETSc Vector Components Using CIVLabstractAbstract This paper presents a modular approach to verifying the vector module of PETSc, a widely used library in scientific computing, using the CIVL model checker. Our approach relies on the creation of stub functions, which serve a dual purpose of specifying the intended behavior of individual PETSc functions and providing an abstraction for called functions to allow efficient verification of callers. This facilitates the use of symbolic execution and model checking to establish the correctness of isolated functions. Our work contributes to the ongoing effort to enhance the reliability of high-performance computing libraries and proposes an effective verification strategy for complex scientific software. Venkata Dhavala, Jan Hückelheim, Paul D. Hovland, Stephen F. Siegel |
CAV (1) | 3 |
| 2025 | Hardware-Software Co-design for Distributed Quantum ComputingabstractDistributed quantum computing (DQC) offers a pathway for scaling up quantum computing architectures beyond the confines of a single chip. Entanglement is a crucial resource for implementing nonlocal operations in DQC, and it is required to allow teleportation of quantum states and gates. Remote entanglement generation in practical systems is probabilistic, has longer duration than that of local operations, and is nondeterministic. Therefore, optimizing the performance of probabilistic remote entanglement generation is critically important for the performance of DQC architectures. In this paper we propose and study a new DQC architecture that combines (1) buffering of successfully generated entanglement, (2) asynchronously attempted entanglement generation, and (3) adaptive scheduling of remote gates based on the entanglement generation pattern. We show that our hardware-software co-design improves both the runtime and the output fidelity under a realistic model of DQC. Ji Liu 0007, Allen Zang, Martin Suchara, Tian Zhong, Paul D. Hovland |
DAC | 5 |
| 2025 | CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank ActivationabstractZiyue Liu, Ruijie Zhang, Zhengyang Wang, Mingsong Yan, Zi Yang, Paul D. Hovland, Bogdan Nicolae, Franck Cappello, Sui Tang, Zheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mingsong Yan, Paul D. Hovland, Bogdan Nicolae, Franck Cappello, Sui Tang |
EMNLP | 6 |
| 2025 | HOPPS: Hardware-Aware Optimal Phase Polynomial Synthesis with Blockwise Optimization for Quantum CircuitsabstractBlocks composed of {CNOT, Rz} are ubiquitous in modern quantum applications, notably in circuits such as QAOA ansatzes and quantum adders. After compilation, many of them exhibit large CNOT counts or depths, which lowers fidelity. Therefore, we introduce HOPPS: a SAT-based hardware-aware optimal phase polynomial synthesis algorithm that could generate {CNOT, Rz} blocks with CNOT count or depth optimality. Sometime {CNOT, Rz} blocks are large, such as in QAOA ansatzes, HOPPS's pursuit of optimality limits its scalability. To address this issue, we introduce an iterative blockwise optimization strategy: large circuits are partitioned into smaller blocks, each block is optimally refined, and the process is repeated for several iterations. Empirical results show that HOPPS is more efficient comparing with existing near-optimal synthesis tools. Used as a peephole optimizer, HOPPS reduces the CNOT count by up to 50.0 % and the CNOT depth by up to 57.1 % under OLSQ. For large QAOA circuit, after mapping by Qiskit, circuit can be reduced CNOT count and depth by up to 44.4 % and 42.4 % by our iterative blockwise optimization. Ji Liu 0007, Paul D. Hovland, Vipin Chaudhary |
HiPC | 4 |
| 2025 | QuCLEAR: Clifford Extraction and Absorption for Quantum Circuit OptimizationabstractQuantum computing carries significant potential for addressing practical problems. However, currently available quantum devices suffer from noisy quantum gates, which degrade the fidelity of executed quantum circuits. Therefore, quantum circuit optimization is crucial for obtaining useful results. In this paper, we present QuCLEAR, a compilation framework designed to optimize quantum circuits. QuCLEAR significantly reduces both the two-qubit gate count and the circuit depth through two novel optimization steps. First, we introduce the concept of Clifford Extraction, which extracts Clifford subcircuits to the end of the circuit while optimizing the gates. Second, since Clifford circuits are classically simulatable, we propose Clifford Absorption, which efficiently processes the extracted Clifford subcircuits classically. We demonstrate our framework on quantum simulation circuits, which have wideranging applications in quantum chemistry simulation, manybody physics, and combinatorial optimization problems. Nearterm algorithms such as VQE and QAOA also fall within this category. Experimental results across various benchmarks show that QuCLEAR achieves up to a 77.7% reduction in CNOT gate count and up to an 84.1% reduction in entangling depth compared with state-of-the-art methods. Ji Liu 0007, Alvin Gonzales, Benchen Huang, Zain H. Saleem, Paul D. Hovland |
HPCA | 5 |
| 2025 | ytopt: Autotuning Scientific Applications for Energy Efficiency at Large ScalesabstractABSTRACT As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low‐overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large‐scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes. Xingfu Wu, Prasanna Balaprakash, Michael Kruse, Jaehoon Koo, Brice Videau, Paul D. Hovland, Valerie Taylor 0001, Brad Geltz, Siddhartha Jana, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 6 |
| 2025 | MITgcm-AD v2: Open source tangent linear and adjoint modeling framework for the oceans and atmosphere enabled by the Automatic Differentiation tool TapenadeabstractThe Massachusetts Institute of Technology General Circulation Model (MITgcm) is widely used by the climate science community to simulate planetary atmosphere and ocean circulations. A defining feature of the MITgcm is that it has been developed to be compatible with an algorithmic differentiation (AD) tool, TAF, enabling the generation of tangent-linear and adjoint models. These provide gradient information which enables dynamics-based sensitivity and attribution studies, state and parameter estimation, and rigorous uncertainty quantification. Importantly, gradient information is essential for computing comprehensive sensitivities and performing efficient large-scale data assimilation, ensuring that observations collected from satellites and in-situ measuring instruments can be effectively used to optimize a large uncertain control space. As a result, the MITgcm forms the dynamical core of a key data assimilation product employed by the physical oceanography research community: Estimating the Circulation and Climate of the Ocean (ECCO) state estimate. Although MITgcm and ECCO are used extensively within the research community, the AD tool TAF is proprietary and hence inaccessible to a large proportion of these users. The new version 2 (MITgcm-AD v2) framework introduced here is based on the source-to-source AD tool Tapenade, which has recently been open-sourced. Another feature of Tapenade is that it stores required variables by default (instead of recomputing them) which simplifies the implementation of efficient, AD-compatible code. The framework has been integrated with the MITgcm model’s main branch and is now freely available. Shreyas Sunil Gaikwad, Sri Hari Krishna Narayanan, Laurent Hascoët, Jean-Michel Campin, Helen Pillar, Jan Hückelheim, Paul D. Hovland, Patrick Heimbach |
Future Gener. Comput. Syst. | 8 |
| 2024 | QuTracer: Mitigating Quantum Gate and Measurement Errors by Tracing Subsets of QubitsabstractQuantum error mitigation plays a crucial role in the current noisy-intermediate-scale-quantum (NISQ) era. As we advance towards achieving a practical quantum advantage in the near term, error mitigation emerges as an indispensable component. One notable prior work, Jigsaw, demonstrates that measurement crosstalk errors can be effectively mitigated by measuring subsets of qubits. Jigsaw operates by running multiple copies of the original circuit, each time measuring only a subset of qubits. The localized distributions yielded from measurement subsetting suffer from less crosstalk and are then used to update the global distribution, thereby achieving improved output fidelity. Inspired by the idea of measurement subsetting, we propose QuTracer, a framework designed to mitigate both gate and measurement errors in subsets of qubits by tracing the states of qubit subsets throughout the computational process. In order to achieve this goal, we introduce a technique, qubit subsetting Pauli checks (QSPC), which utilizes circuit cutting and Pauli Check Sandwiching (PCS) to trace the qubit subsets distribution to mitigate errors. The QuTracer framework can be applied to various algorithms including, but not limited to, VQE, QAOA, quantum arithmetic circuits, QPE, and Hamiltonian simulations. In our experiments, we perform both noisy simulations and real device experiments to demonstrate that QuTracer is scalable and significantly outperforms the state-of-the-art approaches. Peiyi Li 0002, Ji Liu 0007, Alvin Gonzales, Zain H. Saleem, Huiyang Zhou, Paul D. Hovland |
ISCA | 6 |
| 2023 | Model Checking Race-Freedom When "Sequential Consistency for Data-Race-Free Programs" is GuaranteedabstractAbstract Many parallel programming models guarantee that if all sequentially consistent (SC) executions of a program are free of data races, then all executions of the program will appear to be sequentially consistent. This greatly simplifies reasoning about the program, but leaves open the question of how to verify that all SC executions are race-free. In this paper, we show that with a few simple modifications, model checking can be an effective tool for verifying race-freedom. We explore this technique on a suite of C programs parallelized with OpenMP. Jan Hückelheim, Paul D. Hovland, Ziqing Luo, Stephen F. Siegel |
CAV (2) | 3 |
| 2023 | Enhancing Virtual Distillation with Circuit Cutting for Quantum Error MitigationabstractVirtual distillation is a technique that aims to mitigate errors in noisy quantum computers. It works by preparing multiple copies of a noisy quantum state, bridging them through a circuit, and conducting measurements. As the number of copies increases, this process allows for the estimation of the expectation value with respect to a state that approaches the ideal pure state rapidly. However, virtual distillation faces a challenge in realistic scenarios: preparing multiple copies of a quantum state and bridging them through a circuit in a noisy quantum computer will significantly increase the circuit size and introduce excessive noise, which will degrade the performance of virtual distillation. To overcome this challenge, we propose an error mitigation strategy that uses circuit-cutting technology to cut the entire circuit into fragments. With this approach, the fragments responsible for generating the noisy quantum state can be executed on a noisy quantum device, while the remaining fragments are efficiently simulated on a noiseless classical simulator. By running each fragment circuit separately on quantum and classical devices and recombining their results, we can reduce the noise accumulation and enhance the effectiveness of the virtual distillation technique. Our strategy has good scalability in terms of both runtime and computational resources. We demonstrate our strategy’s effectiveness through noisy simulation and experiments on a real quantum device. Peiyi Li 0002, Ji Liu 0007, Hrushikesh Pramod Patil, Paul D. Hovland, Huiyang Zhou |
ICCD | 4 |
| 2023 | Transfer-learning-based Autotuning using Gaussian CopulaabstractAs diverse high-performance computing (HPC) systems are built, many opportunities arise for applications to solve larger problems than ever before. Given the significantly increased complexity of these HPC systems and application tuning, empirical performance tuning, such as autotuning, has emerged as a promising approach in recent years. Despite its effectiveness, autotuning is often a computationally expensive approach. Transfer learning (TL)-based autotuning seeks to address this issue by leveraging the data from prior tuning. Current TL methods for autotuning spend significant time modeling the relationship between parameter configurations and performance, which is ineffective for few-shot (that is, few empirical evaluations) tuning on new tasks. We introduce the first generative TL-based autotuning approach based on the Gaussian copula (GC) to model the high-performing regions of the search space from prior data and then generate high-performing configurations for new tasks. This allows a sampling-based approach that maximizes few-shot performance and provides the first probabilistic estimation of the few-shot budget for effective TL-based autotuning. We compare our generative TL approach with state-of-the-art autotuning techniques on several benchmarks. We find that the GC is capable of achieving 64.37% of peak few-shot performance in its first evaluation. Furthermore, the GC model can determine a few-shot transfer budget that yields up to 33.39× speedup, a dramatic improvement over the 20.58× speedup using prior techniques. Thomas Randall, Jaehoon Koo, Brice Videau, Michael Kruse, Xingfu Wu, Paul D. Hovland, Mary W. Hall, Rong Ge 0002, Prasanna Balaprakash |
ICS | 6 |
| 2023 | QContext: Context-Aware Decomposition for Quantum GatesabstractIn this paper we propose QContext, a new com-piler structure that incorporates context-aware and topology- aware decompositions. Because of circuit equivalence rules and resynthesis, variants of a gate-decomposition template may exist. QContext exploits the circuit information and the hardware topology to select the gate variant that increases circuit optimization opportunities. We study the basis-gate-level context-aware decomposition for Toffoli gates and the native-gate-level context- aware decomposition for CNOT gates. Our experiments show that QContext reduces the number of gates as compared with the state-of-the-art approach, Orchestrated Trios [12]. Ji Liu 0007, Max Bowman, Pranav Gokhale, Siddharth Dangwal, Jeffrey Larson 0001, Fred Chong, Paul D. Hovland |
ISCAS | 7 |
| 2022 | Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler AugmentationabstractDerivatives are key to numerous science, engineering, and machine learning applications. While existing tools generate derivatives of programs in a single language, modern parallel applications combine a set of frameworks and languages to leverage available performance and function in an evolving hardware landscape. We propose a scheme for differentiating arbitrary DAG-based parallelism that preserves scalability and efficiency, implemented into the LLVM-based Enzyme automatic differentiation framework. By integrating with a full-fledged compiler backend, Enzyme can differentiate numerous parallel frameworks and directly control code generation. Combined with its ability to differentiate any LLVM-based language, this flexibility permits Enzyme to leverage the compiler tool chain for parallel and differentiation-specitic optimizations. We differentiate nine distinct versions of the LULESH and miniBUDE applications, written in different programming languages (C++, Julia) and parallel frameworks (OpenMP, MPI, RAJA, Julia tasks, MPI.jl), demonstrating similar scalability to the original program. On benchmarks with 64 threads or nodes, we find a differentiation overhead of 3.4–6.8× on C++ and 5.4–12.5× on Julia. William S. Moses, Sri Hari Krishna Narayanan, Ludger Paehler, Valentin Churavy, Michel Schanen, Jan Hückelheim, Johannes Doerfert, Paul D. Hovland |
SC | 8 |
| 2022 | Verifying Fortran Programs with CIVLabstractAbstract Fortran is widely used in computational science, engineering, and high performance computing. This paper presents an extension to the CIVL verification framework to check correctness properties of Fortran programs. Unlike previous work that translates Fortran to C, LLVM IR, or other intermediate formats before verification, our work allows CIVL to directly consume Fortran source files. We extended the parsing, translation, and analysis phases to support Fortran-specific features such as array slicing and reshaping, and to find program violations that are specific to Fortran, such as argument aliasing rule violations, invalid use of variable and function attributes, or defects due to Fortran’s unspecified expression evaluation order. We demonstrate the usefulness of our tool on a verification benchmark suite and kernels extracted from a real world application. Jan Hückelheim, Paul D. Hovland, Stephen F. Siegel |
TACAS (1) | 3 |
| 2022 | Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimizationabstractAbstract We develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat‐3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd–Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd–Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application. Xingfu Wu, Michael Kruse, Prasanna Balaprakash, Hal Finkel, Paul D. Hovland, Valerie Taylor 0001, Mary W. Hall |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Vector Forward Mode Automatic Differentiation on SIMD/SIMT architecturesabstractAutomatic differentiation, back-propagation, differentiable programming and related methods have received widespread attention, due to their ability to compute accurate gradients of numerical programs for optimization, uncertainty quantification, and machine learning. Two strategies are commonly used. The forward mode, which is easy to implement but has an overhead compared to the original program that grows linearly with the number of inputs, and the reverse mode, which can compute gradients for an arbitrary number of program inputs with a constant factor overhead, although the constant can be large, more memory is required, and the implementation is often challenging. Previous literature has shown that the forward mode can be more easily parallelized and vectorized than the reverse mode, but case studies investigating when either mode is the best choice are lacking, especially for modern CPUs and GPUs. In this paper, we demonstrate that the forward mode can outperform the reverse mode for programs with tens or hundreds of directional derivatives, a number that may yet increase if current hardware trends continue. Jan Hückelheim, Michel Schanen, Sri Hari Krishna Narayanan, Paul D. Hovland |
ICPP | 4 |
| 2019 | Combining Checkpointing and Data Compression to Accelerate Adjoint-Based Optimization Problems
Navjot Kukreja, Jan Hückelheim, Mathias Louboutin, Paul D. Hovland, Gerard Gorman |
Euro-Par | 4 |
| 2019 | Automatic Differentiation for Adjoint Stencil LoopsabstractStencil loops are a common motif in computations including convolutional neural networks, structured-mesh solvers for partial differential equations, and image processing. Stencil loops are easy to parallelise, and their fast execution is aided by compilers, libraries, and domain-specific languages. Reverse-mode automatic differentiation, also known as algorithmic differentiation, autodiff, adjoint differentiation, or back-propagation, is sometimes used to obtain gradients of programs that contain stencil loops. Unfortunately, conventional automatic differentiation results in a memory access pattern that is not stencil-like and not easily parallelisable. Jan Hückelheim, Navjot Kukreja, Sri Hari Krishna Narayanan, Fabio Luporini, Gerard Gorman, Paul D. Hovland |
ICPP | 6 |
| 2018 | Vectorised Computation of Diverging EnsemblesabstractEnsemble computations are used to evaluate a function for multiple inputs, for example in uncertainty quantification. Embedded ensemble computations perform several evaluations within the same program, often enabling a reduced overall runtime by exploiting vectorisation and parallelisation opportunities that are not present in individual ensemble members. This is challenging if members take different control flow paths. We present a source-to-source transformation that turns a given C program into an embedded ensemble program that computes members in a single-instruction-multiple-data fashion using OpenMP SIMD pragmas. We use techniques from whole-function vectorisation, achieving effective vectorisation for moderate amounts of branch divergence, particularly on processors with masked instructions such as recent Xeon Phi or Skylake processors with AVX-512. Jan Hückelheim, Paul D. Hovland, Sri Hari Krishna Narayanan, Paulius Velesko |
ICPP | 2 |
| 2018 | Verifying Properties of Differentiable Programs
Jan Hückelheim, Ziqing Luo, Sri Hari Krishna Narayanan, Stephen F. Siegel, Paul D. Hovland |
SAS | 5 |
| 2015 | Collective I/O Tuning Using Analytical and Machine Learning ModelsabstractThe optimization of parallel I/O has become challenging because of the increasing storage hierarchy, performance variability of shared storage systems, and the number of factors in the hardware and software stacks that impact performance. In this paper, we perform an in-depth study of the complexity involved in I/O autotuning and performance modeling, including the architecture, software stack, and noise. We propose a novel hybrid model combining analytical models for communication and storage operations and black-box models for the performance of the individual operations. The experimental results show that the hybrid approach performs significantly better and shows a higher robustness to noise than state-of-the-art machine learning approaches, at the cost of a higher modeling complexity. Florin Isaila, Prasanna Balaprakash, Stefan M. Wild, Dries Kimpe, Robert Latham, Robert B. Ross, Paul D. Hovland |
CLUSTER | 7 |
| 2015 | Autotuning FPGA Design Parameters for Performance and PowerabstractMany factors affect the performance and power characteristics of FPGA designs. Among them are the optimization parameters for synthesis, map, and place-and-route design tools. Choosing the right combination of these parameters can substantially lower power requirements, while still satisfying timing constraints. Finding such an improvement, however, requires significant experimentation by the FPGA designer. Exhaustive search through the parameter space is an automated alternative to experimentation but is intractable because of the large search space and the long build time of each parameter combination. In this paper, we propose a machine-learning-based approach to tune FPGA design parameters. It performs sampling-based reduction of the parameter space and guides the search toward promising parameter configurations. In our experiments, such selective sampling finds parameter configurations that meet the timing constraints and are within 0.2% of the optimal power consumption. Furthermore, these configurations are found in an order of magnitude less time compared with exhaustive search. Such speedups can substantially alleviate bottlenecks in the FPGA design process. Azamat Mametjanov, Prasanna Balaprakash, Chekuri Choudary, Paul D. Hovland, Stefan M. Wild, Gerald Sabin |
FCCM | 4 |
| 2015 | Generating Efficient Tensor Contractions for GPUsabstractMany scientific and numerical applications, including quantum chemistry modeling and fluid dynamics simulation, require tensor product and tensor contraction evaluation. Tensor computations are characterized by arrays with numerous dimensions, inherent parallelism, moderate data reuse and many degrees of freedom in the order in which to perform the computation. The best-performing implementation is heavily dependent on the tensor dimensionality and the target architecture. In this paper, we map tensor computations to GPUs, starting with a high-level tensor input language and producing efficient CUDA code as output. Our approach is to combine tensor-specific mathematical transformations with a GPU decision algorithm, machine learning and auto tuning of a large parameter space. Generated code shows significant performance gains over sequential and Open MP parallel code, and a comparison with Open ACC shows the importance of auto tuning and other optimizations in our framework for achieving efficient results. Thomas Nelson, Axel Rivera, Prasanna Balaprakash, Mary W. Hall, Paul D. Hovland, Elizabeth R. Jessup, Boyana Norris |
ICPP | 5 |
| 2014 | Energy-performance tradeoffs in multilevel checkpoint strategiesabstractIncreased complexity of computer architectures, consideration of power constraints, and expected failure rates of hardware components make the design and analysis of energy-efficient fault-tolerance schemes an increasingly challenging and important task. We develop run-time and study FTI, a multilevel checkpoint library, on an IBM Blue Gene/Q. We show that FTI has a low energy footprint and that, consequently optimal checkpoint-interval values with respect to time and energy are similar. Leonardo Arturo Bautista-Gomez, Prasanna Balaprakash, Mohamed-Slim Bouguerra, Stefan M. Wild, Franck Cappello, Paul D. Hovland |
CLUSTER | 6 |
| 2013 | Exascale workload characterization and architecture implicationsabstractEmerging exascale architectures bring forth new challenges related to heterogeneous systems power, energy, cost, and resilience. These new challenges require a shift from conventional paradigms in understanding how to best exploit and optimize these features and limitations. Our objective is to identify the top few dominant characteristics in a set of applications. Understanding these characteristics will allow the community to build and exploit customized architectures and tools best suited to optimize each dominant characteristic in the application domain. Every application will typically be composed of multiple characteristics and thus will use several of the customized accelerators and tools during its execution phases, with the eventual goal of using the entire system efficiently. In this poster, we describe a hybrid methodology, based on binary instrumentation, for characterizing scientific applications such as instruction mix and memory access patterns. We apply our methodology to proxy applications that are representative of a broad range of DOE scientific applications. With this empirical basis, we develop and validate statistical models that extrapolate application properties as a function of problem size. These models are then used to project the first quantitative characterization of an exascale computing workload, including computing and memory requirements. We evaluate the potential benefit of processor under memory, a radical new exascale architecture customization and understand how these new customization can impact applications. Prasanna Balaprakash, Darius Buntinas, Apala Guha, Rinku Gupta, Sri Hari Krishna Narayanan, Andrew A. Chien, Paul D. Hovland, Boyana Norris |
ISPASS | 8 |
| 2010 | Speeding up Nek5000 with autotuning and specializationabstractAutotuning technology has emerged recently as a systematic process for evaluating alternative implementations of a computation, in order to select the best-performing solution for a particular architecture. Specialization optimizes code customized to a particular class of input data set. In this paper, we demonstrate how compiler-based autotuning that incorporates specialization for expected data set sizes of key computations can be used to speed up Nek5000, a spectral-element code. Nek5000 makes heavy use of what are effectively Basic Linear Algebra Subroutine (BLAS) calls, but for very small matrices. Through autotuning and specialization, we can achieve significant performance gains over hand-tuned libraries (e.g., Goto, ATLAS, and ACML BLAS). Additional performance gains are obtained from using higher-level compiler optimizations that aggregate multiple BLAS calls. We demonstrate more than 2.2X performance gains on an Opteron over the original manually tuned implementation, and speedups of up to 1.26X on the entire application running on 256 nodes of the Cray XT5 Jaguar system at Oak Ridge. Mary W. Hall, Jacqueline Chame, Chun Chen 0002, Paul F. Fischer, Paul D. Hovland |
ICS | 6 |
| 2009 | Toward adjoinable MPIabstractAutomatic differentiation is the primary means of obtaining analytic derivatives from a numerical model given as a computer program. Therefore, it is an essential productivity tool in numerous computational science and engineering domains. Computing gradients with the adjoint (also called reverse) mode via source transformation is a particularly beneficial but also challenging use of automatic differentiation. To date only ad hoc solutions for adjoint differentiation of MPI programs have been available, forcing automatic differentiation tool users to reason about parallel communication dataflow and dependencies and manually develop adjoint communication code. Using the communication graph as a model we characterize the principal problems of adjoining the most frequently used communication idioms. We propose solutions to cover these idioms and consider the consequences for the MPI implementation, the MPI user and MPI-aware program analysis. The MIT general circulation model serves as a use case to illustrate the viability of our approach. Jean Utke, Laurent Hascoët, Patrick Heimbach, Chris Hill, Paul D. Hovland, Uwe Naumann |
IPDPS | 5 |
| 2006 | Data-Flow Analysis for MPI ProgramsabstractMessage passing via MPI is widely used in single-program, multiple-data (SPMD) parallel programs. Existing data-flow frameworks do not model the semantics of message-passing SPMD programs, which can result in less precise and even incorrect analysis results. We present a data-flow analysis framework for performing interprocedural analysis of message-passing SPMD programs. The framework is based on the MPI-ICFG representation, which is an interprocedural control-flow graph (ICFG) augmented with communication edges between possible send and receive pairs and partial context sensitivity. We show how to formulate nonseparable data-flow analyses within our framework using reaching constants as a canonical example. We also formulate and provide experimental results for the nonseparable analysis, activity analysis. Activity analysis is a domain-specific analysis used to reduce the computation and storage requirements for automatically differentiated MPI programs. Automatic differentiation is important for application domains such as climate modeling, electronic device simulation, oil reservoir simulation, medical treatment planning and computational economics to name a few. Our experimental results show that using the MPI-ICFG data-flow analysis framework improves the precision of activity analysis and as a result significantly reduces memory requirements for the automatically differentiated versions of a set of parallel benchmarks, including some of the NAS parallel benchmarks Michelle Mills Strout, Barbara Kreaseck, Paul D. Hovland |
ICPP | 3 |
| 2005 | Representation-independent program analysisabstractProgram analysis has many applications in software engineering and high-performance computation, such as program understanding, debugging, testing, reverse engineering, and optimization. A ubiquitous compiler infrastructure does not exist; therefore, program analysis is essentially reimplemented for each compiler infrastructure. The goal of the OpenAnalysis toolkit is to separate analysis from the intermediate representation (IR) in a way that allows the orthogonal development of compiler infrastructures and program analysis. Separation of analysis from specific IRs will allow faster development of compiler infrastructures, the ability to share and compare analysis implementations, and in general quicker breakthroughs and evolution in the area of program analysis. This paper presents how we are separating analysis implementations from IRs with analysis-specific, IR-independent interfaces. Analysis-specific IR interfaces for alias/pointer analysis algorithms and reaching constants illustrate that an IR interface designed for language dependence is capable of providing enough information to support the implementation of a broad range of analysis algorithms and also represent constructs within many imperative programming languages. Michelle Mills Strout, John M. Mellor-Crummey, Paul D. Hovland |
PASTE | 3 |
| 2005 | Making automatic differentiation truly automatic: coupling PETSc with ADIC
Paul D. Hovland, Boyana Norris, Barry Smith 0002 |
Future Gener. Comput. Syst. | 1 |
| 2002 | Implementation of automatic differentiation toolsabstractAutomatic differentiation is a semantic transformation that applies the rules of differential calculus to source code. It thus transforms a computer program that computes a mathematical function into a program that computes the function and its derivatives. Derivatives play an important role in a wide variety of scientific computing applications, including optimization, solution of nonlinear equations, sensitivity analysis, and nonlinear inverse problems. We describe a simple component architecture for developing tools for automatic differentiation and other mathematically oriented semantic transformations of scientific software. This architecture consists of a compiler-based, language-specific front-end for source transformation, loosely coupled with one or more language-independent "plug-in" transformation modules. The coupling mechanism between the front-end and transformation modules is provided by the XML Abstract Interface Form (XAIF). XAIF provides an abstract, language-independent representation of language constructs common in imperative languages, such as C and Fortran. We describe the use of this architecture in constructing tools for automatic differentiation of Fortran 77 and ANSI C, and we discuss how access to compiler optimization techniques can enable more efficient derivative augmentation. Christian H. Bischof, Paul D. Hovland, Boyana Norris |
PEPM | 2 |
| 2002 | Parallel components for PDEs and optimization: some issues and experiences
Boyana Norris, Satish Balay, Steven Benson, Lori A. Diachin, Paul D. Hovland, Lois C. McInnes, Barry Smith 0002 |
Parallel Comput. | 5 |
| 2001 | A Distributed Application Server for Automatic Differentiation
Boyana Norris, Paul D. Hovland |
IPDPS | 2 |
| 2001 | Parallel simulation of compressible flow using automatic differentiation and PETSc
Paul D. Hovland, Lois C. McInnes |
Parallel Comput. | 1 |
| 2000 | On Combining Computational Differentiation and Toolkits for Parallel Scientific Computing
Christian H. Bischof, H. Martin Bücker, Paul D. Hovland |
Euro-Par | 3 |
| 1993 | A Model for Automatic Dta PartitioningabstractIn order to efficiently exploit global parallelism, it is essential to find a good way to distribute data among the processors in distributed-memory parallel computer systems. A formal technique utilizing augmented data access descriptors (ADADs) to determine this distribution is presented. This technique differs from previous approaclies in that it views the problem of finding a good distribution as an extension of data dependence analysis. The importance of this difference is demonstrated through an explanation of how ADADs facilitate interprocedural analysis, directed loop transformations, and incremental analysis, which may lead to improvements in the eficieiicy of both program developn~enta nd the program itself. Paul D. Hovland, Lionel M. Ni |
ICPP (2) | 1 |