David Koeplinger

dblp:172/1102 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 59% Electronic design automation · 36% Energy-efficient computing · 4%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
0.522016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Generating Configurable Hardware from Parallel Patterns · ASPLOS 2016
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.422018
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017
Spatial: a language and compiler for application accelerators · PLDI 2018
Compilers and program optimization › code generation › parallel code generation
accelerator code generation
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Compilers and program optimization › hardware compilation
high-level synthesis
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Reconfigurable computing and FPGAs
reconfigurable computing
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Electronic design automation
design space exploration
0.212016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Reconfigurable computing and FPGAs
FPGA accelerator
0.212016
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware · ISCA 2016
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA accelerator design
0.212016
Generating Configurable Hardware from Parallel Patterns · ASPLOS 2016
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112017
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017
Compilers and program optimization
hardware compilation
0.112016
Generating Configurable Hardware from Parallel Patterns · ASPLOS 2016

Methods — techniques the papers use, named apart from their topics

parallel patterns · 0.8high-level synthesis · 0.5parameterized templates · 0.2artificial neural network · 0.2
YearPublicationVenuePosition
2019 Practical Design Space Exploration
abstract
Multi-objective optimization is a crucial matter in computer systems design space exploration because real-world applications often rely on a trade-off between several objectives. Derivatives are usually not available or impractical to compute and the feasibility of an experiment can not always be determined in advance. These problems are particularly difficult when the feasible region is relatively small, and it may be prohibitive to even find a feasible experiment, let alone an optimal one. We introduce a new methodology and corresponding software framework, HyperMapper 2.0, which handles multi-objective optimization, unknown feasibility constraints, and categorical/ordinal variables. This new methodology also supports injection of the user prior knowledge in the search when available. All of these features are common requirements in computer systems but rarely exposed in existing design space exploration systems. The proposed methodology follows a white-box model which is simple to understand and interpret (unlike, for example, neural networks) and can be used by the user to better understand the results of the automatic search. We apply and evaluate the new methodology to the automatic static tuning of hardware accelerators within the recently introduced Spatial programming language, with minimization of design run-time and compute logic under the constraint of the design fitting in a target field-programmable gate array chip. Our results show that HyperMapper 2.0 provides better Pareto fronts compared to state-of-the-art baselines, with better or competitive hypervolume indicator and with 8x improvement in sampling budget for most of the benchmarks explored.
Luigi Nardi, David Koeplinger, Kunle Olukotun
MASCOTS2
2019 HyperMapper: a Practical Design Space Exploration Framework
abstract
Design problems are ubiquitous in scientific and industrial achievements. Scientists design experiments to gain insights into physical and social phenomena, and engineers design machines to execute tasks more efficiently. These design problems are fraught with choices which are often complex and high-dimensional and which include interactions that make them difficult for individuals to reason about. In software/hardware co-design, for example, companies develop libraries with tens or hundreds of free choices and parameters that interact in complex ways. In fact, the level of complexity is often so high that it becomes impossible to find domain experts capable of tuning these libraries [1].
Luigi Nardi, Artur L. F. Souza, David Koeplinger, Kunle Olukotun
MASCOTS3
2018 Spatial: a language and compiler for application accelerators
abstract
Industry is increasingly turning to reconfigurable architectures like FPGAs and CGRAs for improved performance and energy efficiency. Unfortunately, adoption of these architectures has been limited by their programming models. HDLs lack abstractions for productivity and are difficult to target from higher level languages. HLS tools are more productive, but offer an ad-hoc mix of software and hardware abstractions which make performance optimizations difficult.
David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang 0001, Stefan Hadjis, Ruben Fiszel, Tian Zhao 0001, Luigi Nardi, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
PLDI1
2017 Plasticine: A Reconfigurable Architecture For Parallel Paterns
abstract
Reconfigurable architectures have gained popularity in recent years as they allow the design of energy-efficient accelerators. Fine-grain fabrics (e.g. FPGAs) have traditionally suffered from performance and power inefficiencies due to bit-level reconfigurable abstractions. Both fine-grain and coarse-grain architectures (e.g. CGRAs) traditionally require low level programming and suffer from long compilation times. We address both challenges with Plasticine, a new spatially reconfigurable architecture designed to efficiently execute applications composed of parallel patterns. Parallel patterns have emerged from recent research on parallel programming as powerful, high-level abstractions that can elegantly capture data locality, memory access patterns, and parallelism across a wide range of dense and sparse applications.
Raghu Prabhakar, Yaqi Zhang 0001, David Koeplinger, Matthew Feldman, Tian Zhao 0001, Stefan Hadjis, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA3
2016 Generating Configurable Hardware from Parallel Patterns
abstract
In recent years the computing landscape has seen an increasing shift towards specialized accelerators. Field programmable gate arrays (FPGAs) are particularly promising for the implementation of these accelerators, as they offer significant performance and energy improvements over CPUs for a wide class of applications and are far more flexible than fixed-function ASICs. However, FPGAs are difficult to program. Traditional programming models for reconfigurable logic use low-level hardware description languages like Verilog and VHDL, which have none of the productivity features of modern software languages but produce very efficient designs, and low-level software languages like C and OpenCL coupled with high-level synthesis (HLS) tools that typically produce designs that are far less efficient. Functional languages with parallel patterns are a better fit for hardware generation because they provide high-level abstractions to programmers with little experience in hardware design and avoid many of the problems faced when generating hardware from imperative languages. In this paper, we identify two important optimizations for using parallel patterns to generate efficient hardware: tiling and metapipelining. We present a general representation of tiled parallel patterns, and provide rules for automatically tiling patterns and generating metapipelines. We demonstrate experimentally that these optimizations result in speedups up to 39.4× on a set of benchmarks from the data analytics domain.
Raghu Prabhakar, David Koeplinger, Kevin J. Brown, HyoukJoong Lee, Christopher De Sa, Christoforos E. Kozyrakis, Kunle Olukotun
ASPLOS2
2016 Automatic Generation of Efficient Accelerators for Reconfigurable Hardware
abstract
Acceleration in the form of customized datapaths offer large performance and energy improvements over general purpose processors. Reconfigurable fabrics such as FPGAs are gaining popularity for use in implementing application-specific accelerators, thereby increasing the importance of having good high-level FPGA design tools. However, current tools for targeting FPGAs offer inadequate support for high-level programming, resource estimation, and rapid and automatic design space exploration. We describe a design framework that addresses these challenges. We introduce a new representation of hardware using parameterized templates that captures locality and parallelism information at multiple levels of nesting. This representation is designed to be automatically generated from high-level languages based on parallel patterns. We describe a hybrid area estimation technique which uses template-level models and design-level artificial neural networks to account for effects from hardware place-and-route tools, including routing overheads, register and block RAM duplication, and LUT packing. Our runtime estimation accounts for off-chip memory accesses. We use our estimation capabilities to rapidly explore a large space of designs across tile sizes, parallelization factors, and optional coarse-grained pipelining, all at multiple loop levels. We show that estimates average 4.8% error for logic resources, 6.1% error for runtimes, and are 279 to 6533 times faster than a commercial high-level synthesis tool. We compare the best-performing designs to optimized CPU code running on a server-grade 6 core processor and show speedups of up to 16.7×.
David Koeplinger, Raghu Prabhakar, Yaqi Zhang 0001, Christina Delimitrou, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA1