Matthew Feldman

dblp:201/4861 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-2494-6533ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4Systems, architecture and hardware · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 48% Hardware accelerators and domain-specific architectures · 29% Electronic design automation · 17%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Artificial intelligence
1 paper
Optimization for machine learning · 77% Efficient and distributed learning · 23%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
hardware compilation
0.412020
Type-directed scheduling of streaming accelerators · PLDI 2020
Hardware accelerators and domain-specific architectures › spatial architecture
data streaming accelerator
0.412020
Type-directed scheduling of streaming accelerators · PLDI 2020
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.422018
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017
Spatial: a language and compiler for application accelerators · PLDI 2018
Compilers and program optimization › code generation › parallel code generation
accelerator code generation
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Compilers and program optimization › hardware compilation
high-level synthesis
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Reconfigurable computing and FPGAs
reconfigurable computing
0.312018
Spatial: a language and compiler for application accelerators · PLDI 2018
Machine learning › Optimization for machine learning
stochastic gradient descent
0.312017
Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent · ISCA 2017
Electronic design automation
high-level synthesis
0.112020
Type-directed scheduling of streaming accelerators · PLDI 2020
Electronic design automation › high-level synthesis
scheduling
0.112020
Type-directed scheduling of streaming accelerators · PLDI 2020
Machine learning › Efficient and distributed learning
low-precision training
0.112017
Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent · ISCA 2017
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112017
Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017

Methods — techniques the papers use, named apart from their topics

type-directed compilation · 0.9static scheduling · 0.9low-precision arithmetic · 0.3asynchronous SGD · 0.3
YearPublicationVenuePosition
2021 High performance lattice regression on FPGAs via a high level hardware description language
abstract
Lattice regression-based models are highly-constrainable and interpretable machine learning models used in applications such as query classification and path length prediction for maps. To improve their performance and better serve these models to millions of consumers, we accelerate them using field programmable gate arrays. We adopt a library-based approach using a high level hardware description language (HLHDL) to support the broad family of lattice models. HLHDLs improve productivity by providing both control abstraction such as looping, reductions, and memory hierarchies, as well as automatically handling low-level tasks such as retiming. However, these abstractions can lead to performance bottlenecks if not carefully used. We characterize these bottlenecks and implement a lattice regression library using a streaming tensor abstraction which avoids them. On a pair of models trained for network anomaly detection, we achieve a${166\,-\,256\times}$speedup over CPUs even with large batch sizes.
Nathan Zhang, Matthew Feldman, Kunle Olukotun
FPT2
2020 Type-directed scheduling of streaming accelerators
abstract
Designing efficient, application-specialized hardware accelerators requires assessing trade-offs between a hardware module’s performance and resource requirements. To facilitate hardware design space exploration, we describe Aetherling, a system for automatically compiling data-parallel programs into statically scheduled, streaming hardware circuits. Aetherling contributes a space- and time-aware intermediate language featuring data-parallel operators that represent parallel or sequential hardware modules, and sequence data types that encode a module’s throughput by specifying when sequence elements are produced or consumed. As a result, well-typed operator composition in the space-time language corresponds to connecting hardware modules via statically scheduled, streaming interfaces.
David Durst, Matthew Feldman, Dillon Huff, David Akeley, Ross Daly, Gilbert Louis Bernstein, Marco Patrignani, Kayvon Fatahalian, Pat Hanrahan
PLDI2
2018 Spatial: a language and compiler for application accelerators
abstract
Industry is increasingly turning to reconfigurable architectures like FPGAs and CGRAs for improved performance and energy efficiency. Unfortunately, adoption of these architectures has been limited by their programming models. HDLs lack abstractions for productivity and are difficult to target from higher level languages. HLS tools are more productive, but offer an ad-hoc mix of software and hardware abstractions which make performance optimizations difficult.
David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang 0001, Stefan Hadjis, Ruben Fiszel, Tian Zhao 0001, Luigi Nardi, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
PLDI2
2017 Plasticine: A Reconfigurable Architecture For Parallel Paterns
abstract
Reconfigurable architectures have gained popularity in recent years as they allow the design of energy-efficient accelerators. Fine-grain fabrics (e.g. FPGAs) have traditionally suffered from performance and power inefficiencies due to bit-level reconfigurable abstractions. Both fine-grain and coarse-grain architectures (e.g. CGRAs) traditionally require low level programming and suffer from long compilation times. We address both challenges with Plasticine, a new spatially reconfigurable architecture designed to efficiently execute applications composed of parallel patterns. Parallel patterns have emerged from recent research on parallel programming as powerful, high-level abstractions that can elegantly capture data locality, memory access patterns, and parallelism across a wide range of dense and sparse applications.
Raghu Prabhakar, Yaqi Zhang 0001, David Koeplinger, Matthew Feldman, Tian Zhao 0001, Stefan Hadjis, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA4
2017 Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent
Christopher De Sa, Matthew Feldman, Christopher Ré, Kunle Olukotun
ISCA2