VLDB 2026 Research / reviewers in the wild / expert
Matthew Feldman
dblp:201/4861
· DBLP profile ↗
5ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-2494-6533ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4Systems, architecture and hardware · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Reconfigurable computing and FPGAs · 48% Hardware accelerators and domain-specific architectures · 29% Electronic design automation · 17% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Optimization for machine learning · 77% Efficient and distributed learning · 23% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
hardware compilation |
0.4 | 1 | 2020 | Type-directed scheduling of streaming accelerators · PLDI 2020 |
Hardware accelerators and domain-specific architectures › spatial architecture
data streaming accelerator |
0.4 | 1 | 2020 | Type-directed scheduling of streaming accelerators · PLDI 2020 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.4 | 2 | 2018 | Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017 Spatial: a language and compiler for application accelerators · PLDI 2018 |
Compilers and program optimization › code generation › parallel code generation
accelerator code generation |
0.3 | 1 | 2018 | Spatial: a language and compiler for application accelerators · PLDI 2018 |
Compilers and program optimization › hardware compilation
high-level synthesis |
0.3 | 1 | 2018 | Spatial: a language and compiler for application accelerators · PLDI 2018 |
Reconfigurable computing and FPGAs
reconfigurable computing |
0.3 | 1 | 2018 | Spatial: a language and compiler for application accelerators · PLDI 2018 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.3 | 1 | 2017 | Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent · ISCA 2017 |
Electronic design automation
high-level synthesis |
0.1 | 1 | 2020 | Type-directed scheduling of streaming accelerators · PLDI 2020 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2020 | Type-directed scheduling of streaming accelerators · PLDI 2020 |
Machine learning › Efficient and distributed learning
low-precision training |
0.1 | 1 | 2017 | Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent · ISCA 2017 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2017 | Plasticine: A Reconfigurable Architecture For Parallel Paterns · ISCA 2017 |
Methods — techniques the papers use, named apart from their topics
type-directed compilation · 0.9static scheduling · 0.9low-precision arithmetic · 0.3asynchronous SGD · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | High performance lattice regression on FPGAs via a high level hardware description languageabstractLattice regression-based models are highly-constrainable and interpretable machine learning models used in applications such as query classification and path length prediction for maps. To improve their performance and better serve these models to millions of consumers, we accelerate them using field programmable gate arrays. We adopt a library-based approach using a high level hardware description language (HLHDL) to support the broad family of lattice models. HLHDLs improve productivity by providing both control abstraction such as looping, reductions, and memory hierarchies, as well as automatically handling low-level tasks such as retiming. However, these abstractions can lead to performance bottlenecks if not carefully used. We characterize these bottlenecks and implement a lattice regression library using a streaming tensor abstraction which avoids them. On a pair of models trained for network anomaly detection, we achieve a${166\,-\,256\times}$speedup over CPUs even with large batch sizes. Nathan Zhang, Matthew Feldman, Kunle Olukotun |
FPT | 2 |
| 2020 | Type-directed scheduling of streaming acceleratorsabstractDesigning efficient, application-specialized hardware accelerators requires assessing trade-offs between a hardware module’s performance and resource requirements. To facilitate hardware design space exploration, we describe Aetherling, a system for automatically compiling data-parallel programs into statically scheduled, streaming hardware circuits. Aetherling contributes a space- and time-aware intermediate language featuring data-parallel operators that represent parallel or sequential hardware modules, and sequence data types that encode a module’s throughput by specifying when sequence elements are produced or consumed. As a result, well-typed operator composition in the space-time language corresponds to connecting hardware modules via statically scheduled, streaming interfaces. David Durst, Matthew Feldman, Dillon Huff, David Akeley, Ross Daly, Gilbert Louis Bernstein, Marco Patrignani, Kayvon Fatahalian, Pat Hanrahan |
PLDI | 2 |
| 2018 | Spatial: a language and compiler for application acceleratorsabstractIndustry is increasingly turning to reconfigurable architectures like FPGAs and CGRAs for improved performance and energy efficiency. Unfortunately, adoption of these architectures has been limited by their programming models. HDLs lack abstractions for productivity and are difficult to target from higher level languages. HLS tools are more productive, but offer an ad-hoc mix of software and hardware abstractions which make performance optimizations difficult. David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang 0001, Stefan Hadjis, Ruben Fiszel, Tian Zhao 0001, Luigi Nardi, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun |
PLDI | 2 |
| 2017 | Plasticine: A Reconfigurable Architecture For Parallel PaternsabstractReconfigurable architectures have gained popularity in recent years as they allow the design of energy-efficient accelerators. Fine-grain fabrics (e.g. FPGAs) have traditionally suffered from performance and power inefficiencies due to bit-level reconfigurable abstractions. Both fine-grain and coarse-grain architectures (e.g. CGRAs) traditionally require low level programming and suffer from long compilation times. We address both challenges with Plasticine, a new spatially reconfigurable architecture designed to efficiently execute applications composed of parallel patterns. Parallel patterns have emerged from recent research on parallel programming as powerful, high-level abstractions that can elegantly capture data locality, memory access patterns, and parallelism across a wide range of dense and sparse applications. Raghu Prabhakar, Yaqi Zhang 0001, David Koeplinger, Matthew Feldman, Tian Zhao 0001, Stefan Hadjis, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun |
ISCA | 4 |
| 2017 | Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent
Christopher De Sa, Matthew Feldman, Christopher Ré, Kunle Olukotun |
ISCA | 2 |