Pete Ehrett

dblp:229/4367 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-4587-3716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GPUs All Grown-Up: Fully Device-Driven SpMV Using GPU Work Graphs
abstract
Sparse matrix-vector multiplication (SpMV) is a key operation across high-performance computing, graph analytics, and many more applications.In these applications, the matrix characteristics, notably non-zero elements per row, can vary widely and impact which algorithm performs best.Thus, Graphics Processing Unit (GPU) SpMV algorithms often rely on costly preprocessing to determine what per-row algorithm to select to achieve high performance.In this work we combine SpMV preprocessing and the subsequent per-row processing on the GPU by leveraging the novel "Work Graphs" GPU programming model-initially designed for graphics applications-for dynamic on-device self-scheduling.Work Graphs allow for fine-grain dataflow execution of individual workgroups using emerging hardware and firmware support.As soon as preprocessing has generated sufficient work, workgroups of individual processing kernels are self-scheduled and executed, interleaved with those of other kernels.This improves cache locality and eliminates host interaction altogether.Across a suite of 59 sparse matrices, the best of various novel Work Graphs SpMV implementations outperforms state-of-the-art rocSPARSE "LRB" for a single SpMV by up to 7.19× (mean: 3.35×, SD: 1.89).Furthermore, it achieves much more stable performance across various sparsity patterns than the rocSPARSE CSR-General algorithm, and even beats the advanced rocSPARSE CSR-Adaptive algorithm for up to 92 consecutive SpMV calculations.In addition, compared to rocSPARSE LRB, it reduces code complexity by 75%.Its memory footprint for supporting data structures is a fixed ∼25 MiB independent of matrix size, compared to rocSPARSE LRB's data structures that scale with matrix
Fabian Wildgrube, Pete Ehrett, Paul Trojahn, Richard Membarth, Bradford M. Beckmann, Dominik Baumeister, Matthäus G. Chajdas
ISCA2
2021 Chopin: Composing Cost-Effective Custom Chips with Algorithmic Chiplets
abstract
As computational demands rise, the need for specialized hardware has grown acute. However, the immense cost of fully-custom chips has forced many developers to rely on suboptimal solutions like FPGAs, especially for low- to mid-volume applications, in which multi-million-dollar non-recurring engineering (NRE) costs cannot be amortized effectively. We propose to address this problem by composing custom chips out of small, algorithmic chiplets, reusable across diverse designs, such that high NRE costs may be amortized across many different designs. This work models the economics of this paradigm and identifies a cost-optimal granularity for algorithmic chiplets, then demonstrates how those guidelines may be applied to design high-performance, algorithmically-composable hardware components – which may be reused, without modification, across many different processing pipelines. For an example phased-array radar accelerator, our chiplet-centric paradigm improves perf-per-$ by 9.3× over an FPGA, and ∼4× over a conventional ASIC.
Pete Ehrett, Todd M. Austin, Valeria Bertacco
ICCD1
2021 A Defense-Inspired Benchmark Suite
abstract
This work previews the MilSpec suite, a diverse collection of benchmarks targeting heterogeneous embedded/edge systems. These end-to-end and kernel-level benchmarks are inspired by defense-related applications and applicable to the broader system design space. They exercise a wide range of computational capabilities, such as signal processing, image processing/computer vision, and matrix-based computation, which are particularly relevant for modern embedded applications that must process large amounts of data at the edge.
Pete Ehrett, Nathan Block, Bing Schaefer, Adrian Berding, John Paul Koenig, Pranav Srinivasan, Valeria Bertacco, Todd M. Austin
ISPASS1
2019 SiPterposer: A Fault-Tolerant Substrate for Flexible System-in-Package Design
abstract
As Moore's Law scaling slows down, specialized heterogeneous designs are needed to sustain computing performance improvements. Unfortunately, the non-recurring engineering (NRE) costs of chip design-designing interconnects, creating masks, etc.-are often prohibitive. Chiplet-based disintegrated design solutions could address these economic issues, but current technologies lack the flexibility to express a rich variety of designs without redesigning the communication substrate. Moreover, as the number of chiplets increases, yield suffers due to 2.5D assembly defects. This work addresses these problems by presenting a flexible communication fabric that supports construction of arbitrary network topologies and provides robust fault-tolerance, demonstrating near-100% chip assembly yield at typical bonding defect rates. We achieve these goals with less than 3% additional power and zero exposed latency overhead for various real-world applications running on an example SiP.
Pete Ehrett, Todd M. Austin, Valeria Bertacco
DATE1
2018 SWAN: mitigating hardware trojans with design ambiguity
abstract
For the past decade, security experts have warned that malicious engineers could modify hardware designs to include hardware back-doors (trojans), which, in turn, could grant attackers full control over a system. Proposed defenses to detect these attacks have been outpaced by the development of increasingly small, but equally dangerous, trojans. To thwart trojan-based attacks, we propose a novel architecture that maps the security-critical portions of a processor design to a one-time programmable, LUT-free fabric. The programmable fabric is automatically generated by analyzing the HDL of targeted modules. We present our tools to generate the fabric and map functionally equivalent designs onto the fabric. By having a trusted party randomly select a mapping and configure each chip, we prevent an attacker from knowing the physical location of targeted signals at manufacturing time. In addition, we provide decoy options (canaries) for the mapping of security-critical signals, such that hardware trojans hitting a decoy are thwarted and exposed. Using this defense approach, any trojan capable of analyzing the entire configurable fabric must employ complex logic functions with a large silicon footprint, thus exposing it to detection by inspection. We evaluated our solution on a RISC-V BOOM processor and demonstrated that, by providing the ability to map each critical signal to 6 distinct locations on the chip, we can reduce the chance of attack success by an undetectable trojan by 99%, incurring only a 27% area overhead.
Timothy Linscott, Pete Ehrett, Valeria Bertacco, Todd M. Austin
ICCAD2