Steve Margerm

dblp:199/7194 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 27% GPUs and heterogeneous computing · 27% Hardware accelerators and domain-specific architectures · 23%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU programming
dynamic parallelism
0.312018
TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018
Electronic design automation
high-level synthesis
0.312018
TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018
Hardware accelerators and domain-specific architectures
accelerator offloading
0.312017
Needle: Leveraging Program Analysis to Analyze and Extract Accelerators from Whole Programs · HPCA 2017
Parallel and multicore computing › parallel programming models › task parallelism
fork-join parallelism
0.112018
TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018
Parallel and multicore computing
parallel programming models
0.112018
TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.112017
Needle: Leveraging Program Analysis to Analyze and Extract Accelerators from Whole Programs · HPCA 2017

Methods — techniques the papers use, named apart from their topics

program analysis · 0.6dynamic profiling · 0.6LLVM · 0.6high-level synthesis · 0.3compiler intermediate representation · 0.3
YearPublicationVenuePosition
2018 TAPAS: Generating Parallel Accelerators from Parallel Programs
abstract
High-level-synthesis (HLS) tools generate accelerators from software programs to ease the task of building hardware. Unfortunately, current HLS tools have limited support for concurrency, which impacts the speedup achievable with the generated accelerator. Current approaches only target fixed static patterns (e.g., pipeline, data-parallel kernels). This constraints the ability of software programmers to express concurrency. Moreover, the generated accelerator loses a key benefit of parallel hardware, dynamic asynchrony, and the potential to hide long latency and cache misses. We have developed TAPAS, an HLS toolchain for generating parallel accelerators from programs with dynamic parallelism. TAPAS is built on top of Tapir [22], [39], which embeds fork-join parallelism into the compiler's intermediate-representation. TAPAS leverages the compiler IR to identify parallelism and synthesizes the hardware logic. TAPAS provides first-class architecture support for spawning, coordinating and synchronizing tasks during accelerator execution. We demonstrate TAPAS can generate accelerators for concurrent programs with heterogeneous, nested and recursive parallelism. Our evaluation on Intel-Altera DE1-SoC and Arria-10 boards demonstrates that TAPAS generated accelerators achieve 20× the power efficiency of an Intel Xeon, while maintaining comparable performance. We also show that TAPAS enables lightweight tasks that can be spawned in '10 cycles and enables accelerators to exploit available fine-grain parallelism. TAPAS is a complete HLS toolchain for synthesizing parallel programs to accelerators and is open-sourced.
Steve Margerm, Amirali Sharifian, Apala Guha, Arrvindh Shriraman, Gilles Pokam
MICRO1
2017 Needle: Leveraging Program Analysis to Analyze and Extract Accelerators from Whole Programs
abstract
Technology constraints have increasingly led to the adoption of specialized coprocessors, i.e. hardware accelerators. The first challenge that computer architects encounter is identifying “what to specialize in the program”. We demonstrate that this requires precise enumeration of program paths based on dynamic program behavior. We hypothesize that path-based [4] accelerator offloading leads to good coverage of dynamic instructions and improve energy efficiency. Unfortunately, hot paths across programs demonstrate diverse control flow behavior. Accelerators (typically based on dataflow execution), often lack an energy-efficient, complexity effective, and high performance (eg. branch prediction) support for control flow. We have developed NEEDLE, an LLVM based compiler framework that leverages dynamic profile information to identify, merge, and offload acceleratable paths from whole applications. NEEDLE derives insight into what code coverage (and consequently energy reduction) an accelerator can achieve. We also develop a novel program abstraction for offload calledBraid, that merges common code regions across different paths to improve coverage of the accelerator while trading off the increase in dataflow size. This enables coarse grained offloading, reducing interaction with the host CPU core. To prepare the Braids and paths for acceleration, NEEDLE generates software frames. Software frames enable energy efficient speculative execution on accelerators. They are accelerator microarchitecture independent support speculative execution including memory operations. NEEDLE is automated and has been used to analyze 225K paths across 29 workloads. It filtered and ranked 154K paths for acceleration across unmodified SPEC, PARSEC and PERFECT workload suites. We target NEEDLE's offload regions toward a CGRA and demonstrate 34% performance and 20% energy improvement.
Snehasish Kumar, William N. Sumner, Vijayalakshmi Srinivasan, Steve Margerm, Arrvindh Shriraman
HPCA4