Eric Hall

dblp:84/2 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 90% Memory systems · 6% Hardware accelerators and domain-specific architectures · 5%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
vector processor
1.222023
EVE: Ephemeral Vector Engines · HPCA 2023
big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip · MICRO 2022
Processor architecture and microarchitecture › vector processor
vector processing unit
0.712023
EVE: Ephemeral Vector Engines · HPCA 2023
Processor architecture and microarchitecture › multicore design
big.LITTLE
0.612022
big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip · MICRO 2022
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.612022
big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip · MICRO 2022
Memory systems
processing-in-memory
0.212023
EVE: Ephemeral Vector Engines · HPCA 2023
Hardware accelerators and domain-specific architectures
data-parallel acceleration
0.212022
big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip · MICRO 2022

Methods — techniques the papers use, named apart from their topics

on-demand core reconfiguration · 0.6
YearPublicationVenuePosition
2023 EVE: Ephemeral Vector Engines
abstract
There has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead.
Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady O. Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, Christopher Batten
HPCA6
2022 big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip
abstract
Single-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads.
Tuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou, Eric Hall, Courtney Golden, Christopher Batten
MICRO5
2011 Built-In Functional Tests for Silicon Validation and System Integration of Telecom SoC Designs
abstract
Existing silicon validation techniques only address test data capture issues. They all assume the existence of live traffic in the system. Unfortunately, this is not always the case in real life. This paper proposes a novel design methodology for silicon validation and system integration. It uses built-in functional tests to simulate live traffic at full speed when a real one is not available at the arrival of the first silicon. The proposed methodology provides a platform upon which many silicon validation and system integration tasks can be performed before a real traffic is ready. It can also be used to cover logic corner cases that may not be easily achievable in real life. The proposed methodology has been proven effective on time-to-market and quality of verification with multiple complex system-on-chip designs.
Yuejian Wu, Sandy Thomson, Dale Mutcher, Eric Hall
IEEE Trans. Very Large Scale Integr. Syst.4