Olalekan Afuye

dblp:283/1082 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2023
0000-0001-7244-7037ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 88% Memory systems · 12%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › vector processor
vector processing unit
0.712023
EVE: Ephemeral Vector Engines · HPCA 2023
Processor architecture and microarchitecture
vector processor
0.712023
EVE: Ephemeral Vector Engines · HPCA 2023
Memory systems
processing-in-memory
0.212023
EVE: Ephemeral Vector Engines · HPCA 2023
YearPublicationVenuePosition
2023 EVE: Ephemeral Vector Engines
abstract
There has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead.
Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady O. Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, Christopher Batten
HPCA5
2022 LO Synchronization Scheme via Full-Duplex Transceiver for Distributed Beamforming in Wireless Ad hoc Networks
abstract
In this paper, we demonstrate a prototype system for path independent local oscillator (LO) synchronization of a distributed beamformer in wireless ad hoc networks. The system contains a low power full duplex (FD) transceiver IC, a RF phase interpolator IC, and a CDMA encoder/decoder to realize a conjugate loop and synchronize the LO’s of two RF nodes. Both ICs were fabricated in 180nm CMOS technology. The FD transceiver IC consumes 69mW at 700MHz and the RF phase interpolator IC consumes 75mW at 1.4GHz. Using the low power ICs, we demonstrate a simple, lightweight, and robust methodology to synchronize two LO’s with an average phase precision of 2.1° and 94% maximum beamforming gain using RF only transmissions via a single antenna per node.
Olalekan Afuye, Shimin Huang, Ken Ho, Alyosha C. Molnar, Alyssa B. Apsel
ISCAS1
2020 Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator using In-Situ Processing-In-SRAM
abstract
Vector accelerators can efficiently execute regular data-parallel workloads, but they require expensive multi-ported register files to feed large vector ALUs. Recent work on in-situ processing-in-SRAM shows promise in enabling area-efficient vector acceleration. This work explores two different approaches to leveraging in-situ processing-in-SRAM: BS-VRAM, which uses bit-serial execution, and BP-VRAM, which uses bit-parallel execution. The two approaches have very different latency vs. throughput trade-offs. BS-VRAM requires more cycles per operation, but is able to execute thousands of operations in parallel, while BP-VRAM requires fewer cycles per operation, but can only execute hundreds of operations in parallel. This paper is the first work to perform a rigorous evaluation of bit-serial vs. bit-parallel in-situ processing-in-SRAM. Our results show that both approaches have similar area overheads. For 32-bit arithmetic operations, BS-VRAM improves throughput by 1.3-5.0× compared to BP-VRAM, while BP-VRAM improves latency by 3.0-23.0× compared to BS-VRAM.
Khalid Al-Hawaj, Olalekan Afuye, Shady O. Agwa, Alyssa B. Apsel, Christopher Batten
ISCAS2