EDBT 2026 Demo / reviewers in the wild / expert
Courtney Golden
dblp:332/2263
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0009-8874-7873ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quartz: A Reconfigurable, Distributed-Memory Accelerator for Sparse ApplicationsabstractIterative sparse matrix computations lie at the heart of many scientific computing and graph analytics algorithms.On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks.To overcome such bottlenecks, distributed-SRAM architectures are structured as an array of tiles, each with a processing element (PE) and a small local memory, to achieve very high aggregate memory bandwidth.However, current distributed-SRAM architectures suffer from either poor programmability due to over-specialized PEs or poor compute performance due to inefficient general-purpose PEs.We propose Quartz, a new architecture that uses short dataflow tasks and reconfigurable PEs in a distributed-SRAM system to deliver both high performance and high programmability.Unlike traditional sparse CGRAs or on-die reconfigurable engines, Quartz allows reconfigurable compute to be highly utilized and scaled by (1) providing high memory bandwidth to each processing element and (2) introducing a task-level dataflow execution model that fits this new setting.Our execution model dynamically reconfigures each tile's PE in response to inter-tile messages to execute tasks on local data.This execution model enables fine-grained data partitioning across tiles.To make execution efficient, we explore novel data partitioning techniques that use graph and hypergraph partitioning to minimize network traffic and balance load in the face of both static-static and static-dynamic operand sparsity.To ensure programmability, we show how a wide range of Einsum-expressible computations and flexible data distributions can be systematically captured in small tasks for execution on Quartz.Quartz's architecture, data partitioning techniques, and programming model together achieve gmean 21.4× speedup over a prior state-of-the-art system for six different iterative sparse applications from scientific computing and graph analytics. Courtney Golden, Axel Feldmann, Joel S. Emer, Daniel Sánchez 0003 |
MICRO | 1 |
| 2025 | Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM DeviceabstractCompute-in-SRAM architectures offer a promising approach to achieving higher performance and energy efficiency across a range of data-intensive applications.However, prior evaluations have largely relied on simulators or small prototypes, limiting the understanding of their real-world potential.In this work, we present a comprehensive performance and energy characterization of a commercial compute-in-SRAM device, the GSI APU, under realistic workloads.We compare the GSI APU against established architectures, including CPUs and GPUs, to quantify its energy efficiency and performance potential.We introduce an analytical framework for general-purpose compute-in-SRAM devices that reveals fundamental optimization principles by modeling performance trade-offs, thereby guiding program optimizations.Exploiting the fine-grained parallelism of tightly integrated memory-compute architectures requires careful data management.We address this by proposing three optimizations: communicationaware reduction mapping, coalesced DMA, and broadcast-friendly data layouts.When applied to retrieval-augmented generation (RAG) over large corpora (10GB-200GB), these optimizations enable our compute-in-SRAM system to accelerate retrieval by 4.8×-6.6×over an optimized CPU baseline, improving end-to-end RAG latency by 1.1×-1.8×.The shared off-chip memory bandwidth is modeled using a simulated HBM, while all other components are measured on the real compute-in-SRAM device.Critically, this system matches the performance of an NVIDIA A6000 GPU for RAG while being significantly more energy-efficient (54.4×-117.9×reduction).These findings validate the viability of compute-in-SRAM for complex, real-world applications and provide guidance for advancing the technology. Niansong Zhang, Courtney Golden, Dan Ilan, Hongzheng Chen, Christopher Batten, Zhiru Zhang |
MICRO | 3 |
| 2024 | Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip MemoryabstractSolving sparse systems of linear equations is a fundamental primitive in many numeric algorithms. Iterative solvers provide an efficient way of solving large, highly sparse systems. However, iterative solvers are inefficient on existing architectures because they perform computations with (1) poor short-term reuse, which causes frequent off-chip memory traffic; and (2) challenaing data dependences, which limit parallelism. We present Azul, a hardware accelerator that achieves high arithmetic intensity by keeping data in distributed on-chip SRAM. Azul is organized as a grid of tiles, each with a small memory and a simple processing element (PE). This enables keeping solver data on-chip across iterations, achieving high reuse. We present a novel scheduling algorithm that maps data and computation across PEs to avoid communication bottlenecks while achieving high parallelism, and a specialized PE that achieves high utilization of arithmetic units. When tested on a representative set of matrices for sparse iterative solvers, Azul is gmean 217 × faster than state-of-the art GPU implementations, 159× faster than a previously proposed accelerator for sparse iterative solvers, and 90 × faster than a previously proposed distributed-SRAM accelerator. Axel Feldmann, Courtney Golden, Joel S. Emer, Daniel Sánchez 0003 |
MICRO | 2 |
| 2023 | EVE: Ephemeral Vector EnginesabstractThere has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead. Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady O. Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, Christopher Batten |
HPCA | 7 |
| 2022 | big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on ChipabstractSingle-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads. Tuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou, Eric Hall, Courtney Golden, Christopher Batten |
MICRO | 6 |