Johnathan Alsop

dblp:162/9937 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0001-5272-2396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Memory systems · 49% GPUs and heterogeneous computing · 21% High-performance computing · 13%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
1.042022
A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022
Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018
Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015
Memory systems › cache coherence › cache coherence protocol
heterogeneous coherence
0.922022
A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022
Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018
Memory systems › memory consistency
memory consistency model
0.832017
Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017
Lazy release consistency for GPUs · MICRO 2016
Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015
High-performance computing › supercomputing
exascale computing
0.712023
A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023
Hardware accelerators and domain-specific architectures › accelerator architecture
accelerator memory management
0.612022
A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022
GPUs and heterogeneous computing
GPU cache
0.412020
Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling
0.412020
Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020
Integrated circuit design
heterogeneous integration
0.312018
Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018
Concurrent programming
memory models
0.212016
Lazy release consistency for GPUs · MICRO 2016
GPUs and heterogeneous computing › GPU memory
GPU memory model
0.212016
Lazy release consistency for GPUs · MICRO 2016
Memory systems › memory consistency › memory consistency model › release consistency
lazy release consistency
0.212016
Lazy release consistency for GPUs · MICRO 2016
GPUs and heterogeneous computing › GPU computing
GPU synchronization
0.212015
Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015
Memory systems
hybrid memory
0.212015
Stash: have your scratchpad and cache it too · ISCA 2015
High-performance computing
supercomputing
0.212023
A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023
Memory systems › cache coherence › cache coherence protocol
GPU coherence protocol
0.112020
Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.112017
Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017
Parallel and multicore computing
synchronization
0.112017
Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017

Methods — techniques the papers use, named apart from their topics

coherence protocol specialization · 0.6work stealing · 0.4denovo coherence protocol · 0.3performance evaluation · 0.3formal semantics · 0.3scratchpad memory · 0.2FIFO · 0.2
YearPublicationVenuePosition
2023 A Research Retrospective on AMD's Exascale Computing Journey
abstract
The pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing.
Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov
ISCA9
2022 A Case for Fine-grain Coherence Specialization in Heterogeneous Systems
abstract
Hardware specialization is becoming a key enabler of energy-efficient performance. Future systems will be increasingly heterogeneous, integrating multiple specialized and programmable accelerators, each with different memory demands. Traditionally, communication between accelerators has been inefficient, typically orchestrated through explicit DMA transfers between different address spaces. More recently, industry has proposed unified coherent memory which enables implicit data movement and more data reuse, but often these interfaces limit the coherence flexibility available to heterogeneous systems. This paper demonstrates the benefits of fine-grained coherence specialization for heterogeneous systems. We propose an architecture that enables low-complexity independent specialization of each individual coherence request in heterogeneous workloads by building upon a simple and flexible baseline coherence interface, Spandex. We then describe how to optimize individual memory requests to improve cache reuse and performance-critical memory latency in emerging heterogeneous workloads. Collectively, our techniques enable significant gains, reducing execution time by up to 61% or network traffic by up to 99% while adding minimal complexity to the Spandex protocol.
Johnathan Alsop, Weon Taek Na, Matthew D. Sinclair, Samuel Grayson, Sarita V. Adve
ACM Trans. Archit. Code Optim.1
2020 Specializing Coherence, Consistency, and Push/Pull for GPU Graph Analytics
abstract
This work explores the interaction of three communication-centric design dimensions for graph workloads on emerging integrated CPU-GPU systems: update propagation with and without fine-grained synchronization (push vs. pull), emerging coherence protocols (GPU vs. DeNovo coherence), and software-centric consistency models (DRF0, DRF1, and DRFrlx). We show that these dimensions are inter-dependent and the best design depends on the graph algorithm and input. We develop a model to predict this best design, motivating flexible and hardware-software co-designed GPU memory systems.
Giordano Salvador, Wesley H. Darvin, Muhammad Huzaifa, Johnathan Alsop, Matthew D. Sinclair, Sarita V. Adve
ISPASS4
2020 Inter-kernel Reuse-aware Thread Block Scheduling
abstract
As GPUs have become more programmable, their performance and energy benefits have made them increasingly popular. However, while GPU compute units continue to improve in performance, on-chip memories lag behind and data accesses are becoming increasingly expensive in performance and energy. Emerging GPU coherence protocols can mitigate this bottleneck by exploiting data reuse in GPU caches across kernel boundaries. Unfortunately, current GPU thread block schedulers are typically not designed to expose such reuse. This article proposes new hardware thread block schedulers that optimize inter-kernel reuse while using work stealing to preserve load balance. Our schedulers are simple, decentralized, and have extremely low overhead. Compared to a baseline round-robin scheduler, the best performing scheduler reduces average execution time and energy by 19% and 11%, respectively, in regular applications, and 10% and 8%, respectively, in irregular applications.
Muhammad Huzaifa, Johnathan Alsop, Abdulrahman Mahmoud, Giordano Salvador, Matthew D. Sinclair, Sarita V. Adve
ACM Trans. Archit. Code Optim.2
2018 Spandex: A Flexible Interface for Efficient Heterogeneous Coherence
abstract
Recent heterogeneous architectures have trended toward tighter integration and shared memory largely due to the efficient communication and programmability enabled by this shift. However, such integration is complex, because accelerators have widely disparate methods for accessing and keeping data coherent. Some processors use caches %that are backed by hardware coherence protocols like MESI, while others prefer lightweight software coherence protocols or use specialized memories like scratchpads with differing state and communication granularities. Modern solutions tend to build interfaces that extend existing MESI-style CPU coherence protocols, often by adding hierarchical indirection through intermediate shared caches. Although functionally correct, these strategies lack flexibility and generally suffer from performance limitations that make them sub-optimal for some emerging accelerators and workloads. Instead, we need a flexible interface that can efficiently integrate existing and future devices - without requiring intrusive changes to their memory structure. We introduce Spandex, an improved coherence interface based on the simple and scalable DeNovo coherence protocol. Spandex (which takes its name from the flexible material commonly used in one-size-fits-all textiles) directly interfaces devices with diverse coherence properties and memory demands, enabling each device to communicate in a manner appropriate for its specific access properties. We demonstrate the importance of this flexibility by comparing this strategy against a more conventional MESI-based hierarchical solution for a diverse range of heterogeneous applications. On average for the applications studied, Spandex reduces execution time by 16% (max 29%) and network traffic by 27% (max 58%) relative to the MESI-based hierarchical solution.
Johnathan Alsop, Matthew D. Sinclair, Sarita V. Adve
ISCA1
2017 Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems
abstract
An unambiguous and easy-to-understand memory consistency model is crucial for ensuring correct synchronization and guiding future design of heterogeneous systems. In a widely adopted approach, the memory model guarantees sequential consistency (SC) as long as programmers obey certain rules. The popular data-race-free-0 (DRF0) model exemplifies this SC-centric approach by requiring programmers to avoid data races. Recent industry models, however, have extended such SC-centric models to incorporate relaxed atomics. These extensions can improve performance, but are difficult to specify formally and use correctly. This work addresses the impact of relaxed atomics on consistency models for heterogeneous systems in two ways. First, we introduce a new model, Data-Race-Free-Relaxed (DRFrlx), that extends DRF0 to provide SC-centric semantics for the common use cases of relaxed atomics. Second, we evaluate the performance of relaxed atomics in CPU-GPU systems for these use cases. We find mixed results -- for most cases, relaxed atomics provide only a small benefit in execution time, but for some cases, they help significantly (e.g., up to 51% for DRFrlx over DRF0).
Matthew D. Sinclair, Johnathan Alsop, Sarita V. Adve
ISCA2
2016 GSI: A GPU Stall Inspector to characterize the sources of memory stalls for tightly coupled GPUs
abstract
In recent years the power wall has prevented the continued scaling of single core performance. This has lead to the rise of dark silicon and motivated a move toward parallelism and specialization. As a result, energy-efficient high-throughput GPU cores are increasingly favored for accelerating data-parallel applications. However, the best way to efficiently communicate and synchronize across heterogeneous cores remains an important open research question. Many methods have been proposed to improve the efficiency of heterogeneous memory systems, but current methods for evaluating the performance effects of these innovations are limited in their ability to attribute differences in execution time to sources of latency in the memory system. Performance characterization of tightly coupled CPU-GPU systems is complicated by the high levels of parallelism present in GPU codes. Existing simulation tools provide only coarse-grained metrics which can obscure the underlying memory system interactions that cause performance differences. In this work we introduce GPU Stall Inspector (GSI), a method for identifying and visualizing the causes of GPU stalls with a focus on a tightly coupled CPU-GPU memory subsystem. We demonstrate the utility of our approach by evaluating the sources of stalls in several recent architectural innovations for tightly coupled, heterogeneous CPU-GPU systems.
Johnathan Alsop, Matthew D. Sinclair, Rakesh Komuravelli, Sarita V. Adve
ISPASS1
2016 Lazy release consistency for GPUs
abstract
The heterogeneous-race-free (HRF) memory model has been embraced by the Heterogeneous System Architecture (HSA) Foundation and OpenCLTMbecause it clearly and precisely defines the behavior of current GPUs. However, compared to the simpler SC for DRF memory model, HRF has two shortcomings. The first is that HRF requires programmers to label atomic memory operations with the correct scope of synchronization. This explicit labeling can save significant coherence overhead when synchronization is local, but it is tedious and error-prone. The second shortcoming is that HRF restricts important dynamic data sharing patterns like work stealing. Prior work on remote-scope promotion (RSP) attempted to resolve the second shortcoming. However, RSP further complicates the memory model and no scalable implementation of RSP has been proposed. For example, we found that the previously proposed RSP implementation actually results in slowdowns of up to 30% on large GPUs, compared to a naïve baseline system that forgoes work stealing and scopes. Meanwhile, DeNovo has been shown to offer efficient synchronization with an SC for DRF memory model, performing on average 21% better than our baseline system, but it introduces additional overheads to maintain ownership of all modified data. To resolve these deficiencies, we propose to adapt lazy release consistency - previously only proposed for homogeneous CPU systems - to a heterogeneous system. Our approach, called hLRC, uses a DeNovo-like mechanism to track ownership of synchronization variables, lazily performing coherence actions only when a synchronization variable changes locations. hLRC allows GPU programmers to use the simpler SC for DRF memory model without tracking ownership for all modified data. Our evaluation shows that lazy release consistency provides robust performance improvement across a set of work-stealing graph analysis applications - 29% on average versus the baseline system.
Johnathan Alsop, Marc S. Orr, Bradford M. Beckmann, David A. Wood 0001
MICRO1
2015 Stash: have your scratchpad and cache it too
abstract
Heterogeneous systems employ specialization for energy efficiency. Since data movement is expected to be a dominant consumer of energy, these systems employ specialized memories (e.g., scratchpads and FIFOs) for better efficiency for targeted data. These memory structures, however, tend to exist in local address spaces, incurring significant performance and energy penalties due to inefficient data movement between the global and private spaces. We propose an efficient heterogeneous memory system where specialized memory components are tightly coupled in a unified and coherent address space. This paper applies these ideas to a system with CPUs and GPUs with scratchpads and caches.
Rakesh Komuravelli, Matthew D. Sinclair, Johnathan Alsop, Muhammad Huzaifa, Maria Kotsifakou, Prakalp Srivastava, Sarita V. Adve, Vikram S. Adve
ISCA3
2015 Efficient GPU synchronization without scopes: saying no to complex consistency models
abstract
As GPUs have become increasingly general purpose, applications with more general sharing patterns and fine- grained synchronization have started to emerge. Unfortunately, conventional GPU coherence protocols are fairly simplistic, with heavyweight requirements for synchronization accesses. Prior work has tried to resolve these inefficiencies by adding scoped synchronization to conventional GPU coherence protocols, but the resulting memory consistency model, heterogeneous-race-free (HRF), is more complex than the common data-race-free (DRF) model. This work applies the DeNovo coherence protocol to GPUs and compares it with conventional GPU coherence under the DRF and HRF consistency models. The results show that the complexity of the HRF model is neither necessary nor sufficient to obtain high performance. DeNovo with DRF provides a sweet spot in performance, energy, overhead, and memory consistency model complexity.
Matthew D. Sinclair, Johnathan Alsop, Sarita V. Adve
MICRO2