EDBT 2026 Demo / reviewers in the wild / expert
Johnathan Alsop
dblp:162/9937
· DBLP profile ↗
10ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0001-5272-2396ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Memory systems · 49% GPUs and heterogeneous computing · 21% High-performance computing · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Concurrent programming · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
1.0 | 4 | 2022 | A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022 Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018 Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015 |
Memory systems › cache coherence › cache coherence protocol
heterogeneous coherence |
0.9 | 2 | 2022 | A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022 Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018 |
Memory systems › memory consistency
memory consistency model |
0.8 | 3 | 2017 | Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017 Lazy release consistency for GPUs · MICRO 2016 Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015 |
High-performance computing › supercomputing
exascale computing |
0.7 | 1 | 2023 | A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023 |
Hardware accelerators and domain-specific architectures › accelerator architecture
accelerator memory management |
0.6 | 1 | 2022 | A Case for Fine-grain Coherence Specialization in Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2022 |
GPUs and heterogeneous computing
GPU cache |
0.4 | 1 | 2020 | Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020 |
GPUs and heterogeneous computing › GPU scheduling
thread block scheduling |
0.4 | 1 | 2020 | Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020 |
Integrated circuit design
heterogeneous integration |
0.3 | 1 | 2018 | Spandex: A Flexible Interface for Efficient Heterogeneous Coherence · ISCA 2018 |
Concurrent programming
memory models |
0.2 | 1 | 2016 | Lazy release consistency for GPUs · MICRO 2016 |
GPUs and heterogeneous computing › GPU memory
GPU memory model |
0.2 | 1 | 2016 | Lazy release consistency for GPUs · MICRO 2016 |
Memory systems › memory consistency › memory consistency model › release consistency
lazy release consistency |
0.2 | 1 | 2016 | Lazy release consistency for GPUs · MICRO 2016 |
GPUs and heterogeneous computing › GPU computing
GPU synchronization |
0.2 | 1 | 2015 | Efficient GPU synchronization without scopes: saying no to complex consistency models · MICRO 2015 |
Memory systems
hybrid memory |
0.2 | 1 | 2015 | Stash: have your scratchpad and cache it too · ISCA 2015 |
High-performance computing
supercomputing |
0.2 | 1 | 2023 | A Research Retrospective on AMD's Exascale Computing Journey · ISCA 2023 |
Memory systems › cache coherence › cache coherence protocol
GPU coherence protocol |
0.1 | 1 | 2020 | Inter-kernel Reuse-aware Thread Block Scheduling · ACM Trans. Archit. Code Optim. 2020 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems |
0.1 | 1 | 2017 | Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2017 | Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous Systems · ISCA 2017 |
Methods — techniques the papers use, named apart from their topics
coherence protocol specialization · 0.6work stealing · 0.4denovo coherence protocol · 0.3performance evaluation · 0.3formal semantics · 0.3scratchpad memory · 0.2FIFO · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Research Retrospective on AMD's Exascale Computing JourneyabstractThe pace of advancement of the top-end supercomputers historically followed an exponential curve similar to (and driven in part by) Moore's Law. Shortly after hitting the petaflop mark, the community started looking ahead to the next milestone: Exascale. However, many obstacles were already looming on the horizon, such as the slowing of Moore's Law, and others like the end of Dennard Scaling had already arrived. Anticipating significant challenges for the overall high-performance computing (HPC) community to achieve the next 1000x improvement, the U.S. Department of Energy (DOE) launched the Exascale Computing Program to enable and accelerate fundamental research across the many technologies needed to achieve exascale computing. Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Vignesh Adhinarayanan, Shaizeen Aga, Derrick Aguren, Varun Agrawal, Ashwin M. Aji, Johnathan Alsop, Paul T. Bauman, Bradford M. Beckmann, Majed Valad Beigi, Sergey Blagodurov, Travis Boraten, Michael Boyer, William C. Brantley, Noel Chalmers, Shaoming Chen, Michael L. Chu, David Cownie, Nicholas Curtis, Joris Del Pino, Nam Duong, Alexandru Dutu, Yasuko Eckert, Christopher Erb, Chip Freitag, Joseph L. Greathouse, Sudhanva Gurumurthi, Anthony Gutierrez, Khaled Hamidouche, Sachin Hossamani, Wei Huang 0004, Mahzabeen Islam, Nuwan Jayasena, John Kalamatianos, Onur Kayiran, Jagadish Kotra, Alan Lee, Daniel Lowell, Niti Madan, Abhinandan Majumdar, Nicholas Malaya, Srilatha Manne, Susumu Mashimo, Damon McDougall, Elliot Mednick, Michael Mishkin, Mark Nutter, Indrani Paul, Matthew Poremba, Brandon Potter, Kishore Punniyamurthy, Sooraj Puthoor, Steven E. Raasch, Karthik Rao, Gregory Rodgers, Marko Scrbak, Mohammad Seyedzadeh, John Slice, Vilas Sridharan, René van Oostrum, Eric Van Tassell, Abhinav Vishnu, Samuel Wasmundt, Mark Wilkening, Noah Wolfe, Mark Wyse, Adithya Yalavarti, Dmitri Yudanov |
ISCA | 9 |
| 2022 | A Case for Fine-grain Coherence Specialization in Heterogeneous SystemsabstractHardware specialization is becoming a key enabler of energy-efficient performance. Future systems will be increasingly heterogeneous, integrating multiple specialized and programmable accelerators, each with different memory demands. Traditionally, communication between accelerators has been inefficient, typically orchestrated through explicit DMA transfers between different address spaces. More recently, industry has proposed unified coherent memory which enables implicit data movement and more data reuse, but often these interfaces limit the coherence flexibility available to heterogeneous systems. This paper demonstrates the benefits of fine-grained coherence specialization for heterogeneous systems. We propose an architecture that enables low-complexity independent specialization of each individual coherence request in heterogeneous workloads by building upon a simple and flexible baseline coherence interface, Spandex. We then describe how to optimize individual memory requests to improve cache reuse and performance-critical memory latency in emerging heterogeneous workloads. Collectively, our techniques enable significant gains, reducing execution time by up to 61% or network traffic by up to 99% while adding minimal complexity to the Spandex protocol. Johnathan Alsop, Weon Taek Na, Matthew D. Sinclair, Samuel Grayson, Sarita V. Adve |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Specializing Coherence, Consistency, and Push/Pull for GPU Graph AnalyticsabstractThis work explores the interaction of three communication-centric design dimensions for graph workloads on emerging integrated CPU-GPU systems: update propagation with and without fine-grained synchronization (push vs. pull), emerging coherence protocols (GPU vs. DeNovo coherence), and software-centric consistency models (DRF0, DRF1, and DRFrlx). We show that these dimensions are inter-dependent and the best design depends on the graph algorithm and input. We develop a model to predict this best design, motivating flexible and hardware-software co-designed GPU memory systems. Giordano Salvador, Wesley H. Darvin, Muhammad Huzaifa, Johnathan Alsop, Matthew D. Sinclair, Sarita V. Adve |
ISPASS | 4 |
| 2020 | Inter-kernel Reuse-aware Thread Block SchedulingabstractAs GPUs have become more programmable, their performance and energy benefits have made them increasingly popular. However, while GPU compute units continue to improve in performance, on-chip memories lag behind and data accesses are becoming increasingly expensive in performance and energy. Emerging GPU coherence protocols can mitigate this bottleneck by exploiting data reuse in GPU caches across kernel boundaries. Unfortunately, current GPU thread block schedulers are typically not designed to expose such reuse. This article proposes new hardware thread block schedulers that optimize inter-kernel reuse while using work stealing to preserve load balance. Our schedulers are simple, decentralized, and have extremely low overhead. Compared to a baseline round-robin scheduler, the best performing scheduler reduces average execution time and energy by 19% and 11%, respectively, in regular applications, and 10% and 8%, respectively, in irregular applications. Muhammad Huzaifa, Johnathan Alsop, Abdulrahman Mahmoud, Giordano Salvador, Matthew D. Sinclair, Sarita V. Adve |
ACM Trans. Archit. Code Optim. | 2 |
| 2018 | Spandex: A Flexible Interface for Efficient Heterogeneous CoherenceabstractRecent heterogeneous architectures have trended toward tighter integration and shared memory largely due to the efficient communication and programmability enabled by this shift. However, such integration is complex, because accelerators have widely disparate methods for accessing and keeping data coherent. Some processors use caches %that are backed by hardware coherence protocols like MESI, while others prefer lightweight software coherence protocols or use specialized memories like scratchpads with differing state and communication granularities. Modern solutions tend to build interfaces that extend existing MESI-style CPU coherence protocols, often by adding hierarchical indirection through intermediate shared caches. Although functionally correct, these strategies lack flexibility and generally suffer from performance limitations that make them sub-optimal for some emerging accelerators and workloads. Instead, we need a flexible interface that can efficiently integrate existing and future devices - without requiring intrusive changes to their memory structure. We introduce Spandex, an improved coherence interface based on the simple and scalable DeNovo coherence protocol. Spandex (which takes its name from the flexible material commonly used in one-size-fits-all textiles) directly interfaces devices with diverse coherence properties and memory demands, enabling each device to communicate in a manner appropriate for its specific access properties. We demonstrate the importance of this flexibility by comparing this strategy against a more conventional MESI-based hierarchical solution for a diverse range of heterogeneous applications. On average for the applications studied, Spandex reduces execution time by 16% (max 29%) and network traffic by 27% (max 58%) relative to the MESI-based hierarchical solution. Johnathan Alsop, Matthew D. Sinclair, Sarita V. Adve |
ISCA | 1 |
| 2017 | Chasing Away RAts: Semantics and Evaluation for Relaxed Atomics on Heterogeneous SystemsabstractAn unambiguous and easy-to-understand memory consistency model is crucial for ensuring correct synchronization and guiding future design of heterogeneous systems. In a widely adopted approach, the memory model guarantees sequential consistency (SC) as long as programmers obey certain rules. The popular data-race-free-0 (DRF0) model exemplifies this SC-centric approach by requiring programmers to avoid data races. Recent industry models, however, have extended such SC-centric models to incorporate relaxed atomics. These extensions can improve performance, but are difficult to specify formally and use correctly. This work addresses the impact of relaxed atomics on consistency models for heterogeneous systems in two ways. First, we introduce a new model, Data-Race-Free-Relaxed (DRFrlx), that extends DRF0 to provide SC-centric semantics for the common use cases of relaxed atomics. Second, we evaluate the performance of relaxed atomics in CPU-GPU systems for these use cases. We find mixed results -- for most cases, relaxed atomics provide only a small benefit in execution time, but for some cases, they help significantly (e.g., up to 51% for DRFrlx over DRF0). Matthew D. Sinclair, Johnathan Alsop, Sarita V. Adve |
ISCA | 2 |
| 2016 | GSI: A GPU Stall Inspector to characterize the sources of memory stalls for tightly coupled GPUsabstractIn recent years the power wall has prevented the continued scaling of single core performance. This has lead to the rise of dark silicon and motivated a move toward parallelism and specialization. As a result, energy-efficient high-throughput GPU cores are increasingly favored for accelerating data-parallel applications. However, the best way to efficiently communicate and synchronize across heterogeneous cores remains an important open research question. Many methods have been proposed to improve the efficiency of heterogeneous memory systems, but current methods for evaluating the performance effects of these innovations are limited in their ability to attribute differences in execution time to sources of latency in the memory system. Performance characterization of tightly coupled CPU-GPU systems is complicated by the high levels of parallelism present in GPU codes. Existing simulation tools provide only coarse-grained metrics which can obscure the underlying memory system interactions that cause performance differences. In this work we introduce GPU Stall Inspector (GSI), a method for identifying and visualizing the causes of GPU stalls with a focus on a tightly coupled CPU-GPU memory subsystem. We demonstrate the utility of our approach by evaluating the sources of stalls in several recent architectural innovations for tightly coupled, heterogeneous CPU-GPU systems. Johnathan Alsop, Matthew D. Sinclair, Rakesh Komuravelli, Sarita V. Adve |
ISPASS | 1 |
| 2016 | Lazy release consistency for GPUsabstractThe heterogeneous-race-free (HRF) memory model has been embraced by the Heterogeneous System Architecture (HSA) Foundation and OpenCLTMbecause it clearly and precisely defines the behavior of current GPUs. However, compared to the simpler SC for DRF memory model, HRF has two shortcomings. The first is that HRF requires programmers to label atomic memory operations with the correct scope of synchronization. This explicit labeling can save significant coherence overhead when synchronization is local, but it is tedious and error-prone. The second shortcoming is that HRF restricts important dynamic data sharing patterns like work stealing. Prior work on remote-scope promotion (RSP) attempted to resolve the second shortcoming. However, RSP further complicates the memory model and no scalable implementation of RSP has been proposed. For example, we found that the previously proposed RSP implementation actually results in slowdowns of up to 30% on large GPUs, compared to a naïve baseline system that forgoes work stealing and scopes. Meanwhile, DeNovo has been shown to offer efficient synchronization with an SC for DRF memory model, performing on average 21% better than our baseline system, but it introduces additional overheads to maintain ownership of all modified data. To resolve these deficiencies, we propose to adapt lazy release consistency - previously only proposed for homogeneous CPU systems - to a heterogeneous system. Our approach, called hLRC, uses a DeNovo-like mechanism to track ownership of synchronization variables, lazily performing coherence actions only when a synchronization variable changes locations. hLRC allows GPU programmers to use the simpler SC for DRF memory model without tracking ownership for all modified data. Our evaluation shows that lazy release consistency provides robust performance improvement across a set of work-stealing graph analysis applications - 29% on average versus the baseline system. Johnathan Alsop, Marc S. Orr, Bradford M. Beckmann, David A. Wood 0001 |
MICRO | 1 |
| 2015 | Stash: have your scratchpad and cache it tooabstractHeterogeneous systems employ specialization for energy efficiency. Since data movement is expected to be a dominant consumer of energy, these systems employ specialized memories (e.g., scratchpads and FIFOs) for better efficiency for targeted data. These memory structures, however, tend to exist in local address spaces, incurring significant performance and energy penalties due to inefficient data movement between the global and private spaces. We propose an efficient heterogeneous memory system where specialized memory components are tightly coupled in a unified and coherent address space. This paper applies these ideas to a system with CPUs and GPUs with scratchpads and caches. Rakesh Komuravelli, Matthew D. Sinclair, Johnathan Alsop, Muhammad Huzaifa, Maria Kotsifakou, Prakalp Srivastava, Sarita V. Adve, Vikram S. Adve |
ISCA | 3 |
| 2015 | Efficient GPU synchronization without scopes: saying no to complex consistency modelsabstractAs GPUs have become increasingly general purpose, applications with more general sharing patterns and fine- grained synchronization have started to emerge. Unfortunately, conventional GPU coherence protocols are fairly simplistic, with heavyweight requirements for synchronization accesses. Prior work has tried to resolve these inefficiencies by adding scoped synchronization to conventional GPU coherence protocols, but the resulting memory consistency model, heterogeneous-race-free (HRF), is more complex than the common data-race-free (DRF) model. This work applies the DeNovo coherence protocol to GPUs and compares it with conventional GPU coherence under the DRF and HRF consistency models. The results show that the complexity of the HRF model is neither necessary nor sufficient to obtain high performance. DeNovo with DRF provides a sweet spot in performance, energy, overhead, and memory consistency model complexity. Matthew D. Sinclair, Johnathan Alsop, Sarita V. Adve |
MICRO | 2 |