EDBT 2026 Demo / reviewers in the wild / expert
Simon C. Steely Jr.
dblp:58/2198
· DBLP profile ↗
11ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11Software engineering, systems software and programming languages · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Memory systems · 88% Processor architecture and microarchitecture · 4% Electronic design automation · 4% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.6 | 4 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 CRUISE: cache replacement and utility-aware scheduling · ASPLOS 2012 PACMan: prefetch-aware cache management for high performance caching · MICRO 2011 |
Memory systems › cache management
cache replacement |
0.3 | 3 | 2012 | The gradient-based cache partitioning algorithm · ACM Trans. Archit. Code Optim. 2012 High performance cache replacement using re-reference interval prediction (RRIP) · ISCA 2010 Adaptive insertion policies for high performance caching · ISCA 2007 |
Memory systems
cache coherence |
0.3 | 2 | 2013 | Using in-flight chains to build a scalable cache coherence protocol · ACM Trans. Archit. Code Optim. 2013 Achieving Non-Inclusive Cache Performance with Inclusive Caches: Temporal Locality Aware (TLA) Cache Management Policies · MICRO 2010 |
Memory systems
cache management |
0.3 | 2 | 2012 | The gradient-based cache partitioning algorithm · ACM Trans. Archit. Code Optim. 2012 Achieving Non-Inclusive Cache Performance with Inclusive Caches: Temporal Locality Aware (TLA) Cache Management Policies · MICRO 2010 |
Memory systems › memory hierarchy › cache hierarchy management
last-level cache management |
0.2 | 2 | 2011 | PACMan: prefetch-aware cache management for high performance caching · MICRO 2011 SHiP: signature-based hit predictor for high performance caching · MICRO 2011 |
Memory systems › cache management
re-reference interval prediction |
0.2 | 2 | 2011 | SHiP: signature-based hit predictor for high performance caching · MICRO 2011 High performance cache replacement using re-reference interval prediction (RRIP) · ISCA 2010 |
Memory systems › memory hierarchy
cache hierarchy |
0.2 | 1 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 |
Memory systems › memory hierarchy › cache hierarchy management
exclusive caching |
0.2 | 1 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.2 | 1 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 |
Memory systems › cache coherence
directory-based coherence |
0.2 | 1 | 2013 | Using in-flight chains to build a scalable cache coherence protocol · ACM Trans. Archit. Code Optim. 2013 |
Memory systems › cache management
cache partitioning |
0.1 | 1 | 2012 | The gradient-based cache partitioning algorithm · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › cache management › cache replacement
last-level cache replacement |
0.1 | 1 | 2012 | CRUISE: cache replacement and utility-aware scheduling · ASPLOS 2012 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2012 | CRUISE: cache replacement and utility-aware scheduling · ASPLOS 2012 |
Memory systems › cache management
cache insertion policy |
0.1 | 1 | 2011 | SHiP: signature-based hit predictor for high performance caching · MICRO 2011 |
Memory systems › cache design
inclusive cache |
0.1 | 1 | 2010 | Achieving Non-Inclusive Cache Performance with Inclusive Caches: Temporal Locality Aware (TLA) Cache Management Policies · MICRO 2010 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 CRUISE: cache replacement and utility-aware scheduling · ASPLOS 2012 |
Memory systems
cache design |
0.1 | 1 | 2007 | Adaptive insertion policies for high performance caching · ISCA 2007 |
Cloud and datacenter computing
server workloads |
0.1 | 1 | 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive caches · HPCA 2015 |
Parallel and multicore computing › multiprocessor system
scalable multiprocessor |
0.0 | 1 | 2013 | Using in-flight chains to build a scalable cache coherence protocol · ACM Trans. Archit. Code Optim. 2013 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2012 | CRUISE: cache replacement and utility-aware scheduling · ASPLOS 2012 |
Cloud and datacenter computing
quality of service |
0.0 | 1 | 2012 | The gradient-based cache partitioning algorithm · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › cache › multiprocessor cache
shared cache |
0.0 | 1 | 2012 | The gradient-based cache partitioning algorithm · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › cache › prefetching
hardware prefetching |
0.0 | 1 | 2011 | PACMan: prefetch-aware cache management for high performance caching · MICRO 2011 |
Performance modeling and evaluation
design trade-off analysis |
0.0 | 1 | 1979 | Performance simulation as a tool in central processing unit design · SIGMETRICS 1979 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1979 | Performance simulation as a tool in central processing unit design · SIGMETRICS 1979 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.3in-flight chains · 0.2hierarchical tag directory · 0.2utility-aware scheduling · 0.1hardware-software co-design · 0.1gradient-based optimization · 0.1signature-based prediction · 0.1workload characterization · 0.1temporal locality hints · 0.1cache replacement policy · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | High performing cache hierarchies for server workloads: Relaxing inclusion to capture the latency benefits of exclusive cachesabstractIncreasing transistor density enables adding more on-die cache real-estate However, devoting more space to the shared last-level-cache (LLC) causes the memory latency bottleneck to move from memory access latency to shared cache access latency. As such, applications whose working set is larger than the smaller caches spend a large fraction of their execution time on shared cache access latency. To address this problem, this paper investigates increasing the size of smaller private caches in the hierarchy as opposed to increasing the shared LLC. Doing so improves average cache access latency for workloads whose working set fits into the larger private cache while retaining the benefits of a shared LLC. The consequence of increasing the size of private caches is to relax inclusion and build exclusive hierarchies. Thus, for the same total caching capacity, an exclusive cache hierarchy provides better cache access latency. We observe that server workloads benefit tremendously from an exclusive hierarchy with large private caches. This is primarily because large private caches accommodate the large code working-sets of server workloads. For a 16-core CMP, an exclusive cache hierarchy improves server workload performance by 5-12% as compared to an equal capacity inclusive cache hierarchy. The paper also presents directions for further research to maximize performance of exclusive cache hierarchies. Aamer Jaleel, Joseph Nuzman, Adrian Moga, Simon C. Steely Jr., Joel S. Emer |
HPCA | 4 |
| 2013 | Using in-flight chains to build a scalable cache coherence protocolabstractAs microprocessor designs integrate more cores, scalability of cache coherence protocols becomes a challenging problem. Most directory-based protocols avoid races by using blocking tag directories that can impact the performance of parallel applications. In this article, we first quantitatively demonstrate that state-of-the-art blocking protocols significantly constrain throughput at large core counts for several parallel applications. Nonblocking protocols address this throughput concern at the expense of scalability in the interconnection network or in the required resource overheads. To address this concern, we enhance nonblocking directory protocols by migrating the point of service of responses. Our approach uses in-flight chains of cores making parallel memory requests to incorporate scalability while maintaining high-throughput. The proposed cache coherence protocol called chained cache coherence , can outperform blocking protocols by up to 20% on scientific and 12% on commercial applications. It also has low resource overheads and simple address ordering requirements making it both a high-performance and scalable protocol. Furthermore, in-flight chains provide a scalable solution to building hierarchical and nonblocking tag directories as well as optimize communication latencies. Samantika Sury, Simon C. Steely Jr., William Hasenplaugh, Aamer Jaleel, Carl J. Beckmann, Tryggve Fossum, Joel S. Emer |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | CRUISE: cache replacement and utility-aware schedulingabstractWhen several applications are co-scheduled to run on a system with multiple shared LLCs, there is opportunity to improve system performance. This opportunity can be exploited by the hardware, software, or a combination of both hardware and software. The software, i.e., an operating system or hypervisor, can improve system performance by co-scheduling jobs on LLCs to minimize shared cache contention. The hardware can improve system throughput through better replacement policies by allocating more cache resources to applications that benefit from the cache and less to those applications that do not. This study presents a detailed analysis on the interactions between intelligent scheduling and smart cache replacement policies. We find that smart cache replacement reduces the burden on software to provide intelligent scheduling decisions. However, under smart cache replacement, there is still room to improve performance from better application co-scheduling. We find that co-scheduling decisions are a function of the underlying LLC replacement policy. We propose Cache Replacement and Utility-aware Scheduling (CRUISE)-a hardware/software co-designed approach for shared cache management. For 4-core and 8-core CMPs, we find that CRUISE approaches the performance of an ideal job co-scheduling policy under different LLC replacement policies. Aamer Jaleel, Hashem Hashemi Najaf-abadi, Samantika Sury, Simon C. Steely Jr., Joel S. Emer |
ASPLOS | 4 |
| 2012 | The gradient-based cache partitioning algorithmabstractThis paper addresses the problem of partitioning a cache between multiple concurrent threads and in the presence of hardware prefetching. Cache replacement designed to preserve temporal locality (e.g., LRU) will allocate cache resources proportional to the miss-rate of each competing thread irrespective of whether the cache space will be utilized [Qureshi and Patt 2006]. This is clearly suboptimal as applications vary dramatically in their use of recently accessed data. We address this problem by partitioning a shared cache such that a global goodness metric is optimized. This paper introduces the Gradient-based Cache Partitioning Algorithm (GPA), whose variants optimize either hitrate, total instructions per cycle (IPC) or a weighted IPC metric designed to enforce Quality of Service (QoS) [Iyer 2004]. In the context of QoS, GPA enables us to obtain the maximum throughput of low-priority threads, while ensuring high performance on high-priority threads. The GPA mechanism is robust, low-cost, integrates easily with existing cache designs and improves the throughput of an in-order 8-core system sharing an 8MB L3 cache by ∼14%. William Hasenplaugh, Pritpal S. Ahuja, Aamer Jaleel, Simon C. Steely Jr., Joel S. Emer |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | SHiP: signature-based hit predictor for high performance cachingabstractThe shared last-level caches in CMPs play an important role in improving application performance and reducing off-chip memory bandwidth requirements. In order to use LLCs more efficiently, recent research has shown that changing the re-reference prediction on cache insertions and cache hits can significantly improve cache performance. A fundamental challenge, however, is how to best predict the re-reference pattern of an incoming cache line. Carole-Jean Wu, Aamer Jaleel, William Hasenplaugh, Margaret Martonosi, Simon C. Steely Jr., Joel S. Emer |
MICRO | 5 |
| 2011 | PACMan: prefetch-aware cache management for high performance cachingabstractHardware prefetching and last-level cache (LLC) management are two independent mechanisms to mitigate the growing latency to memory. However, the interaction between LLC management and hardware prefetching has received very little attention. This paper characterizes the performance of state-of-the-art LLC management policies in the presence and absence of hardware prefetching. Although prefetching improves performance by fetching useful data in advance, it can interact with LLC management policies to introduce application performance variability. This variability stems from the fact that current replacement policies treat prefetch and demand requests identically. Carole-Jean Wu, Aamer Jaleel, Margaret Martonosi, Simon C. Steely Jr., Joel S. Emer |
MICRO | 4 |
| 2010 | High performance cache replacement using re-reference interval prediction (RRIP)abstractPractical cache replacement policies attempt to emulate optimal replacement by predicting the re-reference interval of a cache block. The commonly used LRU replacement policy always predicts a near-immediate re-reference interval on cache hits and misses. Applications that exhibit a distant re-reference interval perform badly under LRU. Such applications usually have a working-set larger than the cache or have frequent bursts of references to non-temporal data (called scans). To improve the performance of such workloads, this paper proposes cache replacement using Re-reference Interval Prediction (RRIP). We propose Static RRIP (SRRIP) that is scan-resistant and Dynamic RRIP (DRRIP) that is both scan-resistant and thrash-resistant. Both RRIP policies require only 2-bits per cache block and easily integrate into existing LRU approximations found in modern processors. Our evaluations using PC games, multimedia, server and SPEC CPU2006 workloads on a single-core processor with a 2MB last-level cache (LLC) show that both SRRIP and DRRIP outperform LRU replacement on the throughput metric by an average of 4% and 10% respectively. Our evaluations with over 1000 multi-programmed workloads on a 4-core CMP with an 8MB shared LLC show that SRRIP and DRRIP outperform LRU replacement on the throughput metric by an average of 7% and 9% respectively. We also show that RRIP outperforms LFU, the state-of the art scan-resistant replacement algorithm to-date. For the cache configurations under study, RRIP requires 2X less hardware than LRU and 2.5X less hardware than LFU. Aamer Jaleel, Kevin B. Theobald, Simon C. Steely Jr., Joel S. Emer |
ISCA | 3 |
| 2010 | Achieving Non-Inclusive Cache Performance with Inclusive Caches: Temporal Locality Aware (TLA) Cache Management PoliciesabstractInclusive caches are commonly used by processors to simplify cache coherence. However, the trade-off has been lower performance compared to non-inclusive and exclusive caches. Contrary to conventional wisdom, we show that the limited performance of inclusive caches is mostly due to inclusion victims—lines that are evicted from the core caches to satisfy the inclusion property—and not the reduced cache capacity of the hierarchy due to the duplication of data. These inclusion victims are incorrectly chosen for replacement because the last-level cache (LLC) is unaware of the temporal locality of lines in the core caches. We propose Temporal Locality Aware (TLA) cache management policies to allow an inclusive LLC to be aware of the temporal locality of lines in the core caches. We propose three TLA policies: Temporal Locality Hints (TLH), Early Core Invalidation (ECI), and Query Based Selection (QBS). All three policies improve inclusive cache performance without requiring any additional hardware structures. In fact, QBS performs similar to a non-inclusive cache hierarchy. Aamer Jaleel, Eric Borch, Malini Bhandaru, Simon C. Steely Jr., Joel S. Emer |
MICRO | 4 |
| 2008 | Adaptive insertion policies for managing shared cachesabstractChip Multiprocessors (CMPs) allow different applications to concurrently execute on a single chip. When applications with differing demands for memory compete for a shared cache, the conventional LRU replacement policy can significantly degrade cache performance when the aggregate working set size is greater than the shared cache. In such cases, shared cache performance can be significantly improved by preserving the entire working set of applications that can co-exist in the cache and preserving some portion of the working set of the remaining applications. Aamer Jaleel, William Hasenplaugh, Moinuddin K. Qureshi, Julien Sebot, Simon C. Steely Jr., Joel S. Emer |
PACT | 5 |
| 2007 | Adaptive insertion policies for high performance cachingabstractThe commonly used LRU replacement policy is susceptible to thrashing for memory-intensive workloads that have a working set greater than the available cache size. For such applications, the majority of lines traverse from the MRU position to the LRU position without receiving any cache hits, resulting in inefficient use of cache space. Cache performance can be improved if some fraction of the working set is retained in the cache so that at least that fraction of the working set can contribute to cache hits. We show that simple changes to the insertion policy can significantly reduce cache misses for memory-intensive workloads. We propose the LRU Insertion Policy (LIP) which places the incoming line in the LRU position instead of the MRU position. LIP protects the cache from thrashing and results in close to optimal hitrate for applications that have a cyclic reference pattern. We also propose the Bimodal Insertion Policy (BIP) as an enhancement of LIP that adapts to changes in the working set while maintaining the thrashing protection of LIP. We finally propose a Dynamic Insertion Policy (DIP) to choose between BIP and the traditional LRU policy depending on which policy incurs fewer misses. The proposed insertion policies do not require any change to the existing cache structure, are trivial to implement, and have a storage requirement of less than two bytes. We show that DIP reduces the average MPKI of the baseline 1MB 16-way L2 cache by 21%, bridging two-thirds of the gap between LRU and OPT. Moinuddin K. Qureshi, Aamer Jaleel, Yale N. Patt, Simon C. Steely Jr., Joel S. Emer |
ISCA | 4 |
| 1979 | Performance simulation as a tool in central processing unit designabstractPerformance analysis has always been considered important in computer design work. The area of central processing unit (CPU) design is no exception, where the successful development of performance evaluation tools provides valuable information in the analysis of design tradeoffs. Increasing integration of hardware is producing more complicated processor modules which add to the number of alternatives and decisions to be made in the design process. It is important that these modules work together as a balanced unit with no hidden bottlenecks. This paper describes a project to develop performance simulation as an analysis tool in CPU design. The methodology is first detailed as a three part process in which a performance simulation program is realized that executes an instruction trace using command file directions. Discussion follows on the software implemented, applications of this tool in CPU design, and future goals. Cheryl A. Wiecek, Simon C. Steely Jr. |
SIGMETRICS | 2 |