EDBT 2026 Demo / reviewers in the wild / expert
Erik Hagersten
dblp:h/ErikHagersten
· DBLP profile ↗
52ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 1 first-authorSoftware engineering, systems software and programming languages · 16 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
20 papers |
Memory systems · 42% Performance modeling and evaluation · 21% Processor architecture and microarchitecture · 18% |
Topics — the 30 heaviest of 58, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.7 | 6 | 2017 | Building Heterogeneous Unified Virtual Memories (UVMs) without the Overhead · ACM Trans. Archit. Code Optim. 2016 The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015 A Split Cache Hierarchy for Enabling Data-Oriented Optimizations · HPCA 2017 |
Memory systems › memory hierarchy
cache hierarchy |
0.5 | 2 | 2017 | A Split Cache Hierarchy for Enabling Data-Oriented Optimizations · HPCA 2017 Navigating the cache hierarchy with a single lookup · ISCA 2014 |
Processor architecture and microarchitecture
out-of-order execution |
0.4 | 2 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Performance modeling and evaluation › simulation › architectural simulation
sampled simulation |
0.4 | 1 | 2019 | Directed Statistical Warming through Time Traveling · MICRO 2019 |
Performance modeling and evaluation
simulation |
0.4 | 1 | 2019 | Directed Statistical Warming through Time Traveling · MICRO 2019 |
Memory systems › cache coherence
cache coherence protocol |
0.3 | 3 | 2015 | The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015 A case for low-complexity MP architectures · SC 2007 Removing the overhead from software-based shared memory · SC 2001 |
Performance modeling and evaluation
analytical modeling |
0.2 | 1 | 2016 | Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016 |
Energy-efficient computing
power modeling |
0.2 | 1 | 2016 | Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016 |
Performance modeling and evaluation
processor performance modeling |
0.2 | 1 | 2016 | Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016 |
GPUs and heterogeneous computing › GPU memory management
unified virtual memory |
0.2 | 1 | 2016 | Building Heterogeneous Unified Virtual Memories (UVMs) without the Overhead · ACM Trans. Archit. Code Optim. 2016 |
Memory systems
cache design |
0.2 | 2 | 2014 | TLC: a tag-less cache for reducing dynamic first level cache energy · MICRO 2013 Navigating the cache hierarchy with a single lookup · ISCA 2014 |
Memory systems › cache coherence
data classification |
0.2 | 1 | 2015 | The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.2 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Processor architecture and microarchitecture
instruction scheduling |
0.2 | 1 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Embedded and real-time systems › real-time scheduling
mixed-criticality scheduling |
0.2 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Cloud and datacenter computing
resource allocation |
0.2 | 1 | 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processors · MICRO 2015 |
Memory systems › cache coherence › invalidation
self-invalidation |
0.2 | 1 | 2015 | The Effects of Granularity and Adaptivity on Private/Shared Classification for Coherence · ACM Trans. Archit. Code Optim. 2015 |
Processor architecture and microarchitecture › instruction scheduling
speculative scheduling |
0.2 | 1 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 |
Memory systems › memory interference
cache contention |
0.2 | 1 | 2013 | Modeling performance variation due to cache sharing · HPCA 2013 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.2 | 1 | 2013 | TLC: a tag-less cache for reducing dynamic first level cache energy · MICRO 2013 |
Performance modeling and evaluation
performance variability |
0.2 | 1 | 2013 | Modeling performance variation due to cache sharing · HPCA 2013 |
Memory systems › cache design
tag-less cache |
0.2 | 1 | 2013 | TLC: a tag-less cache for reducing dynamic first level cache energy · MICRO 2013 |
Performance modeling and evaluation
workload characterization |
0.2 | 2 | 2019 | Directed Statistical Warming through Time Traveling · MICRO 2019 Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2007 | A case for low-complexity MP architectures · SC 2007 Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Memory systems
cache management |
0.1 | 1 | 2010 | Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory Accesses · SC 2010 |
Memory systems › cache management › cache resource management
cache pollution control |
0.1 | 1 | 2010 | Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory Accesses · SC 2010 |
Memory systems
cache |
0.1 | 2 | 2015 | Cost-effective speculative scheduling in high performance processors · ISCA 2015 Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Memory systems › cache coherence
directory |
0.1 | 1 | 2017 | A Split Cache Hierarchy for Enabling Data-Oriented Optimizations · HPCA 2017 |
Parallel and multicore computing
synchronization |
0.1 | 2 | 2003 | Hierarchical Backoff Locks for Nonuniform Communication Architectures · HPCA 2003 Efficient synchronization for nonuniform communication architectures · SC 2002 |
Electronic design automation
design space exploration |
0.1 | 1 | 2016 | Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.5time traveling · 0.4directed statistical warming · 0.4cycle-level simulation · 0.2coherence protocol · 0.2analytical modeling · 0.2criticality prediction · 0.2set-associative cache design · 0.2native execution profiling · 0.2full-system simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Directed Statistical Warming through Time TravelingabstractImproving the speed of computer architecture evaluation is of paramount importance to shorten the time-to-market when developing new platforms. Sampling is a widely used methodology to speed up workload analysis and performance evaluation by extrapolating from a set of representative detailed regions. Installing an accurate cache state for each detailed region is critical to achieving high accuracy. Prior work requires either huge amounts of storage (checkpoint-based warming), an excessive number of memory accesses to warm up the cache (functional warming), or the collection of a large number of reuse distances (randomized statistical warming) to accurately predict cache warm-up effects. Nikos Nikoleris, Lieven Eeckhout, Erik Hagersten, Trevor E. Carlson |
MICRO | 3 |
| 2017 | POSTER: Putting the G back into GPU/CPU Systems ResearchabstractModern SoCs contain several CPU cores and many GPU cores to execute both general purpose and highly-parallel graphics workloads. In many SoCs, more area is dedicated to graphics than to general purpose compute. Despite this, the micro-architecture research community primarily focuses on GPGPU and CPU-only research, and not on graphics (the primary workload for many SoCs). The main reason for this is the lack of efficient tools and simulators for modern graphics applications.This work focuses on the GPU's memory traffic generated by graphics. We describe a new graphics tracing framework and use it to both study graphics applications' memory behavior as well as how CPUs and GPUs affect system performance. Our results show that graphics applications exhibit a wide range of memory behavior between applications and across time, and slows down co-running SPEC applications by 59% on average. Andreas Sembrant, Trevor E. Carlson, Erik Hagersten, David Black-Schaffer |
PACT | 3 |
| 2017 | A Split Cache Hierarchy for Enabling Data-Oriented OptimizationsabstractToday's caches tightly couple data with metadata (Address Tags) at the cache line granularity. The co-location of data and its identifying metadata means that they require multiple approaches to locate data (associative way searches and level-by-level searches), evict data (coherent writebacks buffers and associative level-by-level searches) and keep data coherent (directory indirections and associative level-by-level searches). This results in complex implementations with many corner cases, increased latency and energy, and limited flexibility for data optimizations. We propose splitting the metadata and data into two separate structures: a metadata hierarchy and a data hierarchy. The metadata hierarchy tracks the location of the data in the data hierarchy. This allows us to easily apply many different optimizations to the data hierarchy, including smart data placement, dynamic coherence, and direct accesses. The new split cache hierarchy, Direct-to-Master (D2M), provides a unified mechanism for cache searching, eviction, and coherence, that eliminates level-by-level data movement and searches, associative cache address tags comparisons and about 90% of the indirections through a central directory. Optimizations such as moving LLC slices to the near-side of the network and private/shared data classification can easily be built on top off D2M to further improve its efficiency. This approach delivers a 54% improvement in cache hierarchy EDP vs. a mobile processor and 40% vs. a server processor, reduces network traffic by an average of 70%, reduces the L1 miss latency by 30% and is especially effective for workloads with high cache pressure. Andreas Sembrant, Erik Hagersten, David Black-Schaffer |
HPCA | 2 |
| 2016 | Data placement across the cache hierarchy: Minimizing data movement with reuse-aware placementabstractModern processors employ multiple levels of caching to address bandwidth, latency and performance requirements. The behavior of these hierarchies is determined by their approach to data placement and data eviction. Recent research has developed many intelligent data eviction policies, but cache hierarchies remain primarily either exclusive or inclusive with regards to data placement. This means that today's cache hierarchies typically install accessed data into all cache levels at one point or another, regardless of whether the data is reused in each level. Such data movement wastes energy by installing data into cache levels where the data is not reused. Andreas Sembrant, Erik Hagersten, David Black-Schaffer |
ICCD | 2 |
| 2016 | Message from the general chairabstractI am delighted to welcome you to the 2016 International Symposium on Performance Analysis of Systems and Software (ISPASS). On its 16th birthday, ISPASS has grown old enough to travel “abroad” for the first time and this 17th ISPASS edition is being held in Uppsala, Sweden. Erik Hagersten |
ISPASS | 1 |
| 2016 | CoolSim: Eliminating traditional cache warming with fast, virtualized profilingabstractSampling (e.g., SMARTS and SimPoint) improves simulation performance by an order of magnitude or more through the reduction of large workloads into a small but representative sample. Virtualized fast-forwarding (e.g., FSA) speeds up simulation further by advancing execution at near-native speed between simulation points, making cache warming the critical limiting factor for simulation performance. CoolSim is an efficient simulation framework that eliminates cache warming. It collects sparse memory reuse information (MRI) while advancing between simulation points using virtualized fast-forwarding. During detailed simulation, a statistical cache model uses the previously acquired MRI to estimate the performance of the caches. CoolSim builds upon KVM and gem5 and runs 19x faster than the state-of-the-art sampled simulation. It estimates the CPI of the SPEC CPU2006 benchmarks with 3.62% error on average, across a wide range of cache sizes. Nikos Nikoleris, Andreas Sandberg, Erik Hagersten, Trevor E. Carlson |
ISPASS | 3 |
| 2016 | Building Heterogeneous Unified Virtual Memories (UVMs) without the OverheadabstractThis work proposes a novel scheme to facilitate heterogeneous systems with unified virtual memory. Research proposals implement coherence protocols for sequential consistency (SC) between central processing unit (CPU) cores and between devices. Such mechanisms introduce severe bottlenecks in the system; therefore, we adopt the heterogeneous-race-free (HRF) memory model. The use of HRF simplifies the coherency protocol and the graphics processing unit (GPU) memory management unit (MMU). Our protocol optimizes CPU and GPU demands separately, with the GPU part being simpler while the CPU is more elaborate and latency aware. We achieve an average 45% speedup and 45% energy-delay product reduction (20% energy) over the corresponding SC implementation. Konstantinos Koukos, Alberto Ros 0001, Erik Hagersten, Stefanos Kaxiras |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent CharacteristicsabstractOptimizing processors for (a) specific application(s) can substantially improve energy-efficiency. With the end of Dennard scaling, and the corresponding reduction in energy-efficiency gains from technology scaling, such approaches may become increasingly important. However, designing application-specific processors requires fast design space exploration tools to optimize for the targeted application(s). Analytical models can be a good fit for such design space exploration as they provide fast performance and power estimates and insight into the interaction between an application's characteristics and the micro-architecture of a processor. Unfortunately, prior analytical models for superscalar out-of-order processors require micro-architecture dependent inputs, such as cache miss rates, branch miss rates and memory-level parallelism. This requires profiling the applications for each cache and branch predictor configuration of interest, which is far more time-consuming than evaluating the analytical performance models. In this work we present amicro-architecture independentprofiler and associated analytical models that allow us to produce performanceandpower estimates across a large superscalar out-of-order processor design space almost instantaneously. We show that using a micro-architecture independent profile leads to a speedup of 300$\times$compared to detailed simulation for our evaluated design space. Over a large design space, the model has a 9.3 percent average error for performance and a 4.3 percent average error for power, compared to detailed cycle-level simulation. The model is able to accurately determine the optimal processor configuration for different applications under power or performance constraints, and provides insight into performance through cycle stacks. Sam Van den Steen, Stijn Eyerman, Sander De Pestel, Moncef Mechri, Trevor E. Carlson, David Black-Schaffer, Erik Hagersten, Lieven Eeckhout |
IEEE Trans. Computers | 7 |
| 2015 | An Efficient, Self-Contained, On-chip Directory: DIR1-SISDabstractDirectory-based cache coherence is the de-facto standard for scalable shared-memory multi/many-cores and significant effort is invested in reducing its overhead. However, directory area and complexity optimizations are often antithetical to each other. Novel directory-less coherence schemes have been introduced to remove the complexity and cost associated with directories in their entirety. However, such schemes introduce new challenges by transferring some of the directory complexity and functionality to the OS and using the page table and the TLBs to store data classification information. In this work we bridge the gap between directory-based and directory-less coherence schemes and propose a hybrid scheme called DIR1-SISD which employs self-invalidation and self-downgrade as directory policies for the shared entries. DIR1-SISD allows simultaneous optimizations in area and complexity without relying on the OS. DIR1-SISD keeps track of a single -- private -- owner, or allows multiple-readers-multiple-writers to exist simultaneously by transferring the responsibility for their coherence to the corresponding cores. A DIR1-SISD self-contained directory cache has a unique ability to minimize eviction-induced complexities by allowing directory entries to be evicted without maintaining inclusion with the cached data (thus avoiding the complexities of broadcasts) and without the need to have a backing store. Using simulation we show that a small, self-contained, DIR1-SISD cache outperforms a traditional DIR16-NB MESI protocol with a directory cache embedded in the LLC (8% in execution time and 15% in traffic) and, further, outperforms a SISD protocol that relies on the OS to provide a persistent page-based directory (4% in execution time and 20% in traffic). Mahdad Davari, Alberto Ros 0001, Erik Hagersten, Stefanos Kaxiras |
PACT | 3 |
| 2015 | AREP: Adaptive Resource Efficient Prefetching for Maximizing Multicore PerformanceabstractModern processors widely use hardware prefetching to hide memory latency. While aggressive hardware prefetchers can improve performance significantly for some applications, they can limit the overall performance in highly-utilized multicore processors by saturating the offchip bandwidth and wasting last-level cache capacity. Co-executing applications can slowdown due to contention over these shared resources. This work introduces Adaptive Resource Efficient Prefetching (AREP) -- a runtime framework that dynamically combines software prefetching and hardware prefetching to maximize throughput in highly utilized multicore processors. AREP achieves better performance by prefetching data in a resource efficient way -- conserving offchip-bandwidth and last-level cache capacity with accurate prefetching and by applying cache-bypassing when possible. AREP dynamically explores a mix of hardware/software prefetching policies, then selects and applies the best performing policy. AREP is phase-aware and re-explores (at runtime) for the best prefetching policy at phase boundaries. A multitude of experiments with workload mixes and parallel applications on a modern high performance multicore show that AREP can increase throughput by up to 49% (8.1% on average). This is complemented by improved fairness, resulting in average quality of service above 94%. Muneeb Khan, Michael Laurenzano, Jason Mars, Erik Hagersten, David Black-Schaffer |
PACT | 4 |
| 2015 | Cost-effective speculative scheduling in high performance processorsabstractTo maximize performance, out-of-order execution processors sometimes issue instructions without having the guarantee that operands will be available in time; e.g. loads are typically assumed to hit in the L1 cache and dependent instructions are issued accordingly. This form of speculation -- that we refer to as speculative scheduling -- has been used for two decades in real processors, but has received little attention from the research community. Arthur Perais, André Seznec, Pierre Michaud, Andreas Sembrant, Erik Hagersten |
ISCA | 5 |
| 2015 | Micro-architecture independent analytical processor performance and power modelingabstractOptimizing processors for specific application(s) can substantially improve energy-efficiency. With the end of Dennard scaling, and the corresponding reduction in energyefficiency gains from technology scaling, such approaches may become increasingly important. However, designing applicationspecific processors require fast design space exploration tools to optimize for the targeted application(s). Analytical models can be a good fit for such design space exploration as they provide fast performance estimations and insight into the interaction between an application's characteristics and the micro-architecture of a processor. Unfortunately, current analytical models require some microarchitecture dependent inputs, such as cache miss rates, branch miss rates and memory-level parallelism. This requires profiling the applications for each cache and branch predictor configuration, which is far more time-consuming than evaluating the actual performance models. In this work we present a micro-architecture independent profiler and associated analytical models that allow us to produce performance and power estimates across a large design space almost instantaneously. We show that using a micro-architecture independent profile leads to a speedup of 25× for our evaluated design space, compared to an analytical model that uses micro-architecture dependent profiles. Over a large design space, the model has a 13% error for performance and a 7% error for power, compared to cycle-level simulation. The model is able to accurately determine the optimal processor configuration for different applications under power or performance constraints, and it can provide insight into performance through cycle stacks. Sam Van den Steen, Sander De Pestel, Moncef Mechri, Stijn Eyerman, Trevor E. Carlson, David Black-Schaffer, Erik Hagersten, Lieven Eeckhout |
ISPASS | 7 |
| 2015 | Long term parking (LTP): criticality-aware resource allocation in OOO processorsabstractModern processors employ large structures (IQ, LSQ, register file, etc.) to expose instruction-level parallelism (ILP) and memory-level parallelism (MLP). These resources are typically allocated to instructions in program order. This wastes resources by allocating resources to instructions that are not yet ready to be executed and by eagerly allocating resources to instructions that are not part of the application's critical path. Andreas Sembrant, Trevor E. Carlson, Erik Hagersten, David Black-Schaffer, Arthur Perais, André Seznec, Pierre Michaud |
MICRO | 3 |
| 2015 | The Effects of Granularity and Adaptivity on Private/Shared Classification for CoherenceabstractClassification of data into private and shared has proven to be a catalyst for techniques to reduce coherence cost, since private data can be taken out of coherence and resources can be concentrated on providing coherence for shared data. In this article, we examine how granularity—page-level versus cache-line level—and adaptivity—going from shared to private—affect the outcome of classification and its final impact on coherence. We create a classification technique, called Generational Classification , and a coherence protocol called Generational Coherence, which treats data as private or shared based on cache-line generations. We compare two coherence protocols based on self-invalidation/self-downgrade with respect to data classification. Our findings are enlightening: (i) Some programs benefit from finer granularity, some benefit further from adaptivity, but some do not benefit from either. (ii) Reducing the amount of shared data has no perceptible impact on coherence misses caused by self-invalidation of shared data, hence no impact on performance. (iii) In contrast, classifying more data as private has implications for protocols that employ write-through as a means of self-downgrade, resulting in network traffic reduction—up to 30%—by reducing write-through traffic. Mahdad Davari, Alberto Ros 0001, Erik Hagersten, Stefanos Kaxiras |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | A Case for Resource Efficient Prefetching in MulticoresabstractModern processors typically employ sophisticated prefetching techniques for hiding memory latency. Hardware prefetching has proven very effective and can speed up some SPEC CPU 2006 benchmarks by more than 40% when running in isolation. However, this speedup often comes at the cost of prefetching a significant volume of useless data (sometimes more than twice the data required) which wastes shared last level cache space and off-chip bandwidth. This paper explores how an accurate resource-efficient prefetching scheme can benefit performance by conserving shared resources in multicores. We present a framework that uses low-overhead runtime sampling and fast cache modeling to accurately identify memory instructions that frequently miss in the cache. We then use this information to automatically insert software prefetches in the application. Our prefetching scheme has good accuracy and employs cache bypassing whenever possible. These properties help reduce off-chip bandwidth consumption and last-level cache pollution. While single-thread performance remains comparable to hardware prefetching, the full advantage of the scheme is realized when several cores are used and demand for shared resources grows. We evaluate our method on two modern commodity multicores. Across 180 mixed workloads that fully utilize a multicore, the proposed software prefetching mechanism achieves up to 24% better throughput than hardware prefetching, and performs 10% better on average. Muneeb Khan, Andreas Sandberg, Erik Hagersten |
ICPP | 3 |
| 2014 | Navigating the cache hierarchy with a single lookupabstractModern processors optimize for cache energy and performance by employing multiple levels of caching that address bandwidth, low-latency and high-capacity. A request typically traverses the cache hierarchy, level by level, until the data is found, thereby wasting time and energy in each level. In this paper, we present the Direct-to-Data (D2D) cache that locates data across the entire cache hierarchy with a single lookup. Andreas Sembrant, Erik Hagersten, David Black-Schaffer |
ISCA | 2 |
| 2014 | A software based profiling method for obtaining speedup stacks on commodity multi-coresabstractA key goodness metric of multi-threaded programs is how their execution times scale when increasing the number of threads. However, there are several bottlenecks that can limit the scalability of a multi-threaded program, e.g., contention for shared cache capacity and off-chip memory bandwidth; and synchronization overheads. In order to improve the scalability of a multi-threaded program, it is vital to be able to quantify how the program is impacted by these scalability bottlenecks. We present a software profiling method for obtaining speedup stacks. A speedup stack reports how much each scalability bottleneck limits the scalability of a multi-threaded program. It thereby quantifies how much its scalability can be improved by eliminating a given bottleneck. A software developer can use this information to determine what optimizations are most likely to improve scalability, while a computer architect can use it to analyze the resource demands of emerging workloads. The proposed method profiles the program on real commodity multi-cores (i.e., no simulations required) using existing performance counters. Consequently, the obtained speedup stacks accurately account for all idiosyncrasies of the machine on which the program is profiled. While the main contribution of this paper is the profiling method to obtain speedup stacks, we present several examples of how speedup stacks can be used to analyze the resource requirements of multi-threaded programs. Furthermore, we discuss how their scalability can be improved by both software developers and computer architects. David Eklov, Nikos Nikoleris, Erik Hagersten |
ISPASS | 3 |
| 2014 | A case for resource efficient prefetching in multicoresabstractHardware prefetching has proven very effective for hiding memory latency and can speed up some applications by more than 40%. However, this speedup comes at the cost of often prefetching a significant volume of useless data which wastes shared last level cache space and off-chip bandwidth. This directly impacts the performance of co-scheduled applications which compete for shared resources in multicores. This paper explores how a resource-efficient prefetching scheme can benefit performance by conserving shared resources in multicores. We present a framework that uses fast cache modeling to accurately identify memory instructions that benefit most from prefetching. The framework inserts software prefetches in the application only when they benefit performance, and employs cache bypassing whenever possible. These properties help reduce off-chip bandwidth consumption and last-level cache pollution. While single-thread performance remains comparable to hardware prefetching, the full advantage of the scheme is realized when several cores are used and demand for shared resources grows. Muneeb Khan, Andreas Sandberg, Erik Hagersten |
ISPASS | 3 |
| 2014 | Extending statistical cache models to support detailed pipeline simulatorsabstractSimulators are widely used in computer architecture research. While detailed cycle-accurate simulations provide useful insights, studies using modern workloads typically require days or weeks. Evaluating many design points, only exacerbates the simulation overhead. Recent works propose methods with good accuracy that reduce the simulated overhead either by sampling the execution (e.g., SMARTS and SimPoint) or by using fast analytical models of the simulated designs (e.g., Interval Simulation). While these techniques reduce significantly the simulation overhead, modeling processor components with large state, such as the last-level cache, requires costly simulation to warm them up. Statistical simulation methods, such as SMARTS, report that the warm-up overhead accounts for 99% of the simulation overhead, while only 1% of the time is spent simulating the target design. This paper proposes WarmSim, a method that eliminates the need to warm up the cache. WarmSim builds on top of a statistical cache modeling technique and extends it to model accurately not only the miss ratio but also the outcome of every cache request. WarmSim uses as input, an application's memory reuse information which is hardware independent. Therefore, different cache configurations can be simulated using the same input data. We demonstrate that this approach can be used to estimate the CPI of the SPEC CPU2006 benchmarks with an average error of 1.77%, reducing the overhead compared to a simulation with a 10M instruction warm-up by a factor of 50x. Nikos Nikoleris, David Eklov, Erik Hagersten |
ISPASS | 3 |
| 2013 | Bandwidth Bandit: Quantitative characterization of memory contentionabstractOn multicore processors, co-executing applications compete for shared resources, such as cache capacity and memory bandwidth. This leads to suboptimal resource allocation and can cause substantial performance loss, which makes it important to effectively manage these shared resources. This, however, requires insights into how the applications are impacted by such resource sharing. While there are several methods to analyze the performance impact of cache contention, less attention has been paid to general, quantitative methods for analyzing the impact of contention for memory bandwidth. To this end we introduce the Bandwidth Bandit, a general, quantitative, profiling method for analyzing the performance impact of contention for memory bandwidth on multicore machines. The profiling data captured by the Bandwidth Bandit is presented in a bandwidth graph. This graph accurately captures the measured application's performance as a function of its available memory bandwidth, and enables us to determine how much the application suffers when its available bandwidth is reduced. To demonstrate the value of this data, we present a case study in which we use the bandwidth graph to analyze the performance impact of memory contention when co-running multiple instances of single threaded application. David Eklov, Nikos Nikoleris, David Black-Schaffer, Erik Hagersten |
CGO | 4 |
| 2013 | Modeling performance variation due to cache sharingabstractShared cache contention can cause significant variability in the performance of co-running applications from run to run. This variability arises from different overlappings of the applications' phases, which can be the result of offsets in application start times or other delays in the system. Understanding this variability is important for generating an accurate view of the expected impact of cache contention. However, variability effects are typically ignored due to the high overhead of modeling or simulating the many executions needed to expose them. This paper introduces a method for efficiently investigating the performance variability due to cache contention. Our method relies on input data captured from native execution of applications running in isolation and a fast, phase-aware, cache sharing performance model. This allows us to assess the performance interactions and bandwidth demands of co-running applications by quickly evaluating hundreds of overlappings. We evaluate our method on a contemporary multicore machine and show that performance and bandwidth demands can vary significantly across runs of the same set of co-running applications. We show that our method can predict application slowdown with an average relative error of 0.41% (maximum 1.8%) as well as bandwidth consumption. Using our method, we can estimate an application pair's performance variation 213× faster, on average, than native execution. Andreas Sandberg, Andreas Sembrant, Erik Hagersten, David Black-Schaffer |
HPCA | 3 |
| 2013 | TLC: a tag-less cache for reducing dynamic first level cache energyabstractFirst level caches are performance-critical and are therefore optimized for speed. To do so, modern processors reduce the miss ratio by using set-associative caches and optimize latency by reading all ways in parallel with the TLB and tag lookup. However, this wastes energy since only data from one way is actually used. Andreas Sembrant, Erik Hagersten, David Black-Schaffer |
MICRO | 2 |
| 2012 | Bandwidth bandit: quantitative characterization of memory contentionabstractApplications that are co-scheduled on a multi-core compete for shared resources, such as cache capacity and memory bandwidth. The performance degradation resulting from this contention can be substantial, which makes it important to effectively manage these shared resources. This, however, requires quantitative insight into how applications are impacted by such contention. David Eklov, Nikos Nikoleris, David Black-Schaffer, Erik Hagersten |
PACT | 4 |
| 2012 | Efficient techniques for predicting cache sharing and throughputabstractThis work addresses the modeling of shared cache contention in multicore systems and its impact on throughput and bandwidth. We develop two simple and fast cache sharing models for accurately predicting shared cache allocations for random and LRU caches. Andreas Sandberg, David Black-Schaffer, Erik Hagersten |
PACT | 3 |
| 2012 | Phase guided profiling for fast cache modelingabstractStatistical cache models are powerful tools for understanding application behavior as a function of cache allocation. However, previous techniques have modeled only the average application behavior, which hides the effect of program variations over time. Without detailed time-based information, transient behavior, such as exceeding bandwidth or cache capacity, may be missed. Yet these events, while short, often play a disproportionate role and are critical to understanding program behavior. Andreas Sembrant, David Black-Schaffer, Erik Hagersten |
CGO | 3 |
| 2012 | Bandwidth bandit: Understanding memory contentionabstractApplications that are co-scheduled on a multicore compete for shared resources, such as cache capacity and memory bandwidth. The performance degradation resulting from this contention can be substantial, which makes it important to effectively manage these shared resources. This, however, requires insight into how applications are impacted by such contention. In this paper we present a quantitative method to measure applications' sensitivities to different degrees of contention for off-chip memory bandwidth on real hardware. This method is then used to demonstrate the varying contention sensitivity across a selection of benchmarks, and explains why some of them experience substantial slowdowns long before the overall memory bandwidth saturates. David Eklov, Nikos Nikoleris, David Black-Schaffer, Erik Hagersten |
ISPASS | 4 |
| 2012 | Low Overhead Instruction-Cache Modeling Using Instruction Reuse ProfilesabstractPerformance loss caused by L1 instruction cache misses varies between different architectures and cache sizes. For processors employing power-efficient in-order execution with small caches, performance can be significantly affected by instruction cache misses. The growing use of low-power multi-threaded CPUs (with shared L1 caches) in general purpose computing platforms requires new efficient techniques for analyzing application instruction cache usage. Such insight can be achieved using traditional simulation technologies modeling several cache sizes, but the overhead of simulators may be prohibitive for practical optimization usage. In this paper we present a statistical method to quickly model application instruction cache performance. Most importantly we propose a very low-overhead sampling mechanism to collect runtime data from the application's instruction stream. This data is fed to the statistical model which accurately estimates the instruction cache miss ratio for the sampled execution. Our sampling method is about 10x faster than previously suggested sampling approaches, with average runtime overhead as low as 25% over native execution. The architecturally-independent data collected is used to accurately model miss ratio for several cache sizes simultaneously, with average absolute error of 0.2%. Finally, we show how our tool can be used to identify program phases with large instruction cache footprint. Such phases can then be targeted to optimize for reduced code footprint. Muneeb Khan, Andreas Sembrant, Erik Hagersten |
SBAC-PAD | 3 |
| 2011 | Fast modeling of shared caches in multicore systemsabstractThis work presents StatCC, a simple and efficient model for estimating the shared cache miss ratios of co-scheduled applications on architectures with a hierarchy of private and shared caches. StatCC leverages the StatStack cache model to estimate the co-scheduled applications' cache miss ratios from their individual memory reuse distance distributions, and a simple performance model that estimates their CPIs based on the shared cache miss ratios. These methods are combined into a system of equations that explicitly models the CPIs in terms of the shared miss ratios and can be solved to determine both. The result is a fast algorithm with a 2% error across the SPEC CPU2006 benchmark suite compared to a simulated in-order processor and a hierarchy of private and shared caches. David Eklov, David Black-Schaffer, Erik Hagersten |
HiPEAC | 3 |
| 2011 | Cache Pirating: Measuring the Curse of the Shared CacheabstractWe present a low-overhead method for accurately measuring application performance (CPI) and off-chip bandwidth (GB/s) as a function of available shared cache capacity. The method is implemented on real hardware, with no modifications to the application or operating system. We accomplish this by co-running a Pirate application that "steals" cache space with the Target application. By adjusting how much space the Pirate steals during the Target's execution, and using hardware performance counters to record the Target's performance, we can accurately and efficiently capture performance data for the Target application as a function of its available shared cache. At the same time we use performance counters to monitor the Pirate to ensure that it is successfully stealing the desired amount of cache. To evaluate this approach, we show that 1) the cache available to the Target behaves as expected, 2) the Pirate steals the desired amount of cache, and ) the Pirate does not bias the Target's performance. As a result, we are able to accurately measure the Target's performance while stealing up to an average of 6.8MB of the 8MB of cache on our Nehalem based test system with an average measurement overhead of only 5.5%. David Eklov, Nikos Nikoleris, David Black-Schaffer, Erik Hagersten |
ICPP | 4 |
| 2010 | StatCC: a statistical cache contention modelabstractChip multiprocessor (CMP) architectures sharing on chip resources, such as last-level caches, have recently become a mainstream computing platform. The performance of such systems can vary greatly depending on how co-scheduled applications compete for these shared resources. This work presents StatCC, a simple and efficient model for estimating the contention for shared cache resources between co-scheduled applications on chip multiprocessor architectures. David Eklov, David Black-Schaffer, Erik Hagersten |
PACT | 3 |
| 2010 | StatStack: Efficient modeling of LRU cachesabstractEfficient execution on modern architectures requires good data locality, which can be measured by the powerful stack distance abstraction. Based on this abstraction, the miss rate for LRU caches of any size can be predicted. However, measuring stack distance requires the number of unique memory objects to be counted between successive accesses to the same data object, which requires complex and inefficient data collection. This paper presents a new efficient way of estimating the stack distances of an application. Instead of counting the number of unique memory objects touched between successive accesses to the same data, our scheme only requires the number of memory accesses to be counted, a task efficiently handled by existing builtin hardware counters. Furthermore, this information only needs to be captured for a small fraction of the memory accesses. A new efficient off-line algorithm is proposed to estimate the corresponding stack distance based on this sparse information. We evaluate the accuracy of the proposed estimation compared with full stack distance measurements for 28 of the applications in the SPEC CPU2006 benchmark suite. The estimation shows excellent accuracy based on information about only every 10,000th memory access. David Eklov, Erik Hagersten |
ISPASS | 2 |
| 2010 | Reducing Cache Pollution Through Detection and Elimination of Non-Temporal Memory AccessesabstractContention for shared cache resources has been recognized as a major bottleneck for multicores--especially for mixed workloads of independent applications. While most modern processors implement instructions to manage caches, these instructions are largely unused due to a lack of understanding of how to best leverage them. This paper introduces a classification of applications into four cache usage categories. We discuss how applications from different categories affect each other's performance indirectly through cache sharing and devise a scheme to optimize such sharing. We also propose a low-overhead method to automatically find the best per-instruction cache management policy. We demonstrate how the indirect cache-sharing effects of mixed workloads can be tamed by automatically altering some instructions to better manage cache resources. Practical experiments demonstrate that our software-only method can improve application performance up to 35% on x86 multicore hardware. Andreas Sandberg, David Eklov, Erik Hagersten |
SC | 3 |
| 2007 | Conserving Memory Bandwidth in Chip Multiprocessors with Runahead ExecutionabstractThe introduction of chip multiprocessors (CMPs) presents new challenges and trade-offs to computer architects. Architects must now strike a balance between the number of cores per chip versus the amount of on-chip cache and the cost-efficient amount of pin bandwidth. Technology projections indicate that the cost of pin bandwidth would increase significantly and may therefore inhibit the number of processor cores per CMP. Runahead execution is a very promising approach to tolerate long memory latencies. In this paper we study the memory access characteristics of runahead execution. We show that temporal and data dependency aspects of runahead execution makes it possible to conserve bandwidth through the use of smaller cache blocks in the cache. We demonstrate, using execution-driven full system simulation, that our method of fine-grained fetching can obtain significant performance speedups in bandwidth constrained systems but also yield stable performance in systems that are not bandwidth limited. Martin Karlsson, Erik Hagersten |
IPDPS | 2 |
| 2007 | A case for low-complexity MP architecturesabstractAdvances in semiconductor technology have driven sharedmemory servers toward processors with multiple cores per die and multiple threads per core. This paper presents simple hardware primitives enabling flexible and low-complexity multi-chip designs supporting an efficient inter-node coherence protocol implemented in software. We argue that our primitives and the example design presented in this paper have lower hardware overhead, have easier (and later) verification requirements, and provide the opportunity for flexible coherence protocols and simpler protocol bug corrections than traditional designs. Our evaluation is based on detailed full-system simulations of modern chip-multiprocessors and both commercial and HPC workloads. We compare a low-complexity system based on the proposed primitives with aggressive hardware multi-chip shared-memory systems and show that the performance is competitive across a large design space. 1. Håkan Zeffer, Erik Hagersten |
SC | 2 |
| 2006 | Multigrid and Gauss-Seidel smoothers revisited: parallelization on chip multiprocessorsabstractEfficient solution of partial differential equations require a match between the algorithm and the target architecture. Many recent chip multiprocessors, CMPs (a.k.a. multi-core), feature low intra-thread communication costs and smaller per-thread caches compared to previous shared memory multi-processor systems. From an algorithmic point of view this means that data locality issues become more important than communication overheads. A fact that may require a re-evaluation of many existing algorithms.We have investigated parallel implementations of multi-grid methods using a parallel temporally blocked, naturally ordered smoother. Compared to the standard multigrid solution based on a red-black ordering, we improve the data locality often as much as ten times, while our use of a fine-grained locking scheme keeps the parallel efficiency high.Our algorithm was initially inspired by CMPs and it was surprising to see that our OpenMP multigrid implementation ran up to 40 percent faster than the standard red-black algorithm on a contemporary 8-way SMP system. Thanks to the temporal blocking introduced, our smoother implementation often allowed us to apply the smoother two times at the same cost as a single application of a red-black smoother. By executing our smoother on a 32-thread UltraSPARC T1 (Niagara) SMT/CMP and a simulated 32-way CMP we demonstrate that such architectures can tolerate the increased communication costs implied by the tradeoffs made in our implementation. Dan Wallin, Henrik Löf, Erik Hagersten, Sverker Holmgren |
ICS | 3 |
| 2006 | TMA: a trap-based memory architectureabstractThe advances in semiconductor technology have set the shared-memory server trend towards processors with multiple cores per die and multiple threads per core. We believe that this technology shift forces a reevaluation of how to interconnect multiple such chips to form larger systems.This paper argues that by adding support for coherence traps in future chip multiprocessors, large-scale server systems can be formed at a much lower cost. This is due to shorter design time, verification and time to market when compared to its traditional all-hardware counter part. In the proposed trap-based memory architecture (TMA), software trap handlers are responsible for obtaining read/write permission, whereas the coherence trap hardware is responsible for the actual permission check.In this paper we evaluate a TMA implementation (called TMA Lite) with a minimal amount of hardware extensions, all contained within the processor. The proposed mechanisms for coherence trap processing should not affect the critical path and have a negligible cost in terms of area and power for most processor designs.Our evaluation is based on detailed full system simulation using out-of-order processors with one or two dual-threaded cores per die as processing nodes. The results show that a TMA based distributed shared memory system can perform on par with a highly optimized hardware based design. Håkan Zeffer, Zoran Radovic, Martin Karlsson, Erik Hagersten |
ICS | 4 |
| 2006 | Exploiting locality: a flexible DSM approachabstractNo single coherence strategy suits all applications well. Many promising adaptive protocols and coherence predictors, capable of dynamically modifying the coherence strategy, have been suggested over the years. While most dynamic detection schemes rely on plentiful of dedicated hardware, the customization technique suggested in this paper requires no extra hardware support for its per-application coherence strategy. Instead, each application is profiled using a low-overhead profiling tool. The appropriate coherence flag setting, suggested by the profiling, is specified when the application is launched. We have compared the performance of a hardware DSM (Sun WildFire) to a software DSM (distributed shared memory) built with identical interconnect hardware and coherence strategy. With no support for flexibility, the software DSM runs on average 45 percent slower than the hardware DSM on the 12 studied applications, while the flexibility can get the software DSM within 11 percent. Our all-software system outperforms the hardware DSM on four applications Håkan Zeffer, Zoran Radovic, Erik Hagersten |
IPDPS | 3 |
| 2006 | A statistical multiprocessor cache modelabstractThe introduction of general-purpose microprocessors running multiple threads will put a focus on methods and tools helping a programmer to write efficient parallel applications. Such a tool should be fast enough to meet a software developer's need for short turn-around time, but also be accurate and flexible enough to provide trend-correct and intuitive feedback. This paper presents a novel sample-based method for analyzing the data locality of a multithreaded application. Very sparse data is collected during a single execution of the studied application. The architectural-independent information collected during the execution is fed to a mathematical memory-system model for predicting the cache miss ratio. The sparse data can be used to characterize the application's data locality with respect to almost any possible memory system, such as complicated multiprocessor multilevel cache hierarchies. Any combination of cache size, cache-line size and degree of sharing can be modeled. Each modeled design point takes only a fraction of a second to evaluate, even though the application from which the sampled data was collected may have executed for hours. This makes the tool not just usable for software developers, but also for hardware developers who need to evaluate a huge memory-system design space. The accuracy of the method is evaluated using a large number of commercial and technical multi-threaded applications. The result produced by the algorithm is shown to be consistent with results from a traditional (and much slower) architecture simulation. Erik Berg, Håkan Zeffer, Erik Hagersten |
ISPASS | 3 |
| 2005 | Exploring Processor Design Options for Java-Based MiddlewareabstractJava-based middleware is a rapidly growing workload for high-end server processors, particularly chip multiprocessors (CMP). To help architects design future microprocessors to run this important new workload, we provide a detailed characterization of two popular Java server benchmarks, ECperf and SPECjbb2000. We first estimate the amount of instruction-level parallelism in these workloads by simulating a very wide issue processor with perfect caches and perfect branch predictors. We then identify performance bottlenecks for these workloads on a more realistic processor by selectively idealizing individual processor structures. Finally, we combine our findings on available ILP in Java middleware with results from previous papers that characterize the availibility of TLP to investigate the optimal balance between ILP and TLP in CMPs. We find that, like other commercial workloads, Java middleware has only a small amount of instruction-level parallelism, even when run on very aggressive processors. When run on processors resembling currently available processors, the performance of Java middleware is limited by frequent traps, address translation and stalls in the memory system. We find that SPECjbb2000 differs from ECperf in two meaningful ways: (1) the performance of ECperf is affected much more by cache and TLB misses during instruction fetch and (2) SPECjbb2000 has more memory-level parallelism. Martin Karlsson, Erik Hagersten, Kevin E. Moore, David A. Wood 0001 |
ICPP | 2 |
| 2005 | Fast data-locality profiling of native executionabstractPerformance tools based on hardware counters can efficiently profile the cache behavior of an application and help software developers improve its cache utilization. Simulator-based tools can potentially provide more insights and flexibility and model many different cache configurations, but have the drawback of large run-time overhead.We present StatCache, a performance tool based on a statistical cache model. It has a small run-time overhead while providing much of the flexibility of simulator-based tools. A monitor process running in the background collects sparse memory access statistics about the analyzed application running natively on a host computer. Generic locality information is derived and presented in a code-centric and/or data-centric view.We evaluate the accuracy and performance of the tool using ten SPEC CPU2000 benchmarks. We also exemplify how the flexibility of the tool can be used to better understand the characteristics of cache-related performance problems. Erik Berg, Erik Hagersten |
SIGMETRICS | 2 |
| 2004 | Exploiting Spatial Store Locality Through Permission Caching in Software DSMs
Håkan Zeffer, Zoran Radovic, Oskar Grenholm, Erik Hagersten |
Euro-Par | 4 |
| 2004 | Bundling: Reducing the Overhead of Multiprocessor PrefetchersabstractSummary form only given. Prefetching has proven to be a useful technique for reducing cache misses in multiprocessors at the cost of increased coherence traffic. This is especially trouble some for snoop-based systems, where the available coherence bandwidth often is the scalability bottleneck. The bundling technique reduces the overhead caused by prefetching in two ways: piggybacking prefetches with normal requests, and requiring only one device to perform the snoop lookup for each prefetch transaction. This can reduce both the address bandwidth and the number of snoop lookups compared with a nonprefetching system. We describe bundling implementations for two important transaction types: reads and upgrades. While bundling could reduce the overhead of most existing prefetch schemes, the evaluation of bundling performed has been limited to two of them: sequential prefetching and Dahlgren's adaptive sequential prefetching. Both schemes have their snoop bandwidth halved for all commercial and scientific benchmarks in the study. The combined effect of bundling applied to these prefetch schemes lowers the cache miss rate, the address bandwidth and the snoop bandwidth, compared with a system with no prefetching, for all applications. Bundling, will not reduce the data bandwidth introduced by a prefetch scheme. However, we argue that the data bandwidth is more easily scaled than the snoop bandwidth for snoop-based coherence systems. Dan Wallin, Erik Hagersten |
IPDPS | 2 |
| 2004 | StatCache: a probabilistic approach to efficient and accurate data locality analysisabstractThe widening memory gap reduces performance of applications with poor data locality. Therefore, there is a need for methods to analyze data locality and help application optimization. In this paper we present StatCache, a novel sampling-based method for performing data-locality analysis on realistic workloads. StatCache is based on a probabilistic model of the cache, rather than a functional cache simulator. It uses statistics from a single run to accurately estimate miss ratios of fully-associative caches of arbitrary sizes and generate working-set graphs. We evaluate StatCache using the SPEC CPU2000 benchmarks and show that StatCache gives accurate results with a sampling rate as low as 10/sup -4/. We also provide a proof-of-concept implementation, and discuss potentially very fast implementation alternatives. Erik Berg, Erik Hagersten |
ISPASS | 2 |
| 2003 | THROOM - Supporting POSIX Multithreaded Binaries on a Cluster
Henrik Löf, Zoran Radovic, Erik Hagersten |
Euro-Par | 3 |
| 2003 | Memory System Behavior of Java-Based MiddlewareabstractIn this paper, we present a detailed characterization of the memory system, behavior of ECperf and SPECjbb using both commercial server hardware and Simics full-system simulation. We find that the memory footprint and primary working sets of these workloads are small compared to other commercial workloads (e.g. on-line transaction processing), and that a large fraction of the working sets are shared between processors. We observed two key differences between ECperf and SPECjbb that highlight the importance of isolating the behavior of the middle tier. First, ECperf has a larger instruction footprint, resulting in much higher miss rates for intermediate-size instruction caches. Second, SPECjbb's data set size increases linearly as the benchmark scales up, while ECperf's remains roughly constant. This difference can lead to opposite conclusions on the design of multiprocessor memory systems, such as the utility of moderate sized (i.e. 1 MB) shared caches in a chip multiprocessor. Martin Karlsson, Kevin E. Moore, Erik Hagersten, David A. Wood 0001 |
HPCA | 3 |
| 2003 | Hierarchical Backoff Locks for Nonuniform Communication ArchitecturesabstractThis paper identifies node affinity as an important property for scalable general-purpose locks. Nonuniform communication architectures (NUCA), for example CC-NUMA built from a few large nodes or from chip multiprocessors (CMP), have a lower penalty for reading data from a neighbor's cache than from a remote cache. Lock implementations that encourages handing over locks to neighbors will improve the lock handover time, as well as the access to the critical data guarded by the lock, but will also be vulnerable to starvation. We propose a set of simple software-based hierarchical backoff locks (HBO) that create node affinity in NUCA. A solution for lowering the risk of starvation is also suggested. The HBO locks are compared with other software-based lock implementations using simple benchmarks, and are shown to be very competitive for uncontested locks while being more than twice as fast for contended locks. An application study also demonstrates superior performance for applications with high lock contention and competitive performance for other programs. Zoran Radovic, Erik Hagersten |
HPCA | 2 |
| 2002 | SIP: Performance Tuning through Source Code Interdependence
Erik Berg, Erik Hagersten |
Euro-Par | 2 |
| 2002 | Efficient synchronization for nonuniform communication architecturesabstractScalable parallel computers are often nonuniform communication architectures (NUCAs), where the access time to other processor’s caches vary with their physical location. Still, few attempts of exploring cache-to-cache communication locality have been made. This paper introduces a new kind of synchronization primitives (lock-unlock) that favor neighboring processors when a lock is released. This improves the lock handover time as well as access time to the shared data of the critical region. A critical section guarded by our new RH lock takes less than half the time to execute compared with the same critical section guarded by any other lock on our NUCA hardware. The execution time for Raytrace with 28 processors was improved 2.23 - 4.68 times, while global traffic was dramatically decreased compared with all the other locks. The average execution time was improved 7 - 24% while the global traffic was decreased 8 - 28% for an average over the seven applications studied. Zoran Radovic, Erik Hagersten |
SC | 2 |
| 2001 | Removing the overhead from software-based shared memoryabstractThe implementation presented in this paper---DSZOOM-WF---is a sequentially consistent, fine-grained distributed software-based shared memory. It demonstrates a protocol-handling overhead below a microsecond for all the actions involved in a remote load operation, to be compared to the fastest implementation to date of around ten microseconds.The all-software protocol is implemented assuming some basic low-level primitives in the cluster interconnect and an operating system bypass functionality, similar to the emerging InfiniBand standard. All interrupt- and/or poll-based asynchronous protocol processing is completely removed by running the entire coherence protocol in the requesting processor. This not only removes the asynchronous overhead, but also makes use of a processor that otherwise would stall. The technique is applicable to both page-based and fine-grain software-based shared memory.DSZOOM-WF consistently demonstrates performance comparable to hardware-based distributed shared memory implementations. Zoran Radovic, Erik Hagersten |
SC | 2 |
| 1999 | WildFire: A Scalable Path for SMPsabstractResearchers have searched for scalable alternatives to the symmetric multiprocessor (SMP) architecture since it was first introduced in 1982. The paper introduces an alternative view of the relationship between scalable technologies and SMPs. Instead of replacing large SMPs with scalable technology, we propose new scalable techniques that allow large SMPs to be tied together efficiently, while maintaining the compatibility with, and performance characteristics of, an SMP. The trade-offs of such an architecture differ from those of traditional, scalable, Non-Uniform Memory Architecture (cc-NUMA) approaches. WildFire is a distributed shared memory (DSM) prototype implementation based on large SMPs. It relies on two techniques for creating application-transparent locality: Coherent Memory Replication (CMR), which is a variation of Simple COMA/Reactive NUMA, and Hierarchical Affinity Scheduling (HAS). These two optimizations create extra node locality, which blurs the node boundaries to an application such that SMP-like performance can be achieved with no NUMA-specific optimizations. We present a performance study of a large OLTP benchmark running on DSMs built from various sized nodes and with varying amounts of application-transparent locality. WildFire's measured performance is shown to be more than two times that of an unoptimized NUMA implementation built from small nodes and within 13% of the performance of the ideal implementation: a large SMP with the same access time to its entire shared memory as the local memory access time of WildFire. Erik Hagersten, Michael Koster |
HPCA | 1 |
| 1999 | Parallel computing in the commercial marketplace: research and innovation at workabstractThis is a good time for parallel computer development and research in both academia and industry. The performance improvements predicted by Moore's Law have proven to be quite accurate over many years. However, the doubling of processor performance every 18 months cannot keep up with the growing demand of many applications. The performance of database applications has been doubling every nine to ten months. At last, parallel computer technology has come to play an important role in the commercial marketplace. Multiprocessing has been an active research area for almost 40 years and commercial parallel computers have been available for more than 35 years. After getting off to a slow start, this area has now taken off. Shared memory multiprocessors have dominated this development. This is an area that has sprung out of tireless research and numerous published breakthrough results. The article analyzes some of the reasons for the sudden acceptance of the relatively old parallel computing field, outlines the key properties of a successful parallel computer of the 1990's, and identifies some important research areas and key technologies for the future. Erik Hagersten, Greg Papadopoulos |
Proc. IEEE | 1 |
| 1991 | Race-Free Interconnection Networks and Multiprocessor ConsistencyabstractModernshared-memory multiprocessors require complex interconnection networks to provide sufficient communication bandwidth between processors.They also rely on advanced memory systems that allow multiple memory operations to be made in parallel.It is expensive to maintain a high consistency level in a machine based on a general network, but for special interconnection topologies, some of these costs can be reduced.We define and study one class of interconnection networks, race-free net works.New conditions for sequential consistency are presented which show that sequential consistency can be maintained if all accesses in a multiprocessor can be ordered in an acyclic graph.We show that this can be done in race-free networks without the need for a transaction to be globally performed before the next transaction can be issued.We also investigate what is required to maintain processor consistency in race-free networks.In a race-free network which maintains processor consistency, writes may be pipelined, and reads may bypass writes.The proposed methods reduce the latencies associated with processor write-misses to shared data. Anders Landin, Erik Hagersten, Seif Haridi |
ISCA | 2 |