VLDB 2026 Research / reviewers in the wild / expert
Magnus Jahre
dblp:88/2818
· DBLP profile ↗
39ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0001-9147-5228ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 4 first-author · 14 since 2021Software engineering, systems software and programming languages · 14 · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at Dispatch
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre |
ASPLOS (2) | 4 |
| 2026 | Exploring the Energy Storage and Voltage Control Unit Design Space in Battery-Less IoTabstractBattery-less Internet of Things (IoT) devices are emerging because it will be environmentally and economically unsustainable to power billions of IoT devices with batteries. Battery-less devices harvest energy from their environment, but the energy supplied by many such sources varies considerably over time, and there is hence a need for buffering surplus energy. This is enabled by the Energy Storage and Voltage Control Unit (ESVU), and the state-of-the-art ESVUs are REACT and CapDYN. We observe that an ideal ESVU (i) wastes minimal energy (efficiency), (ii) enables the application System-on-Chip (SoC) quickly when energy becomes available (responsiveness), (iii) can support a large variety of SoCs (applicability), and (iv) requires few or cheap components and occupies minimal area and volume (overhead). Through detailed circuit-level simulations, we demonstrate that no existing ESVU ticks all the boxes. More specifically, we find that CapDYN and REACT can provide high efficiency and responsiveness, but they fall short in applicability and overhead, respectively, and we thus propose Coulombix to fill this gap. To understand how ESVU design affects performance at the system level, we conduct a case study in which we implement hardware prototypes of REACT and Coulombix and measure the throughput and response time of a diverse set of IoT benchmarks with a solar energy harvester across three seasons. The case study supports the insights of the simulation-based study, while also exposing interesting second-order effects, highlighting that end-to-end analysis is critical when evaluating ESVUs. Lukas Liedtke, Espen Holsen, Per Gunnar Kjeldsberg, Frank Alexander Kraemer, Magnus Jahre |
ISPASS | 5 |
| 2026 | FYI: A Foundational Counter Architecture for Building CPI Stacks on Out-of-Order ProcessorsabstractThe end of Dennard scaling and the imminent end of Moore’s law is making it increasingly critical to write software that fully utilizes the hardware compute resources. Software developers hence need to gain insight into the inefficiencies of their applications, and a typical first step is to obtain applicationlevel Cycles-Per-Instruction (CPI) stacks, for example by using a Top-Down methodology. Unfortunately, existing approaches have only been validated against processors with idealized microarchitectural resources. Although this approach adds confidence that the generated CPI stacks make intuitive sense, it does not conclusively document that the CPI stacks break down total execution time into components in an architecturally sound manner. Our goal in this paper is to bridge this knowledge gap. We first define a complete and mutually exclusive reference that accounts for all dispatch slots in all clock cycles across all microoperations ($\mu$ ops), i.e., a $\mu$ op is either (i) dispatching, (ii) stalled while waiting for a structural back-end stall, (iii) not available due to a front-end issue, or (iv) mistakenly executed due to misspeculation. If the reference is also architecturally sound, i.e., it attributes each dispatch slot to the CPI stack component that accurately represents the impact on effective dispatch bandwidth, we refer to it as foundational; second-order overlap effects typically mean that an architecture has multiple foundational references. We derive foundational references for the BOOM out-of-order core, and, surprisingly perhaps, find that they can be implemented in hardware with less than 1 KB of additional state. We hence propose the Foundational Yet Implementable (FYI) counter architecture for CPI stacks that are, by design, an exact match for the target reference. Silvio Heverton Campelo de Santana, Lieven Eeckhout, Magnus Jahre |
ISPASS | 3 |
| 2026 | Pesto: Diagnosing Performance Pathologies in Out-of-Order ProcessorsabstractHigh-performance parallel acceleration necessitates high-performance CPUs to stave off Amdahl’s law. Getting the most out of a CPU requires instruction-level performance profiling to understand microarchitecture performance bottlenecks, guiding hardware design for future generations. This depends on effective methods for using instruction profilers to identify the key ways performance losses occur in out-of-order processors. Identifying the root causes of these losses is challenging, but can be greatly assisted by comprehending how they manifest behaviorally in hardware, which we refer to in this work as ‘performance pathologies’. Historically, characterizing these behaviors has been a labor-intensive process that required a detailed understanding of software and microarchitecture. However, the advent of profilers capable of producing accurate instruction-level cycle stacks at both ends of the execution window, i.e., at dispatch and commit, dramatically simplifies this task. In this work, we present Pesto, an instruction-level methodology for diagnosing microarchitecture performance pathologies, i.e., identifying the key microarchitectural manifestations of performance loss in out-of-order processors. Pesto uses k-means clustering to group per-instruction cycle stack components produced at both ends of the execution window. The resulting centroids for each cluster characterize exactly one kind of pathology (e.g., pipeline flush due to branch misprediction, or commit stall due to data cache misses). We apply Pesto to an FPGA-accelerated BOOM core in FireSim running SPEC2017 and discover that the BOOM core has 20 pathologies, corresponding to 20 fundamental ways its microarchitecture can lose performance. Furthermore, we observe that individual instruction types (e.g., loads) lead to performance loss in at most four ways. Finally, we show that Pesto works using statistical sampling, making it easy to adopt. Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre |
ISPASS | 4 |
| 2026 | EStacker: Explaining Battery-Less IoT System Performance with Energy StacksabstractThe number of Internet of Things (IoT) devices is increasing exponentially, and it is environmentally and economically unsustainable to power all these devices with batteries. The key alternative is energy harvesting, but battery-less IoT systems require extensive evaluation to demonstrate that they are sufficiently performant across the full range of expected operating conditions. IoT developers thus need an evaluation platform that (i) ensures that each evaluated application and configuration is exposed to exactly the same energy environment and events, and (ii) provides a detailed account of what the application spends the harvested energy on. We therefore developed the EStacker evaluation platform which (i) enables fair and repeatable evaluation, and (ii) generates energy stacks. Energy stacks break down the total energy consumption of an application across hardware components and application activities, thereby explaining what the application specifically uses energy on. We augment EStacker with the ST-SP optimization which, in our experiments, reduces evaluation time by 6.3× on average while retaining the temporal behavior of the battery-less IoT system (average throughput error of 7.7%) by proportionally scaling time and power. We demonstrate the utility of EStacker through two case studies. In the first case study, we use energy stack profiles to identify a performance problem that, once addressed, improves performance by 3.3×. The second case study focuses on ST-SP, and we use it to explore the design space required to dimension the harvester and energy storage sizes of a smart parking application in roughly one week (7.7 days). Without ST-SP, sweeping this design space would have taken well over one month (41.7 days). Lukas Liedtke, Per Gunnar Kjeldsberg, Frank Alexander Kraemer, Magnus Jahre |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | HILP: Accounting for Workload-Level Parallelism in System-on-Chip Design Space ExplorationabstractHigh-performance System-on-Chip (SoC) architectures are becoming increasingly complex and heterogeneous, and the days when a single application could utilize all of an SoC’s hardware resources are all but over. The SoC’s workload, i.e., the set of independent applications that the SoC typically executes, therefore has a significant impact on its efficiency. Accounting for Workload-Level Parallelism (WLP) in early-stage design space exploration is thus critical as later-stage analysis steps must focus on favorable design points to yield optimal results. Unfortunately, state-of-the-art MultiAmdahl and Gables fall short because they only model the extremes of minimal and maximal WLP. We hence propose HILP, the first early-stage design space exploration approach for heterogeneous SoCs that fully accounts for WLP. Our key observation is that scheduling a workload of independent multi-phase applications on a heterogeneous SoC is an instance of the classic job-shop scheduling optimization problem and thus can be solved using integer linear programming. HILP therefore uses a high-performance integer linear programming solver to find a near-optimal schedule that minimizes the overall execution time of the workload, i.e., it schedules the dependent phases of all applications in the workload on the cores and accelerators of the target SoC to maximize performance while respecting power consumption and memory bandwidth constraints. We validate HILP by demonstrating that it captures the performance effects of Amdahl’s law, the memory wall, and dark silicon, and then use it to explore the impact of WLP across a large SoC design space, yielding multiple insights. The key takeaway is that modeling WLP is necessary to ensure that more detailed, later-stage design tasks focus on the most favorable parts of the vast design space of heterogeneous SoCs. Joseph Rogers, Lieven Eeckhout, Magnus Jahre |
HPCA | 3 |
| 2025 | Neoscope: How Resilient Is My SoC to Workload Churn?abstractThe lifetime of hardware is increasing, but the lifetime of software is not.This leads to devices that, while performant when released, have fall-off due to changing workload suitability.To ensure that performance is maintained, computer architects must begin considering the effects of evolving workloads early in the design process.The ways that workloads can change, i.e., churn, over time profoundly affect hardware's ability to maintain performance.To better understand churn, we introduce the concepts Magnitude and Disruption, which enable how workloads change to be quantitatively described.By leveraging these terms, we present churn as a span, where the churn of a given workload can be categorized as either Minimal, Perturbing, Escalating, or Volatile.To account for how churn affects performance, we propose Neoscope, the first multi-objective pre-silicon design space exploration tool for investigating System-on-Chip (SoC) architectures that are resilient to workload churn.Neoscope uses integer linear programming and concepts from job-shop scheduling to construct near-optimal SoCs for a given workload.Unlike previous methods, Neoscope approaches (and often finds) the globally optimal SoC with a single invocation, instead of needing to parameter sweep.Neoscope is also multi-objective, i.e., it can optimize for many kinds of metrics other than absolute performance, including area, energy, cost, and carbon efficiencies.Using Neoscope, we explore near-optimal SoCs for many workload-churn configurations, and investigate how changing the objective function changes the ideal SoC.Neoscope shows that small SoCs need high levels of specialization, but that this is risky, as it increases susceptibility to churn. Joseph Rogers, Lieven Eeckhout, Taha Soliman, Magnus Jahre |
ISCA | 4 |
| 2025 | FOMOsim: An open-source simulator for rigorous analysis of micromobility planning problemsabstractExisting simulation models for micromobility systems often face significant limitations: they are typically custom-built for specific contexts, lack generalizability , are not open-source, and undergo limited testing, which again question their reliability and applicability in broader settings. To address these challenges, this paper presents FOMOsim, an open-source simulator developed to analyze micromobility systems. The simulator’s modular design allows for better scalability and usability, facilitating detailed and realistic simulations of shared transport operations. New benchmark instances, utilizing data from bike-sharing systems (BSSs) in major American and Norwegian cities, are created and embedded within the simulator. Additionally, we introduce a new greedy heuristic for the dynamic bike rebalancing problem, integrated within FOMOsim. Rigorous testing demonstrates that the proposed heuristic, when combined with FOMOsim, outperforms state-of-the-art methods, significantly improving BSS performance by reducing bike shortages and surpluses. Furthermore, a comprehensive experimental study is designed to assess the impact of key strategic and operational decisions on BSS performance. The findings underscore the simulator’s adaptability in addressing various planning challenges and its potential to improve BSS management through informed decision-making. Consequently, FOMOsim provides the research community with a robust platform for replicating experiments, generating new instances, and exploring innovative solutions in BSS management, thereby substantially enriching the field’s body of knowledge. Steffen J. S. Bakker, Benahmed Mohammed, Asbjørn Djupdal, Lasse Natvig, Henrik Andersson 0001, Magnus Jahre, Kjetil Fagerholt |
Expert Syst. Appl. | 6 |
| 2024 | ECM: Improving IoT Throughput with Energy-Aware Connection ManagementabstractDesigning Internet of Things (IoT) devices that solely rely on energy harvesting is the most promising approach towards achieving a scalable and sustainable IoT. The power output of energy harvesters can however vary significantly and maximizing throughput hence requires adapting application behavior to match the harvester's current power output. In this work, we focus on the connection policy of the IoT device and find that the on-demand connect policy - which is used by state-of-the-art IoT runtime systems - and the aggressive maintain connection policy both fall short across a broad range of harvester power outputs. We therefore propose Energy-aware Connection Management (ECM) which tunes the connection policy and sampling frequency to consistently achieve high throughput. ECM accomplishes this by predicting both the average power output of the harvester and the energy consumed by the IoT device with a lightweight analytical model that only requires tracking six energy thresholds. Our evaluation demonstrates that ECM can improve throughput substantially, i.e., by up to 9.5 × and 3.0 × compared to the on-demand connect and maintain connection policies, respectively. Lukas Liedtke, Magnus Jahre |
DATE | 3 |
| 2024 | AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesabstractSpecialization is the key approach for continued performance growth beyond the end of Dennard scaling. Academics and industry are hence continuously proposing new accelerator architectures, including conventional Domain-Specific Accelerators (DSAs) and emerging Processing in Memory (PIM) accelerators. We are thus fast approaching an era in which earlystage accelerator analysis is critical for maintaining the productivity of software developers, system software designers, and computer architects — to ensure that they focus time-consuming implementation and optimization efforts on the most favorable class of accelerators for the problem at hand. Unfortunately, existing approaches fall short because they either adopt a level of abstraction that is too high — and therefore are unable to account for key performance phenomena — or too low — because they focus on details that do not generalize across diverse accelerators. Our Architecture-Independent Operation (AIO) abstraction addresses this issue by leveraging that accelerators typically focus on data-level parallelism, and an AIO is hence a key piece of algorithm-level data-parallel work that remains the same across diverse accelerators. To demonstrate that the AIO abstraction can be accurate and useful, we create the AccMe performance model which predicts kernel performance by estimating the number of clock cycles spent on compute, memory, and invocation overhead while accounting for overlap between compute and memory cycles as well as finite memory bandwidth. We demonstrate that AccMe can be accurate, i.e., it yields an average performance prediction error of $5.6 \%$ across our diverse kernels and accelerators. This is a significant improvement over the $\mathbf{2 0. 6 \%}$ average error of curve-fitted Roofline which provides the best-case accuracy of Roofline’s operational intensity abstraction. We further demonstrate that AccMe is useful through three case studies that illustrate (i) how developers can use AccMe for accelerator selection under uncertainty; (ii) how system software can use AccMe for scheduling — and thereby improve throughput by $2.8 \times$ on average compared to Roofline-driven scheduling; and (iii) how computer architects can use AccMe for architectural exploration. Joseph Rogers, Taha Soliman, Magnus Jahre |
ISCA | 3 |
| 2024 | Temporarily Unauthorized Stores: Write First, Ask for Permission Laterabstractx86 processors implement a total store order (x86-TSO) consistency model, which requires stores to update memory in a sequenced manner. The latency of stores is then hidden by the store buffer (SB), which holds stores until the write is performed. On a long latency cache miss, however, stores block the SB, eventually stalling the processor and degrading performance. Contemporary industrial high-performance processors deal with this situation by overprovisioning the size of the SB, but this comes at the cost of energy and latency overheads. In this work, we remove the stalls caused by stores blocked at the head of the SB while reusing existing processor resources, either improving performance when SB size is kept constant or maintaining performance while reducing SB size. Our proposal, Temporarily Unauthorized Stores (TUS), achieves this by extending the functionality of 1) the write combining buffers, to allow them to coalesce stores while maintaining x86- TSO consistency, and 2) immediately write data to the first-level cache upon a miss (i.e., providing an always-hit illusion) but temporarily keeping the written data invisible to the cache coherence protocol, i.e., these stores are temporarily unauthorized. TUS makes temporarily unauthorized stores visible in x86- TSO order without speculation or rollbacks once write permission is obtained. In essence, TUS logically transforms the write combining buffers and the first-level cache into an “extension” of the SB. TUS improves performance by up to 26 % (3.2 % on average) while reducing the total energy-delay-product (EDP) by up to 35.9% (6.4% on average) for SB-bound benchmarks with a 114-entry SB compared to our baseline architecture with an SB of the same size. When configured with a 32-entry SB, TUS yields a performance improvement of 2 % over a 114-entry SB baseline while reducing SB energy per search by a factor of 2 x, SB area by 21 %, and store-to-Ioad forwarding latency from 5 to 3 cycles. Juan M. Cebrian, Magnus Jahre, Alberto Ros 0001 |
MICRO | 2 |
| 2023 | NUBA: Non-Uniform Bandwidth GPUsabstractThe parallel execution model of GPUs enables scaling to hundreds of thousands of threads, which is a key capability that many modern high-performance applications exploit. GPU vendors are hence increasing the compute and memory resources with every GPU generation — resulting in the need to efficiently stitch together a plethora of Symmetric Multiprocessors (SMs), Last-Level Cache (LLC) slices and memory controllers while maximizing bandwidth and keeping power consumption and design complexity in check. Conventional GPUs are Uniform Bandwidth Architectures (UBAs) as they provide equal bandwidth between all SMs and all LLC slices. UBA GPUs require a uniform high-bandwidth Network-on-Chip (NoC), and our key observation is that provisioning a NoC to match the LLC slice bandwidth incurs a hefty power and complexity overhead. We propose the Non-Uniform Bandwidth Architecture (NUBA), a GPU system architecture aimed at fully utilizing LLC slice bandwidth. A NUBA GPU consists of partitions that each feature a few SMs and LLC slices as well as a memory controller — hence exposing the complete LLC bandwidth to the SMs within a partition since they can be connected with point-to-point links — and a NoC between partitions — to enable access to remote data.Exploiting the potential of NUBA GPUs however requires carefully co-designing system software, the compiler and architectural policies. The critical system software component is our Local-And-Balanced (LAB) page placement policy which enables the GPU driver to place data in local partitions while avoiding load imbalance. Moreover, we propose Model-Driven Replication (MDR) which identifies read-only shared data with data-flow analysis at compile time. At run time, MDR leverages an architectural mechanism that replicates read-only shared data across LLC slices when this can be done without pressuring cache capacity. With LAB and MDR, our NUBA GPU improves average performance by 23.1% and 22.2% (and up to 183.9% and 182.4%) compared to iso-resource memory-side and SM-side UBA GPUs, respectively. When the NUBA concept is leveraged to reduce overhead while maintaining similar performance, NUBA reduces NoC power consumption by 12.1× and 9.4×, respectively. Xia Zhao 0004, Magnus Jahre, Yuhua Tang, Guangda Zhang, Lieven Eeckhout |
ASPLOS (2) | 2 |
| 2023 | TEA: Time-Proportional Event AnalysisabstractAs computer architectures become increasingly complex and heterogeneous, it becomes progressively more difficult to write applications that make good use of hardware resources. Performance analysis tools are hence critically important as they are the only way through which developers can gain insight into the reasons why their application performs as it does. State-of-the-art performance analysis tools capture a plethora of performance events and are practically non-intrusive, but performance optimization is still extremely challenging. We believe that the fundamental reason is that current state-of-the-art tools in general cannot explain why executing the application's performance-critical instructions take time. Björn Gottschall, Lieven Eeckhout, Magnus Jahre |
ISCA | 3 |
| 2023 | SAC: Sharing-Aware Caching in Multi-Chip GPUsabstractBandwidth non-uniformity in multi-chip GPUs poses a major design challenge for its last-level cache (LLC) architecture. Whereas a memory-side LLC caches data from the local memory partition while being accessible by all chips, an SM-side LLC is private to a chip while caching data from all memory partitions. We find that some workloads prefer a memory-side LLC while others prefer an SM-side LLC, and this preference solely depends on which organization maximizes the effective LLC bandwidth. In contrast to prior work which optimizes bandwidth beyond the LLC, we make the observation that the effective bandwidth ahead of the LLC is critical to end-to-end application performance. We propose Sharing-Aware Caching (SAC) to adopt either a memory-side or SM-side LLC organization by dynamically reconfiguring the routing policies in the intra-chip interconnection network and LLC controllers. SAC is driven by a simple and lightweight analytical model that predicts the impact of data sharing across chips on the effective LLC bandwidth. SAC improves average performance by 76% and 12% (and up to 157% and 49%) compared to a memory-side and SM-side LLC, respectively. We demonstrate significant performance improvements across the design space and across workloads. Shiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven Eeckhout |
ISCA | 3 |
| 2023 | PES: An Energy and Throughput Model for Energy Harvesting IoT SystemsabstractThe Internet of Things (IoT) requires ultra-low-power sensor platforms that can be deployed at scale. Scalable systems however cannot be battery-powered because replacing batteries at scale is costly, impractical and has a negative impact on the environment. The key alternative is to rely on energy harvesting, but this is challenging because the developer needs to ensure that the application achieves sufficient throughput, i.e., the sensor platform delivers information to the back-end system at a sufficient rate when provided with a certain amount of energy. We hence propose the analytical Periodic Energy Harvesting Systems (PES) model which enables developers to explore energy versus throughput trade-offs early in the design process – thereby enabling developers to select a reasonable ultralow-power platform and energy harvesting technology before incurring the (significant)) overhead of adapting their application to the particularities of the platform. PES faithfully models the energy consumed by an IoT application during sampling and communication as well as while idle between samples. If the average power output of the energy harvesting subsystem is insufficient to sustain the application, PES uses the mismatch to predict the number of shutdowns required to harvest sufficient energy. PES achieves an average error of 3.6% across the IoT applications we consider in this work; a significant improvement over the 87.3% average error of state-of-the art EH. Lukas Liedtke, Magnus Jahre |
ISPASS | 3 |
| 2023 | Near-optimal multi-accelerator architectures for predictive maintenance at the edge
Mostafa Koraei, Juan M. Cebrian, Magnus Jahre |
Future Gener. Comput. Syst. | 3 |
| 2023 | Characterizing Multi-Chip GPU Data SharingabstractMulti-chip Graphics Processing Unit (GPU) systems are critical to scale performance beyond a single GPU chip for a wide variety of important emerging applications. A key challenge for multi-chip GPUs, though, is how to overcome the bandwidth gap between inter-chip and intra-chip communication. Accesses to shared data, i.e., data accessed by multiple chips, pose a major performance challenge as they incur remote memory accesses possibly congesting the inter-chip links and degrading overall system performance. This article characterizes the shared dataset in multi-chip GPUs in terms of (1) truly versus falsely shared data, (2) how the shared dataset scales with input size, (3) along which dimensions the shared dataset scales, and (4) how sensitive the shared dataset is with respect to the input’s characteristics, i.e., node degree and connectivity in graph workloads. We observe significant variety in scaling behavior across workloads: some workloads feature a shared dataset that scales linearly with input size, whereas others feature sublinear scaling (following a \(\sqrt {2}\) or \(\sqrt [3]{2}\) relationship). We further demonstrate how the shared dataset affects the optimum last-level cache organization (memory-side versus SM-side) in multi-chip GPUs, as well as optimum memory page allocation and thread scheduling policy. Sensitivity analyses demonstrate the insights across the broad design space. Shiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven Eeckhout |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Delegated Replies: Alleviating Network Clogging in Heterogeneous ArchitecturesabstractHeterogeneous architectures with latency-sensitive CPU cores and bandwidth-intensive accelerators are attractive as they deliver high performance at favorable cost. These architectures typically have significantly more compute cores than memory nodes. The many bandwidth-intensive accelerators hence overwhelm the few memory nodes, resulting in suboptimal accelerator performance — as their bandwidth needs are not met — and poor CPU performance — because memory node blocking creates high latencies. We call this phenomenon network clogging. Since network clogging is a widespread issue in heterogeneous architectures, we first investigate if existing state-of-the-art approaches can address it. We find that the most effective prior approach, called Realistic Probing (RP), is suboptimal because it searches the local caches of other cores for missing data.We propose Delegated Replies which lets memory nodes speculatively delegate the responsibility of replying to last-level cache hits to the private cache that last accessed the requested cache block, hence avoiding the search that fundamentally limits RP. Moreover, Delegated Replies uses the (typically) under-utilized request network for delegation; it is the reply network links of the memory nodes that commonly clog because replies include complete cache blocks in addition to metadata. We evaluate Delegated Replies in the context of heterogeneous architectures with latency-sensitive CPU cores and bandwidth-intensive GPU cores and find that it improves GPU (CPU) performance by 14.2% (5.2%) and 25.7% (8.8%) on average compared to RP and our baseline, respectively. Xia Zhao 0004, Lieven Eeckhout, Magnus Jahre |
HPCA | 3 |
| 2021 | TIP: Time-Proportional Instruction ProfilingabstractA fundamental part of developing software is to understand what the application spends time on. This is typically determined using a performance profiler which essentially captures how execution time is distributed across the instructions of a program. At the same time, the highly parallel execution model of modern high-performance processors means that it is difficult to reliably attribute time to instructions — resulting in performance analysis being unnecessarily challenging. Björn Gottschall, Lieven Eeckhout, Magnus Jahre |
MICRO | 3 |
| 2020 | HSM: A Hybrid Slowdown Model for Multitasking GPUsabstractGraphics Processing Units (GPUs) are increasingly widely used in the cloud to accelerate compute-heavy tasks. However, GPU-compute applications stress the GPU architecture in different ways --- leading to suboptimal resource utilization when a single GPU is used to run a single application. One solution is to use the GPU in a multitasking fashion to improve utilization. Unfortunately, multitasking leads to destructive interference between co-running applications which causes fairness issues and Quality-of-Service (QoS) violations. Xia Zhao 0004, Magnus Jahre, Lieven Eeckhout |
ASPLOS | 2 |
| 2020 | Selective Replication in Memory-Side GPU CachesabstractData-intensive applications put immense strain on the memory systems of Graphics Processing Units (GPUs). To cater to this need, GPU memory systems distribute requests across independent units to provide high bandwidth by servicing requests (mostly) in parallel. We find that this strategy breaks down for shared data structures because the shared Last-Level Cache (LLC) organization used by contemporary GPUs stores shared data in a single LLC slice. Shared data requests are hence serialized - resulting in data-intensive applications not being provided with the bandwidth they require. A private LLC organization can provide high bandwidth, but it is often undesirable since it significantly reduces the effective LLC capacity. In this work, we propose the Selective Replication (SelRep) LLC which selectively replicates shared read-only data across LLC slices to improve bandwidth supply while ensuring that the LLC retains sufficient capacity to keep shared data cached. The compile-time component of SelRep LLC uses dataflow analysis to identify read-only shared data structures and uses a special-purpose load instruction for these accesses. The runtime component of SelRep LLC then monitors the caching behavior of these loads. Leveraging an analytical model, SelRep LLC chooses a replication degree that carefully balances the effective LLC bandwidth benefits of replication against its capacity cost. SelRep LLC consistently provides high performance to replication-sensitive applications across different data set sizes. More specifically, SelRep LLC improves performance by 19.7% and 11.1% on average (and up to 61.6% and 31.0%) compared to the shared LLC baseline and the state-of-the-art Adaptive LLC, respectively. Xia Zhao 0004, Magnus Jahre, Lieven Eeckhout |
MICRO | 2 |
| 2020 | MDM: The GPU Memory Divergence ModelabstractAnalytical models enable architects to carry out early-stage design space exploration several orders of magnitude faster than cycle-accurate simulation by capturing first-order performance phenomena with a set of mathematical equations. However, this speed advantage is void if the conclusions obtained through the model are misleading due to model inaccuracies. Therefore, a practical analytical model needs to be sufficiently accurate to capture key performance trends across a broad range of applications and architectural configurations.In this work, we focus on analytically modeling the performance of emerging memory-divergent GPU-compute applications which are common in domains such as machine learning and data analytics. The poor spatial locality of these applications leads to frequent L1 cache blocking due to the application issuing significantly more concurrent cache misses than the cache can support, which cripples the GPU’s ability to use Thread-Level Parallelism (TLP) to hide memory latencies. We propose the GPU Memory Divergence Model (MDM) which faithfully captures the key performance characteristics of memory-divergent applications, including memory request batching and excessive NoC/DRAM queueing delays. We validate MDM against detailed simulation and real hardware, and report substantial improvements in (1) scope: the ability to model prevalent memory-divergent applications in addition to non-memory divergent applications; (2) practicality: 6.1× faster by computing model inputs using binary instrumentation as opposed to functional simulation; and (3) accuracy: 13.9% average prediction error versus 162% for the state-of-the-art GPUMech model. Lu Wang 0019, Magnus Jahre, Almutaz Adileh, Lieven Eeckhout |
MICRO | 2 |
| 2020 | DCMI: A Scalable Strategy for Accelerating Iterative Stencil Loops on FPGAsabstractIterative Stencil Loops (ISLs) are the key kernel within a range of compute-intensive applications. To accelerate ISLs with Field Programmable Gate Arrays, it is critical to exploit parallelism (1) among elements within the same iteration and (2) across loop iterations. We propose a novel ISL acceleration scheme called Direct Computation of Multiple Iterations (DCMI) that improves upon prior work by pre-computing the effective stencil coefficients after a number of iterations at design time—resulting in accelerators that use minimal on-chip memory and avoid redundant computation. This enables DCMI to improve throughput by up to 7.7× compared to the state-of-the-art cone-based architecture. Mostafa Koraei, Omid Fatemi, Magnus Jahre |
ACM Trans. Archit. Code Optim. | 3 |
| 2020 | Scalability analysis of AVX-512 extensions
Juan M. Cebrian, Lasse Natvig, Magnus Jahre |
J. Supercomput. | 3 |
| 2018 | GDP: Using Dataflow Properties to Accurately Estimate Interference-Free Performance at RuntimeabstractMulti-core memory systems commonly share resources between processors. Resource sharing improves utilization at the cost of increased inter-application interference which may lead to priority inversion, missed deadlines and unpredictable interactive performance. A key component to effectively manage multi-core resources is performance accounting which aims to accurately estimate interference-free application performance. Previously proposed accounting systems are either invasive or transparent. Invasive accounting systems can be accurate, but slow down latency-sensitive processes. Transparent accounting systems do not affect performance, but tend to provide less accurate performance estimates. We propose a novel class of performance accounting systems that achieve both performance-transparency and superior accuracy. We call the approach dataflow accounting, and the key idea is to track dynamic dataflow properties and use these to estimate interference-free performance. Our main contribution is Graph-based Dynamic Performance (GDP) accounting. GDP dynamically builds a dataflow graph of load requests and periods where the processor commits instructions. This graph concisely represents the relationship between memory loads and forward progress in program execution. More specifically, GDP estimates interference-free stall cycles by multiplying the critical path length of the dataflow graph with the estimated interference-free memory latency. GDP is very accurate with mean IPC estimation errors of 3.4% and 9.8% for our 4- and 8-core processors, respectively. When GDP is used in a cache partitioning policy, we observe average system throughput improvements of 11.9% and 20.8% compared to partitioning using the state-of-the-art Application Slowdown Model. Magnus Jahre, Lieven Eeckhout |
HPCA | 1 |
| 2018 | Get Out of the Valley: Power-Efficient Address Mapping for GPUsabstractGPU memory systems adopt a multi-dimensional hardware structure to provide the bandwidth necessary to support 100s to 1000s of concurrent threads. On the software side, GPU-compute workloads also use multi-dimensional structures to organize the threads. We observe that these structures can combine unfavorably and create significant resource imbalance in the memory subsystem - causing low performance and poor power-efficiency. The key issue is that it is highly application-dependent which memory address bits exhibit high variability. To solve this problem, we first provide an entropy analysis approach tailored for the highly concurrent memory request behavior in GPU-compute workloads. Our window-based entropy metric captures the information content of each address bit of the memory requests that are likely to co-exist in the memory system at runtime. Using this metric, we find that GPU-compute workloads exhibit entropy valleys distributed throughout the lower order address bits. This indicates that efficient GPU-address mapping schemes need to harvest entropy from broad address-bit ranges and concentrate the entropy into the bits used for channel and bank selection in the memory subsystem. This insight leads us to propose the Page Address Entropy (PAE) mapping scheme which concentrates the entropy of the row, channel and bank bits of the input address into the bank and channel bits of the output address. PAE maps straightforwardly to hardware and can be implemented with a tree of XOR-gates. PAE improves performance by 1.31X and power-efficiency by 1.25X compared to state-of-the-art permutation-based address mapping. Xia Zhao 0004, Magnus Jahre, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ISCA | 3 |
| 2017 | Towards efficient quantized neural network inference on mobile devices: work-in-progressabstractFrom voice recognition to object detection, Deep Neural Networks (DNNs) are steadily getting better at extracting information from complex raw data. Combined with the popularity of mobile computing and the rise of the Internet-of-Things (IoT), there is enormous potential for widespread deployment of intelligent devices, but a computational challenge remains. A modern DNN can require billions of floating point operations to classify a single image, which is far too costly for energy-constrained mobile devices. Offloading DNNs to powerful servers in the cloud is only a limited solution, as it requires significant energy for data transfer and cannot address applications with low-latency requirements such as augmented reality or navigation for autonomous drones. Yaman Umuroglu, Magnus Jahre |
CASES | 2 |
| 2017 | Towards Efficient Design Space Exploration of FPGA-based Accelerators for Streaming HPC Applications (Abstract Only)
Mostafa Koraei, Magnus Jahre, Omid Fatemi |
FPGA | 2 |
| 2017 | FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong, Magnus Jahre, Kees A. Vissers |
FPGA | 6 |
| 2015 | Hybrid breadth-first search on a single-chip FPGA-CPU heterogeneous platformabstractLarge and sparse small-world graphs are ubiquitous across many scientific domains from bioinformatics to computer science. As these graphs grow in scale, traversal algorithms such as breadth-first search (BFS), fundamental to many graph processing applications and metrics, become more costly to compute. The cause is attributed to poor temporal and spatial locality due to the inherently irregular memory access patterns of these algorithms. A large body of research has targeted accelerating and parallelizing BFS on a variety of computing platforms, including hybrid CPU-GPU approaches for exploiting the small-world property. In the same spirit, we show how a single-die FPGA-CPU heterogeneous device can be used to leverage the varying degree of parallelism in small-world graphs. Additionally, we demonstrate how dense rather than sparse treatment of the BFS frontier vector yields simpler memory access patterns for BFS, trading redundant computation for DRAM bandwidth utilization and faster graph exploration. On a range of synthetic small-world graphs, our hybrid approach performs 7.8× better than software-only and 2× better than accelerator-only implementations. We achieve an average traversal speed of 172 MTEPS (millions of traversed edges per second) on the ZedBoard platform, which is more than twice as effective as the best previously published FPGA BFS implementation in terms of traversals per bandwidth. Yaman Umuroglu, Donn Morrison, Magnus Jahre |
FPL | 3 |
| 2015 | Tuning the victim selection policy of Intel TBB
Alexandru C. Iordan, Magnus Jahre, Lasse Natvig |
J. Syst. Archit. | 2 |
| 2014 | Graph-based performance accounting for chip multiprocessor memory systemsabstractChip Multiprocessor (CMP) memory systems share memory system resources between processor cores. While this sharing enables good resource utilization and fast inter-processor communication, it also makes the performance of an application depend on its co-runners. This breaks the system software assumption that a process has the same rate of progress regardless of the co-schedule, potentially leading to priority inversion, missed deadlines, unpredictable interactive performance and non-compliance with service level agreements. In this work, we present a novel graph-based technique that accurately estimates the performance an application would experience without memory system interference. Dynamic interference-free performance estimates can enable scheduling algorithms and management policies that optimize directly for system performance metrics. Magnus Jahre |
PACT | 1 |
| 2014 | An energy efficient column-major backend for FPGA SpMV acceleratorsabstractFPGAs are promising candidates for energy efficient acceleration of sparse matrix-vector multiplication (SpMV), a kernel with important applications in scientific computing and engineering. SpMV is characterized by matrix-dependent performance and high external memory bandwidth demands, which makes bandwidth utilization an important performance indicator. Existing FPGA SpMV accelerators focus on datapath optimizations instead of memory behavior, and exhibit matrix-dependent bandwidth utilization. In this work, we propose to decouple the SpMV computation and memory behavior, and focus on the backend which handles the latter. We describe a scalable backend architecture that exploits column-major traversal and interleaving to achieve high bandwidth utilization. Our experiments show that a single backend is able to sustain 96% of its assigned memory port bandwidth on average, and scales well with increased bandwidth by instantiating multiple parallel units. Compared to a baseline scheme, our scheme offers up to 1.5× higher DRAM power efficiency and up to 20% higher aggregate bandwidth. The results indicate that our scheme improves the average bandwidth utilization of existing FPGA SpMV accelerators by 15 to 77%. Yaman Umuroglu, Magnus Jahre |
ICCD | 2 |
| 2014 | Optimized hardware for suboptimal software: The case for SIMD-aware benchmarksabstractEvaluation of new architectural proposals against real applications is a necessary step in academic research. However, providing benchmarks that keep up with new architectural changes has become a real challenge. If benchmarks don't cover the most common architectural features, architects may end up under/over estimating the impact of their contributions. In this work, we extend the PARSEC benchmark suite with SIMD capabilities to provide an enhanced evaluation framework for new academic/industry proposals. We then perform a detailed energy and performance evaluation of this commonly used application set on different platforms (Intel®and ARM®processors). Our results show how SIMD code alters scalability, energy efficiency and hardware requirements. Performance and energy efficiency improvements depend greatly on the fraction of code that we can actually vectorize (up to 50×). Our enhancements are based in a custom built wrapper library compatible with SSE, AVX and NEON to facilitate general vectorization. We aim to distribute the source code to reinforce the evaluation process of new proposals for computing systems. Juan M. Cebrian, Magnus Jahre, Lasse Natvig |
ISPASS | 2 |
| 2014 | Perfect Reconstructability of Control Flow from Demand Dependence GraphsabstractDemand-based dependence graphs (DDGs), such as the (Regionalized) Value State Dependence Graph ((R)VSDG), are intermediate representations (IRs) well suited for a wide range of program transformations. They explicitly model the flow of data and state, and only implicitly represent a restricted form of control flow. These features make DDGs especially suitable for automatic parallelization and vectorization, but cannot be leveraged by practical compilers without efficient construction and destruction algorithms. Construction algorithms remodel the arbitrarily complex control flow of a procedure to make it amenable to DDG representation, whereas destruction algorithms reestablish control flow for generating efficient object code. Existing literature presents solutions to both problems, but these impose structural constraints on the generatable control flow, and omit qualitative evaluation. The key contribution of this article is to show that there is no intrinsic structural limitation in the control flow directly extractable from RVSDGs. This fundamental result originates from an interpretation of loop repetition and decision predicates as computed continuations, leading to the introduction of the predicate continuation normal form. We provide an algorithm for constructing RVSDGs in predicate continuation form, and propose a novel destruction algorithm for RVSDGs in this form. Our destruction algorithm can generate arbitrarily complex control flow; we show this by proving that the original CFG an RVSDG was derived from can, apart from overspecific detail, be reconstructed perfectly. Additionally, we prove termination and correctness of these algorithms. Furthermore, we empirically evaluate the performance, the representational overhead at compile time, and the reduction in branch instructions compared to existing solutions. In contrast to previous work, our algorithms impose no additional overhead on the control flow of the produced object code. To our knowledge, this is the first scheme that allows the original control flow of a procedure to be recovered from a DDG representation. Helge Bahmann, Nico Reissmann, Magnus Jahre, Jan Christian Meyer |
ACM Trans. Archit. Code Optim. | 3 |
| 2010 | Multi-level Hardware Prefetching Using Low Complexity Delta Correlating Prediction Tables with Partial Matching
Marius Grannæs, Magnus Jahre, Lasse Natvig |
HiPEAC | 2 |
| 2010 | DIEF: An Accurate Interference Feedback Mechanism for Chip Multiprocessor Memory Systems
Magnus Jahre, Marius Grannæs, Lasse Natvig |
HiPEAC | 1 |
| 2009 | A Quantitative Study of Memory System Interference in Chip Multiprocessor ArchitecturesabstractThe potential for destructive interference between running processes is increased as Chip Multiprocessors (CMPs) share more on-chip resources. We believe that understanding the nature of memory system interference is vital to achieve good fairness/complexity/performance trade-offs in CMPs. Our goal in this work is to quantify the latency penalties due to interference in all hardware-controlled, shared units (i.e. the on-chip interconnect, shared cache and memory bus). To achieve this, we simulate a wide variety of realistic CMP architectures. In particular, we vary the number of cores, interconnect topology, shared cache size and off-chip memory bandwidth. We observe that interference in the off-chip memory bus accounts for between 63% and 87% of the total interference impact while the impact of cache capacity interference can be lower than indicated by previous studies (between 5% and 32% of the total impact). In addition, as much as 11% of the total impact can be due to uncontrolled allocation of shared cache Miss Status Holding Registers (MSHRs). Magnus Jahre, Marius Grannæs, Lasse Natvig |
HPCC | 1 |
| 2008 | Low-cost open-page prefetch scheduling in chip multiprocessorsabstractThe pressure on off-chip memory increases significantly as more cores compete for the same resources. A CMP deals with the memory wall by exploiting thread level parallelism (TLP), shifting the focus from reducing overall memory latency to memory throughput. This extends to the memory controller where the 3D structure of modern DRAM is exploited to increase throughput. Traditionally, prefetching reduces latency by fetching data before it is needed. In this paper we explore how prefetching can be used to increase memory throughput. We present our own low-cost open-page prefetch scheduler that exploits the 3D structure of DRAM when issuing prefetches. We show that because of the complex structure of modern DRAM, prefetches can be made cheaper than ordinary reads, thus making prefetching beneficial even when prefetcher accuracy is low. As a result, prefetching with good coverage is more important than high accuracy. By exploiting this observation our low-cost open page scheme increases performance and QoS. Furthermore, we explore how prefetches should be scheduled in a state of the art memory controller by examining sequential, scheduled region, CZone/delta correlation and reference prediction table prefetchers. Marius Grannæs, Magnus Jahre, Lasse Natvig |
ICCD | 2 |