Stephan Diestelhorst

dblp:72/8290 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-authorSoftware engineering, systems software and programming languages · 6Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Memory systems · 47% Processor architecture and microarchitecture · 21% Energy-efficient computing · 11%
Software engineering, system software, and programming languages
5 papers
Concurrent programming · 87% Programming languages and type systems · 8% Operating systems · 3%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 26 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › non-volatile memory
persistent memory
1.752020
Relaxed Persist Ordering Using Strand Persistency · ISCA 2020
Software Wear Management for Persistent Memories · FAST 2019
Persistency for synchronization-free regions · PLDI 2018
Memory systems
non-volatile memory
1.032019
Software Wear Management for Persistent Memories · FAST 2019
Persistency for synchronization-free regions · PLDI 2018
Delegated persist ordering · MICRO 2016
Memory systems › non-volatile memory › persistent memory
persistency model
0.722020
Relaxed Persist Ordering Using Strand Persistency · ISCA 2020
Language-level persistency · ISCA 2017
Concurrent programming
memory models
0.622018
Persistency for synchronization-free regions · PLDI 2018
Language-level persistency · ISCA 2017
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering
0.412020
Relaxed Persist Ordering Using Strand Persistency · ISCA 2020
Hardware security and side channels › microarchitectural side channel
branch predictor side channel
0.412019
BRB: Mitigating Branch Predictor Side-Channels · HPCA 2019
Hardware security and side channels
microarchitectural side channel
0.412019
BRB: Mitigating Branch Predictor Side-Channels · HPCA 2019
Processor architecture and microarchitecture
branch prediction
0.412019
BRB: Mitigating Branch Predictor Side-Channels · HPCA 2019
Concurrent programming › memory models
persistency models
0.312018
Persistency for synchronization-free regions · PLDI 2018
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.312018
SynchroTrace: Synchronization-Aware Architecture-Agnostic Traces for Lightweight Multicore Simulation of CMP and HPC Workloads · ACM Trans. Archit. Code Optim. 2018
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.312017
Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Energy-efficient computing
power modeling
0.312017
Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Energy-efficient computing › power modeling
run-time power estimation
0.312017
Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.212014
The TURBO Diaries: Application-controlled Frequency Scaling Explained · USENIX ATC 2014
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.112019
BRB: Mitigating Branch Predictor Side-Channels · HPCA 2019
Concurrent programming › transactional memory
hardware transactional memory
0.112010
Evaluation of AMD's advanced synchronization facility within a complete transactional memory stack · EuroSys 2010
Concurrent programming
transactional memory
0.112010
Evaluation of AMD's advanced synchronization facility within a complete transactional memory stack · EuroSys 2010
Parallel and multicore computing › transactional memory
hardware transactional memory
0.112010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010
Processor architecture and microarchitecture
instruction set architecture
0.112010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010
Processor architecture and microarchitecture › instruction set architecture
instruction set extension
0.112010
Evaluation of AMD's advanced synchronization facility within a complete transactional memory stack · EuroSys 2010
Processor architecture and microarchitecture
speculative execution
0.112010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010
Parallel and multicore computing
transactional memory
0.112010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010
Operating systems › resource management › process management
CPU scheduling
0.112014
The TURBO Diaries: Application-controlled Frequency Scaling Explained · USENIX ATC 2014
Compilers and program optimization › compiler construction
compiler support for transactional memory
0.012010
Evaluation of AMD's advanced synchronization facility within a complete transactional memory stack · EuroSys 2010
Parallel and multicore computing › concurrent data structures
lock-free data structures
0.012010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010
Parallel and multicore computing
synchronization
0.012010
ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory · MICRO 2010

Methods — techniques the papers use, named apart from their topics

performance simulation · 0.8memory consistency model · 0.6synchronization-aware tracing · 0.3statistical regression · 0.3multicollinearity reduction · 0.3simulation · 0.2software fallback · 0.2cycle-accurate simulation · 0.2compiler extension · 0.2out-of-order simulation · 0.1
YearPublicationVenuePosition
2020 Relaxed Persist Ordering Using Strand Persistency
abstract
Emerging persistent memory (PM) technologies promise the performance of DRAM with the durability of Flash. Several language-level persistency models have emerged recently to aid programming recoverable data structures in PM. Unfortunately, these persistency models are built upon hardware primitives that impose stricter ordering constraints on PM operations than the persistency models require. Alternative solutions use fixed and inflexible hardware logging techniques to relax ordering constraints on PM operations, but do not readily apply to general synchronization primitives employed by language-level persistency models. Instead, we propose StrandWeaver, a hardware strand persistency model, to minimally constrain ordering on PM operations. StrandWeaver manages PM order within a strand, a logically independent sequence of operations within a thread. PM operations that lie on separate strands are unordered and may drain concurrently to PM. StrandWeaver implements primitives under strand persistency to allow programmers to improve concurrency and relax ordering constraints on updates as they drain to PM. Furthermore, we design mechanisms that map persistency semantics in high-level language persistency models to the primitives implemented by StrandWeaver. We demonstrate that StrandWeaver can enable greater concurrency of PM operations than existing ISA-level ordering mechanisms, improving performance by up to $1.97 \times (1.45 \times avg.)$.
Vaibhav Gogte, Stephan Diestelhorst, Peter M. Chen, Satish Narayanasamy, Thomas F. Wenisch
ISCA3
2019 Software Wear Management for Persistent Memories
Vaibhav Gogte, Stephan Diestelhorst, Aasheesh Kolli, Peter M. Chen, Satish Narayanasamy, Thomas F. Wenisch
FAST3
2019 BRB: Mitigating Branch Predictor Side-Channels
abstract
Modern processors use branch prediction as an optimization to improve processor performance. Predictors have become larger and increasingly more sophisticated in order to achieve higher accuracies which are needed in high performance cores. However, branch prediction can also be a source of side channel exploits, as one context can deliberately change the branch predictor state and alter the instruction flow of another context. Current mitigation techniques either sacrifice performance for security, or fail to guarantee isolation when retaining the accuracy. Achieving both has proven to be challenging. In this work we address this by, (1) introducing the notions of steady-state and transient branch predictor accuracy, and (2) showing that current predictors increase their misprediction rate by as much as 90% on average when forced to flush branch prediction state to remain secure. To solve this, (3) we introduce the branch retention buffer, a novel mechanism that partitions only the most useful branch predictor components to isolate separate contexts. Our mechanism makes thread isolation practical, as it stops the predictor from executing cold with little if any added area and no warm-up overheads. At the same time our results show that, compared to the state-of-the-art, average misprediction rates are reduced by 15-20% without increasing area, leading to a 2% performance increase.
Ilias Vougioukas, Nikos Nikoleris, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett
HPCA4
2019 Persistent Atomics for Implementing Durable Lock-Free Data Structures for Non-Volatile Memory (Brief Announcement)
abstract
This brief announcement presents a persist ordering problem uncovered in implementing durable lock-free data structures for non-volatile memory, and proposes a hardware solution with persistent atomics in the Arm instruction set architecture.
Stephan Diestelhorst
SPAA2
2018 Hardware-Validated CPU Performance and Energy Modelling
abstract
Full-system simulation frameworks such as gem5 are used extensively to evaluate research ideas and for design-space exploration. Moreover, energy-efficiency has become the key design constraint in recent years and many works use a separate power modelling framework to evaluate energy consumption. While such tools are convenient and flexible, they are known to contain sources of error which are often not fully understood and potentially impact the conclusions drawn from investigations. This work enables accurate, hardware-validated performance, power, and energy modelling of CPUs by first presenting a methodology to evaluate and identify sources of error in CPU performance models, and secondly developing empirical power models optimised for use with such performance models. Hierarchical clustering, correlation analysis, and regression techniques are used to identify sources of error without requiring detailed CPU specifications and enable existing models to be improved, new models to be developed, validation of simulator changes, and testing of model suitability for specific use-cases. Furthermore, the GemStone open-source software tool is presented, which automates the process of characterising hardware platforms, identifying sources of error in gem5 models, applying power analysis, and quantifying the effect of errors on the performance, power, and energy estimations. In addition, the mean percentage error in execution time was found to swing from -51% to +10% between two versions of the same gem5 model, underlining the need for an automated tool to validate models against reference hardware, ensuring accuracy and consistency.
Matthew J. Walker, Sascha Bischoff, Stephan Diestelhorst, Geoff V. Merrett, Bashir M. Al-Hashimi
ISPASS3
2018 Persistency for synchronization-free regions
abstract
Nascent persistent memory (PM) technologies promise the performance of DRAM with the durability of disk, but how best to integrate them into programming systems remains an open question. Recent work extends language memory models with a persistency model prescribing semantics for updates to PM. These semantics enable programmers to design data structures in PM that are accessed like memory and yet are recoverable upon crash or failure. Alas, we find the semantics and performance of existing approaches unsatisfying. Existing approaches require high-overhead mechanisms, are restricted to certain synchronization constructs, provide incomplete semantics, and/or may recover to state that cannot arise in fault-free execution.
Vaibhav Gogte, Stephan Diestelhorst, Satish Narayanasamy, Peter M. Chen, Thomas F. Wenisch
PLDI2
2018 SynchroTrace: Synchronization-Aware Architecture-Agnostic Traces for Lightweight Multicore Simulation of CMP and HPC Workloads
abstract
Trace-driven simulation of chip multiprocessor (CMP) systems offers many advantages over execution-driven simulation, such as reducing simulation time and complexity, allowing portability, and scalability. However, trace-based simulation approaches have difficulty capturing and accurately replaying multithreaded traces due to the inherent nondeterminism in the execution of multithreaded programs. In this work, we present SynchroTrace, a scalable, flexible, and accurate trace-based multithreaded simulation methodology. By recording synchronization events relevant to modern threading libraries (e.g., Pthreads and OpenMP) and dependencies in the traces, independent of the host architecture, the methodology is able to accurately model the nondeterminism of multithreaded programs for different hardware platforms and threading paradigms. Through capturing high-level instruction categories, the SynchroTrace average CPI trace Replay timing model offers fast and accurate simulation of many-core in-order CMPs. We perform two case studies to validate the SynchroTrace simulation flow against the gem5 full-system simulator: (1) a constraint-based design space exploration with traditional CMP benchmarks and (2) a thread-scalability study with HPC-representative applications. The results from these case studies show that (1) our trace-based approach with trace filtering has a peak speedup of up to 18.7× over simulation in gem5 full-system with an average of 9.6× speedup, (2) SynchroTrace maintains the thread-scaling accuracy of gem5 and can efficiently scale up to 64 threads, and (3) SynchroTrace can trace in one platform and model any platform in early stages of design.
Karthik Sangaiah, Michael Lui, Radhika Jagtap, Stephan Diestelhorst, Siddharth Nilakantan, Ankit More, Baris Taskin, Mark Hempstead
ACM Trans. Archit. Code Optim.4
2017 Language-level persistency
abstract
The commercial release of byte-addressable persistent memories, such as Intel/Micron 3D XPoint memory, is imminent. Ongoing research has sought mechanisms to allow programmers to implement recoverable data structures in these new main memories. Ensuring recoverability requires programmer control of the order of persistent stores; recent work proposes persistency models as an extension to memory consistency to specify such ordering. Prior work has considered persistency models at the abstraction of the instruction set architecture. Instead, we argue for extending the language-level memory model to provide guarantees on the order of persistent writes.
Aasheesh Kolli, Vaibhav Gogte, Ali G. Saidi, Stephan Diestelhorst, Peter M. Chen, Satish Narayanasamy, Thomas F. Wenisch
ISCA4
2017 dist-gem5: Distributed simulation of computer clusters
abstract
When analyzing a distributed computer system, we often observe that the complex interplay among processor, node, and network sub-systems can profoundly affect the performance and power efficiency of the distributed computer system. Therefore, to effectively cross-optimize hardware and software components of a distributed computer system, we need a full-system simulation infrastructure that can precisely capture the complex interplay. Responding to the aforementioned need, we present dist-gem5, a flexible, detailed, and open-source full-system simulation infrastructure that can model and simulate a distributed computer system using multiple simulation hosts. Then we validate dist-gem5 against a physical cluster and show that the latency and bandwidth of the simulated network sub-system are within 18% of the physical one. Compared with the single threaded and parallel versions of gem5, dist-gem5 speeds up the simulation of a 63-node computer cluster by 83.1x and 12.8x, respectively.
Mohammad Alian, Umur Darbaz, Gábor Dózsa, Stephan Diestelhorst, Daehoon Kim 0001, Nam Sung Kim
ISPASS4
2017 Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUs
abstract
Modern mobile and embedded devices are required to be increasingly energy-efficient while running more sophisticated tasks, causing the CPU design to become more complex and employ more energy-saving techniques. This has created a greater need for fast and accurate power estimation frameworks for both run-time CPU energy management and design-space exploration. We present a statistically rigorous and novel methodology for building accurate run-time power models using performance monitoring counters (PMCs) for mobile and embedded devices, and demonstrate how our models make more efficient use of limited training data and better adapt to unseen scenarios by uniquely considering stability. Our robust model formulation reduces multicollinearity, allows separation of static and dynamic power, and allows a 100× reduction in experiment time while sacrificing only 0.6% accuracy. We present a statistically detailed evaluation of our model, highlighting and addressing the problem of heteroscedasticity in power modeling. We present software implementing our methodology and build power models for ARM Cortex-A7 and Cortex-A15 CPUs, with 3.8% and 2.8% average error, respectively. We model the behavior of the nonideal CPU voltage regulator under dynamic CPU activity to improve modeling accuracy by up to 5.5% in situations where the voltage cannot be measured. To address the lack of research utilizing PMC data from real mobile devices, we also present our data acquisition method and experimental platform software. We support this paper with online resources including software tools, documentation, raw data and further results.
Matthew J. Walker, Stephan Diestelhorst, Andreas Hansson 0001, Anup Das 0001, Sheng Yang 0003, Bashir M. Al-Hashimi, Geoff V. Merrett
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Nucleus: Finding the Sharing Limit of Heterogeneous Cores
abstract
Heterogeneous multi-processors are designed to bridge the gap between performance and energy efficiency in modern embedded systems. This is achieved by pairing Out-of-Order (OoO) cores, yielding performance through aggressive speculation and latency masking, with In-Order (InO) cores, that preserve energy through simpler design. By leveraging migrations between them, workloads can therefore select the best setting for any given energy/delay envelope. However, migrations introduce execution overheads that can hurt performance if they happen too frequently. Finding the optimal migration frequency is critical to maximize energy savings while maintaining acceptable performance. We develop a simulation methodology that can 1) isolate the hardware effects of migrations from the software, 2) directly compare the performance of different core types, 3) quantify the performance degradation and 4) calculate the cost of migrations for each case. To showcase our methodology we run mibench, a microbenchmark suite, and show that migrations can happen as fast as every 100k instructions with little performance loss. We also show that, contrary to numerous recent studies, hypothetical designs do not need to share all of their internal components to be able to migrate at that frequency. Instead, we propose a feasible system that shares level 2 caches and a translation lookaside buffer that matches performance and efficiency. Our results show that there are phases comprising up to 10% that a migration to the OoO core leads to performance benefits without any additional energy cost when running on the InO core, and up to 6% of phases where a migration to the InO core can save energy without affecting performance. When considering a policy that focuses on improving the energy-delay product, results show that on average 66% of the phases can be migrated to deliver equal or better system operation without having to aggressively share the entire memory system or to revert to migration periods finer than 100k instructions.
Ilias Vougioukas, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett
ACM Trans. Embed. Comput. Syst.3
2016 Elastic traces for fast and accurate system performance exploration
abstract
As computer systems become increasingly complex, the need for fast and accurate simulation tools increases. Accurate but slow processor core models are often substituted with simple trace players to achieve faster memory-system simulation. However, existing trace-driven simulation techniques are limited in their applicability and availability. In this work, we capture elastic traces containing out-of-order core dependencies and effects of speculative execution, which overcome limitations of existing work. Additionally, we make our capture and replay modelling available in the gem5 simulator. Our trace-driven CPU achieves a speed-up of 6-8x compared to the reference core and predicts the performance with less than 1% error on average when the memory-system is changed.
Radhika Jagtap, Stephan Diestelhorst, Andreas Hansson 0001
ISPASS2
2016 Delegated persist ordering
abstract
Systems featuring a load-store interface to persistent memory (PM) are expected soon, making in-memory persistent data structures feasible. Ensuring persistent data structure recoverability requires constraints on the order PM writes become persistent. But, current memory systems reorder writes, providing no such guarantees. To complement their upcoming 3D XPoint memory, Intel has announced new instructions to enable programmer control of data persistence. We describe the semantics implied by these instructions, an ordering model we call synchronous ordering. Synchronous ordering (SO) enforces order by stalling execution when PM write ordering is required, exposing PM write latency on the execution critical path. It incurs an average slowdown of 7.21x over volatile execution without ordering in PM-write-intensive benchmarks. SO tightly couples enforcing order and flushing writes to PM, but this tight coupling is unneeded in many recoverable software systems. Instead, we propose delegated ordering, wherein ordering requirements are communicated explicitly to the PM controller, fully decoupling PM write ordering from volatile execution and cache management. We demonstrate that delegated ordering can bring performance within 1.93x of volatile execution, improving over SO by 3.73x.
Aasheesh Kolli, Jeff Rosen, Stephan Diestelhorst, Ali G. Saidi, Steven Pelley, Sihang Liu 0001, Peter M. Chen, Thomas F. Wenisch
MICRO3
2014 The TURBO Diaries: Application-controlled Frequency Scaling Explained
Jons-Tobias Wamhoff, Stephan Diestelhorst, Christof Fetzer, Patrick Marlier, Pascal Felber, David Dice
USENIX ATC2
2013 Brief announcement: between all and nothing - versatile aborts in hardware transactional memory
abstract
Hardware Transactional Memory (HTM) implementations are becoming available in commercial, off-the-shelf components. While generally comparable, some implementations deviate from the strict all-or-nothing property of pure Transactional Memory. We analyse these deviations and find that with small modifications, they can be used to accelerate and simplify both transactional and non-transactional programming constructs. At the heart of our extensions we enable access to the transaction's full register state in the abort handler in an existing HTM without extending the architectural register state. Access to the full register state enables applications in both transactional and non-transactional parallel programming: hybrid transactional memory; transactional escape actions; transactional suspend/resume; and alert-on-update.
Stephan Diestelhorst, Martin Nowack, Michael F. Spear, Christof Fetzer
SPAA1
2012 Delegation and nesting in best-effort hardware transactional memory
abstract
The guiding design principle behind best-effort hardware transactional memory (BEHTM) is simplicity of implementation and verification. Only minimal modifications to the base processor architecture are allowed, thereby reducing the burden of verification and long-term support. In exchange, the hardware can support only relatively simple multiword atomic operations, and must fall back to a software run-time for any operation that exceeds the abilities of the hardware.
Yujie Liu 0003, Stephan Diestelhorst, Michael F. Spear
SPAA2
2010 Evaluation of AMD's advanced synchronization facility within a complete transactional memory stack
abstract
AMD's Advanced Synchronization Facility (ASF) is an x86 instruction set extension proposal intended to simplify and speed up the synchronization of concurrent programs. In this paper, we report our experiences using ASF for implementing transactional memory. We have extended a C/C++ compiler to support language-level transactions and generate code that takes advantage of ASF. We use a software fallback mechanism for transactions that cannot be committed within ASF (e.g., because of hardware capacity limitations). Our evaluation uses a cycle-accurate x86 simulator that we have extended with ASF support. Building a complete ASF-based software stack allows us to evaluate the performance gains that a user-level program can obtain from ASF. Our measurements on a wide range of benchmarks indicate that the overheads traditionally associated with software transactional memories can be significantly reduced with the help of ASF.
David Christie, Jae-Woong Chung, Stephan Diestelhorst, Michael Hohmuth, Martin Pohlack, Christof Fetzer, Martin Nowack, Torvald Riegel, Pascal Felber, Patrick Marlier, Etienne Rivière
EuroSys3
2010 ASF: AMD64 Extension for Lock-Free Data Structures and Transactional Memory
abstract
Advanced Synchronization Facility (ASF) is an AMD64 hardware extension for lock-free data structures and transactional memory. It provides a speculative region that atomically executes speculative accesses in the region. Five new instructions are added to demarcate the region, use speculative accesses selectively, and control the speculative hardware context. Programmers can use speculative regions to build flexible multi-word atomic primitives with no additional software support by relying on the minimum guarantee of available ASF hardware resources for lock-free programming. Transactional programs with high-level TM language constructs can either be compiled directly to the ASF code or be linked to software TM systems that use ASF to accelerate transactional execution. In this paper we develop an out-of-order hardware design to implement ASF on a future AMD processor and evaluate it with an in-house simulator. The experimental results show that the combined use of the L1 cache and the LS unit is very helpful for the performance robustness of ASF-based lock free data structures, and that the selective use of speculative accesses enables transactional programs to scale with limited ASF hardware resources.
Jae-Woong Chung, Luke Yen, Stephan Diestelhorst, Martin Pohlack, Michael Hohmuth, David Christie, Dan Grossman
MICRO3