Sam Van den Steen

dblp:162/1150 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0003-3630-2214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 68% Performance modeling and evaluation · 20% Energy-efficient computing · 10%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › branch prediction
branch misprediction recovery
0.512021
Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions · MICRO 2021
Processor architecture and microarchitecture
branch prediction
0.512021
Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions · MICRO 2021
Processor architecture and microarchitecture
speculative execution
0.512021
Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions · MICRO 2021
Performance modeling and evaluation
analytical modeling
0.212016
Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016
Energy-efficient computing
power modeling
0.212016
Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016
Performance modeling and evaluation
processor performance modeling
0.212016
Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016
Processor architecture and microarchitecture
reorder buffer
0.112021
Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions · MICRO 2021
Electronic design automation
design space exploration
0.112016
Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016
Processor architecture and microarchitecture › superscalar processor
superscalar out-of-order processor
0.112016
Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics · IEEE Trans. Computers 2016

Methods — techniques the papers use, named apart from their topics

selective flushing · 0.5cycle-level simulation · 0.2analytical modeling · 0.2
YearPublicationVenuePosition
2023 Simulating Wrong-Path Instructions in Decoupled Functional-First Simulation
abstract
Wrong-path speculative execution on an out-oforder processor core has no impact on an application’s functionality and correctness, but it can impact performance by changing the state of caches and predictors. Not modeling wrong-path execution in performance simulation leads to performance projection errors up to 22% for our setup. However, wrong-path execution is challenging to model for common functional-first simulators, because the functional simulator is not aware of branch predictor misses and only provides correct-path instructions. We propose and evaluate multiple wrong-path modeling techniques for functional-first simulators, each with a different accuracy versus simulation speed balance. The novel instruction reconstruction with convergence exploitation technique proves to be the best balanced technique, with about $3 \times$ lower error than no wrong path modeling and about 2 to $3 \times$ faster simulation than full wrong path emulation.
Stijn Eyerman, Sam Van den Steen, Wim Heirman, Ibrahim Hur
ISPASS2
2021 Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions
abstract
Conventionally, branch mispredictions are resolved by flushing wrongly speculated instructions from the reorder buffer and refetching instructions along the correct path. However, a large part of the misspeculated instructions could have reconverged with the correct path and executed correctly. Yet, they are flushed to ensure in-order commit. This inefficiency has been recognized in prior work, which proposes either complex additions to a core to reuse the correctly executed instructions, or less intrusive solutions that only reuse part of the converged instructions.
Stijn Eyerman, Wim Heirman, Sam Van den Steen, Ibrahim Hur
MICRO3
2019 RPPM: Rapid Performance Prediction of Multithreaded Workloads on Multicore Processors
abstract
Analytical performance modeling is a useful complement to detailed cycle-level simulation to quickly explore the design space in an early design stage. Mechanistic analytical modeling is particularly interesting as it provides deep insight and does not require expensive offline profiling as empirical modeling. Previous work in mechanistic analytical modeling, unfortunately, is limited to single-threaded applications running on single-core processors. This work proposes RPPM, a mechanistic analytical performance model for multi-threaded applications on multicore hardware. RPPM collects microarchitecture-independent characteristics of a multi-threaded workload to predict performance on a previously unseen multicore architecture. The profile needs to be collected only once to predict a range of processor architectures. We evaluate RPPM's accuracy against simulation and report a performance prediction error of 11.2% on average (23% max). We demonstrate RPPM's usefulness for conducting design space exploration experiments as well as for analyzing parallel application performance.
Sander De Pestel, Sam Van den Steen, Shoaib Akram 0001, Lieven Eeckhout
ISPASS2
2016 Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics
abstract
Optimizing processors for (a) specific application(s) can substantially improve energy-efficiency. With the end of Dennard scaling, and the corresponding reduction in energy-efficiency gains from technology scaling, such approaches may become increasingly important. However, designing application-specific processors requires fast design space exploration tools to optimize for the targeted application(s). Analytical models can be a good fit for such design space exploration as they provide fast performance and power estimates and insight into the interaction between an application's characteristics and the micro-architecture of a processor. Unfortunately, prior analytical models for superscalar out-of-order processors require micro-architecture dependent inputs, such as cache miss rates, branch miss rates and memory-level parallelism. This requires profiling the applications for each cache and branch predictor configuration of interest, which is far more time-consuming than evaluating the analytical performance models. In this work we present amicro-architecture independentprofiler and associated analytical models that allow us to produce performanceandpower estimates across a large superscalar out-of-order processor design space almost instantaneously. We show that using a micro-architecture independent profile leads to a speedup of 300$\times$compared to detailed simulation for our evaluated design space. Over a large design space, the model has a 9.3 percent average error for performance and a 4.3 percent average error for power, compared to detailed cycle-level simulation. The model is able to accurately determine the optimal processor configuration for different applications under power or performance constraints, and provides insight into performance through cycle stacks.
Sam Van den Steen, Stijn Eyerman, Sander De Pestel, Moncef Mechri, Trevor E. Carlson, David Black-Schaffer, Erik Hagersten, Lieven Eeckhout
IEEE Trans. Computers1
2015 Micro-architecture independent analytical processor performance and power modeling
abstract
Optimizing processors for specific application(s) can substantially improve energy-efficiency. With the end of Dennard scaling, and the corresponding reduction in energyefficiency gains from technology scaling, such approaches may become increasingly important. However, designing applicationspecific processors require fast design space exploration tools to optimize for the targeted application(s). Analytical models can be a good fit for such design space exploration as they provide fast performance estimations and insight into the interaction between an application's characteristics and the micro-architecture of a processor. Unfortunately, current analytical models require some microarchitecture dependent inputs, such as cache miss rates, branch miss rates and memory-level parallelism. This requires profiling the applications for each cache and branch predictor configuration, which is far more time-consuming than evaluating the actual performance models. In this work we present a micro-architecture independent profiler and associated analytical models that allow us to produce performance and power estimates across a large design space almost instantaneously. We show that using a micro-architecture independent profile leads to a speedup of 25× for our evaluated design space, compared to an analytical model that uses micro-architecture dependent profiles. Over a large design space, the model has a 13% error for performance and a 7% error for power, compared to cycle-level simulation. The model is able to accurately determine the optimal processor configuration for different applications under power or performance constraints, and it can provide insight into performance through cycle stacks.
Sam Van den Steen, Sander De Pestel, Moncef Mechri, Stijn Eyerman, Trevor E. Carlson, David Black-Schaffer, Erik Hagersten, Lieven Eeckhout
ISPASS1