Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jaime Roelandts

dblp:362/9014 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-8937-6888ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 60% Processor architecture and microarchitecture · 35% Energy-efficient computing · 5%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory access optimization
memory-level parallelism
1.422024
Scalar Vector Runahead · MICRO 2024
Decoupled Vector Runahead · MICRO 2023
Processor architecture and microarchitecture › microprocessor design › processor core design
in-order core
0.812024
Scalar Vector Runahead · MICRO 2024
Processor architecture and microarchitecture › latency hiding
runahead execution
0.812024
Scalar Vector Runahead · MICRO 2024
Memory systems › memory access patterns
indirect memory access
0.712023
Decoupled Vector Runahead · MICRO 2023
Memory systems › cache
prefetching
0.712023
Decoupled Vector Runahead · MICRO 2023
Energy-efficient computing › energy-efficient architecture
energy-efficient processor
0.212024
Scalar Vector Runahead · MICRO 2024
Memory systems › memory access patterns
irregular memory access
0.212024
Scalar Vector Runahead · MICRO 2024
Processor architecture and microarchitecture › out-of-order execution
out-of-order core
0.212023
Decoupled Vector Runahead · MICRO 2023

Methods — techniques the papers use, named apart from their topics

hardware prefetching · 0.7
YearPublicationVenuePosition
2026 ASI: A Unifying Metric to Assess and Improve Processor Microarchitecture Sustainability
abstract
Computing systems contribute significantly to global carbon emissions. Designing sustainable processor microarchitectures is particularly challenging due to the huge design space, the complex objective function, and the inherent data uncertainty.We introduce the Architectural Sustainability Indicator (ASI), a novel metric for evaluating the sustainability of processor microarchitectures. ASI enables classifying a design in different sustainability regions: strongly, weakly or unsustainable. ASI’s key strength lies in its ability to offer insights for how to transform unsustainable and weakly sustainable designs into strongly sustainable ones. Moreover, ASI can be used to guide automated design space exploration towards microarchitecture configurations that optimally trade off sustainability for performance. We analyze the effect of design decisions, such as the dominance of embodied versus operation footprint on ASI, and demonstrate that optimizing ASI differs from optimizing other metrics such as energy or energy-delay-area-product (EDAP).
Jaime Roelandts, Ajeya Naithani, Lieven Eeckhout
ISPASS1
2024 Scalar Vector Runahead
abstract
Modern graph and database processing typically takes place on high-end servers in data centers. However, with growing concerns of data privacy, trustworthiness, and all-time connectivity, there has been a shift toward increased analytics processing on edge devices such as mobile phones. In an ideal scenario, we would run these applications on the energy-efficient in-order cores available in these systems rather than the power-hungry out-of-order cores. However, these applications typically feature an extremely low computation-to-communication ratio and irregular memory accesses, meaning their performance is memory-bound, and out-of-order cores provide significant performance advantages over their in-order counterparts. Although prior work on Vector Runahead has substantially improved the performance of graph applications on very large out-of-order cores, it incurs high complexity and power consumption, so is unsuitable for energy-efficient in-order processors. Scalar Vector Runahead (SVR) extracts high memory-level parallelism on simple in-order cores by piggybacking on existing instructions executed on the processor leading to future irregular memory accesses. SVR executes multiple transient, independent, parallel instances of memory accesses and their chains initiated from different values of a predicted induction variable to move mutually independent memory accesses next to each other to hide dependent stalls. With a hardware overhead of only 2 KiB, SVR delivers$\mathbf{3.2}\times$higher performance than a baseline 3-wide in-order core, and$\mathbf{1.3}\times$higher performance than a full out-of-order core, while halving energy consumption. Increasing the overhead to 9 KiB to account for a larger register file, SVR can extend the speedup relative to an out-of-order core to$\mathbf{1.7}\times$.
Jaime Roelandts, Ajeya Naithani, Sam Ainsworth 0001, Timothy M. Jones 0001, Lieven Eeckhout
MICRO1
2023 Decoupled Vector Runahead
abstract
We present Decoupled Vector Runahead (DVR), an in-core prefetching technique, executing separately to the main application thread, that exploits massive amounts of memory-level parallelism to improve the performance of applications featuring indirect memory accesses. DVR dynamically infers loop bounds at run-time, recognizing striding loads, and vectorizing subsequent instructions that are part of an indirect chain. It proactively issues memory accesses for the resulting loads far into the future, even when the out-of-order core has not yet stalled, bringing their data into the L1 cache, and thus providing timely prefetches for the main thread. DVR can adjust the degree of vectorization at run-time, vectorize the same chain of indirect memory accesses across multiple invocations of an inner loop, and efficiently handle branch divergence along the vectorized chain. DVR runs as an on-demand, speculative, in-order, lightweight hardware subthread alongside the main thread within the core and incurs a minimal hardware overhead of only 1139 bytes. Relative to a large superscalar 5-wide out-of-order baseline and Vector Runahead — a recent microarchitectural technique to accelerate indirect memory accesses on out-of-order processors — DVR delivers 2.4 × and 2 × higher performance, respectively, for a set of graph analytics, database, and HPC workloads.
Ajeya Naithani, Jaime Roelandts, Sam Ainsworth 0001, Timothy M. Jones 0001, Lieven Eeckhout
MICRO2