Rik van Riel

dblp:55/1143 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0009-8010-3635ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 50% Hardware reliability and fault tolerance · 33% Performance modeling and evaluation · 10%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance › soft errors
silent data corruption
0.912025
Hardware Sentinel: Protecting Software Applications from Hardware Silent Data Corruptions · ASPLOS (2) 2025
Memory systems › memory management
fragmentation
0.712023
Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters · ISCA 2023
Memory systems › memory management
virtual memory
0.712023
Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters · ISCA 2023
Performance modeling and evaluation
workload characterization
0.312025
Hardware Sentinel: Protecting Software Applications from Hardware Silent Data Corruptions · ASPLOS (2) 2025
Operating systems › resource management
memory management
0.212023
Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters · ISCA 2023
Operating systems › resource management › memory management
page migration
0.212023
Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters · ISCA 2023
Cloud and datacenter computing › resource management
datacenter memory management
0.212023
Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters · ISCA 2023

Methods — techniques the papers use, named apart from their topics

page migration · 1.3hardware-software co-design · 1.3TLB shootdown · 1.3software failure indicator analysis · 0.9large-scale fleet deployment · 0.9
YearPublicationVenuePosition
2025 Hardware Sentinel: Protecting Software Applications from Hardware Silent Data Corruptions
abstract
Silent Data Corruptions (SDCs) pose a significant challenge in large-scale infrastructures, affecting data center applications unpredictably and reducing service reliability. Primarily caused by silicon defects, traditional hardware testing methods are insufficient to prevent SDC propagation. SDCs are influenced by various factors, including data randomization, workload characteristics, environmental conditions, and aging, necessitating top-down approaches from the application layer. In this paper, we introduce Hardware Sentinel, a novel framework that detects SDCs through typical software failure indicators such as segmentation faults, core dumps, application crashes, and logs. We have validated our framework in a large-scale data center fleet, across diverse application, kernel, and hardware configurations, achieving a high success rate of SDC detection. Hardware Sentinel has uncovered novel instances of SDCs, surpassing the detection capabilities of published testing techniques. Our analysis of over 6 years' worth of application and system failure data within a large-scale infrastructure has successfully identified hundreds of defective CPUs that triggered SDCs. Notably, the Hardware Sentinel flow increases effective coverage over existing hardware-testing methods like Fleetscanner (out-of-production testing) by 1.74x and Ripple (in-production testing) by 1.92x. We share the top kernel exceptions with the highest correlation to silent data corruption failures. We present results spanning 7 CPU generations from multiple semiconductor manufacturers, 13 large-scale workloads, and 27 data center regions, providing insights into the trade-offs involved in detection and fleet deployment.
Rhea Dutta, Harish Dattatraya Dixit, Rik van Riel, Gautham Vunnam, Sriram Sankar
ASPLOS (2)3
2023 Contiguitas: The Pursuit of Physical Memory Contiguity in Datacenters
abstract
The unabating growth of the memory needs of emerging datacenter applications has exacerbated the scalability bottleneck of virtual memory. However, reducing the excessive overhead of address translation will remain onerous until the physical memory contiguity predicament gets resolved. To address this problem, this paper presents Contiguitas, a novel redesign of memory management in the operating system and hardware that provides ample physical memory contiguity. We identify that the primary cause of memory fragmentation in Meta's datacenters is unmovable allocations scattered across the address space that impede large contiguity from being formed. To provide ample physical memory contiguity by design, Contiguitas first separates regular movable allocations from unmovable ones by placing them into two different continuous regions in physical memory and dynamically adjusts the boundary of the two regions based on memory demand. Drastically reducing unmovable allocations is challenging because the majority of unmovable pages cannot be moved with software alone given that access to the page cannot be blocked for a migration to take place. Furthermore, page migration is expensive as it requires a long downtime to (a) perform TLB shootdowns that scale poorly with the number of victim TLBs, and (b) copy the page. To this end, Contiguitas eliminates the primary source of unmovable allocations by introducing hardware extensions in the last-level cache to enable the transparent and efficient migration of unmovable pages even while the pages remain in use.
Kaiyang Zhao 0002, Ziqi Wang 0007, Dan Schatzberg, Leon Yang, Antonis Manousis, Johannes Weiner, Rik van Riel, Bikash Sharma, Chunqiang Tang, Dimitrios Skarlatos 0002
ISCA8