Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marc de Kruijf

dblp:54/8277 · also Marc Asher de Kruijf · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-authorComputer networks · 2 · 1 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Cloud and datacenter computing · 53% Hardware reliability and fault tolerance · 21% Processor architecture and microarchitecture · 16%
Software engineering, system software, and programming languages
5 papers
Concurrent programming · 26% Debugging and program repair · 26% Software maintenance and evolution · 15%
Computer networks
1 paper
Datacenter networks · 100%

Topics — the 20 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
concurrency bugs
0.422015
Fixing, preventing, and recovering from concurrency bugs · Sci. China Inf. Sci. 2015
ConAir: featherweight concurrency bug recovery via single-threaded idempotent execution · ASPLOS 2013
Cloud and datacenter computing
datacenter network
0.412019
Snap: a microkernel approach to host networking · SOSP 2019
Cloud and datacenter computing › virtualization › network virtualization
network function virtualization
0.412019
Snap: a microkernel approach to host networking · SOSP 2019
Cloud and datacenter computing
userspace networking
0.412019
Snap: a microkernel approach to host networking · SOSP 2019
Cloud and datacenter computing › virtualization
network virtualization
0.312018
Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization · NSDI 2018
Processor architecture and microarchitecture
speculative execution
0.322012
iGPU: Exception support and speculative execution on GPUs · ISCA 2012
Idempotent processor architecture · MICRO 2011
Software maintenance and evolution › software maintenance
bug fixing
0.212015
Fixing, preventing, and recovering from concurrency bugs · Sci. China Inf. Sci. 2015
Debugging and program repair › automated program repair
concurrency bug fixing
0.212013
ConAir: featherweight concurrency bug recovery via single-threaded idempotent execution · ASPLOS 2013
Compilers and program optimization
compiler construction
0.112012
Static analysis and compiler design for idempotent processing · PLDI 2012
Program analysis
static analysis
0.112012
Static analysis and compiler design for idempotent processing · PLDI 2012
GPUs and heterogeneous computing
GPU architecture
0.112012
iGPU: Exception support and speculative execution on GPUs · ISCA 2012
Hardware reliability and fault tolerance › redundancy
dual modular redundancy
0.112011
Sampling + DMR: practical and low-overhead permanent fault detection · ISCA 2011
Distributed systems
fault tolerance
0.112011
Sampling + DMR: practical and low-overhead permanent fault detection · ISCA 2011
Hardware reliability and fault tolerance
permanent fault
0.112011
Sampling + DMR: practical and low-overhead permanent fault detection · ISCA 2011
Operating systems › kernel › kernel design › microkernel
microkernel design
0.112019
Snap: a microkernel approach to host networking · SOSP 2019
Hardware reliability and fault tolerance
fault-tolerant architecture
0.112010
Relax: an architectural framework for software recovery of hardware faults · ISCA 2010
Hardware reliability and fault tolerance
error recovery
0.122012
Static analysis and compiler design for idempotent processing · PLDI 2012
Idempotent processor architecture · MICRO 2011
Runtime systems and virtual machines
dynamic compilation
0.012012
iGPU: Exception support and speculative execution on GPUs · ISCA 2012
Processor architecture and microarchitecture
chip multiprocessor
0.012011
Sampling + DMR: practical and low-overhead permanent fault detection · ISCA 2011
Hardware reliability and fault tolerance › soft errors
transient fault
0.012010
Relax: an architectural framework for software recovery of hardware faults · ISCA 2010

Methods — techniques the papers use, named apart from their topics

kernel bypass · 0.8RDMA · 0.8idempotent code regions · 0.3checkpointing · 0.3single-threaded idempotent execution · 0.2sampling · 0.1re-execution · 0.1dual modular redundancy · 0.1compiler-constructed idempotent regions · 0.1
YearPublicationVenuePosition
2022 Understanding host interconnect congestion
abstract
We present evidence and characterization of host congestion in production clusters: adoption of high-bandwidth access links leading to emergence of bottlenecks within the host interconnect (NIC-to-CPU data path). We demonstrate that contention on existing IO memory management units and/or the memory subsystem can significantly reduce the available NIC-to-CPU bandwidth, resulting in hundreds of microseconds of queueing delays and eventual packet drops at hosts (even when running a state-of-the-art congestion control protocol that accounts for CPU-induced host congestion). We also discuss implications of host interconnect congestion to design of future host architecture, network stacks and network protocols.
Saksham Agarwal, Rachit Agarwal 0001, Behnam Montazeri, Masoud Moshref, Khaled Elmeleegy, Luigi Rizzo, Marc de Kruijf, Gautam Kumar 0001, Sylvia Ratnasamy, David E. Culler, Amin Vahdat
HotNets7
2019 Snap: a microkernel approach to host networking
abstract
This paper presents our design and experience with a microkernel-inspired approach to host networking called Snap. Snap is a userspace networking system that supports Google's rapidly evolving needs with flexible modules that implement a range of network functions, including edge packet switching, virtualization for our cloud platform, traffic shaping policy enforcement, and a high-performance reliable messaging and RDMA-like service. Snap has been running in production for over three years, supporting the extensible communication needs of several large and critical systems.
Michael Marty, Marc de Kruijf, Jacob Adriaens, Christopher Alfeld, Sean Bauer, Carlo Contavalli, Michael Dalton, Nandita Dukkipati, William C. Evans, Steve D. Gribble, Nicholas Kidd, Roman Kononov, Gautam Kumar 0001, Carl Mauer, Emily Musick, Lena E. Olson, Erik Rubow, Michael Ryan, Kevin Springborn, Valas Valancius, Amin Vahdat
SOSP2
2018 Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization
Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshuman Gupta, Brian Fahs, Dima Rubinstein, Enrique Cauich Zermeno, Erik Rubow, James Alexander Docauer, Jesse Alpert, Jing Ai, Jon Olson, Kevin DeCabooter, Marc de Kruijf, Nan Hua, Nathan Lewis, Nikhil Kasinadhuni, Riccardo Crepaldi, Srinivas Krishnan, Subbaiah Venkata, Yossi Richter, Uday Naik, Amin Vahdat
NSDI15
2015 Fixing, preventing, and recovering from concurrency bugs
Dongdong Deng, Guoliang Jin, Marc de Kruijf, Ben Liblit, Shan Lu 0001, Shanxiang Qi, Jinglei Ren, Karthikeyan Sankaralingam, Linhai Song, Yongwei Wu 0001, Wei Zhang 0022
Sci. China Inf. Sci.3
2013 ConAir: featherweight concurrency bug recovery via single-threaded idempotent execution
abstract
Many concurrency bugs are hidden in deployed software and cause severe failures for end-users. When they finally manifest and become known by developers, they are difficult to fix correctly. To support end-users, we need techniques that help software survive hidden concurrency bugs during production runs. To help developers, we need techniques that fix exposed concurrency bugs.
Wei Zhang 0022, Marc de Kruijf, Shan Lu 0001, Karthikeyan Sankaralingam
ASPLOS2
2013 Idempotent code generation: Implementation, analysis, and evaluation
abstract
Leveraging idempotence for efficient recovery is of emerging interest in compiler design. In particular, identifying semantically idempotent code and then compiling such code to preserve the semantic idempotence property enables recovery with substantially lower overheads than competing software techniques. However, the efficacy of this technique depends on application-, architecture-, and compiler-specific factors that are not well understood. In this paper, we develop algorithms for the code generation of idempotent code regions and evaluate these algorithms considering how they are impacted by these factors. Without optimizing for these factors, we find that typical performance overheads fall in the range of roughly 10-15%. However, manipulating application idempotent region size typically improves the run-time performance of compiled code by 2-10%, differences in the architecture instruction set affect performance by up to 15%, and knowing in the compiler whether control flow side-effects can or cannot occur can impact performance by up to 10%. Overall, we find that, with small idempotent region and careful architecture- and application-specific tuning, it is possible to bring compiler performance overheads consistently down into the single-digit percentage range. The absolute best performance occurs when constructing the largest possible idempotent regions; to this end, however, better compiler support is needed. In the interest of spurring development in this area, we open-source our LLVM compiler implementation and make it available as a research tool.
Marc de Kruijf, Karthikeyan Sankaralingam
CGO1
2012 Mechanisms and Evaluation of Cross-Layer Fault-Tolerance for Supercomputing
abstract
Reliability is emerging as an important constraint for future microprocessors. Cooperative hardware and software approaches for error tolerance can solve this hardware reliability challenge. Cross-layer fault tolerance frameworks expose hardware failures to upper-layers, like the compiler, to help correct faults. Such cooperative approaches require less hardware complexity than masking all faults at the hardware level and are generally more energy efficient. This paper provides a detailed design and an implementation study of cross-layer fault tolerance for supercomputing. Since supercomputers necessarily involve large component counts, they have more frequent failures than consumer electronics and small systems. Conventionally, these systems use redundancy and check pointing to achieve reliable computing. However, redundancy increases acquisition as well as recurring energy costs. This paper describes a simple language-level mechanism coupled with complementary compilation and lightweight hardware error detection that provides efficient reliability and cross-layer fault-tolerance for supercomputers. Our evaluation focuses on strong scaling problems for which we can trade computing power for redundancy. Our results show a range of 1.07× to 2.5× speedup when employing cross-layer error-tolerance compared to conventional full dual modular redundancy (DMR) to contain all errors within hardware. Further, we demonstrate the approach can sustain 7% to 50% lower energy. The most important result of this work is qualitative: we can use a simplified hardware design with relaxed architectural correctness guarantees.
Chen-Han Ho, Marc de Kruijf, Karthikeyan Sankaralingam, Barry Rountree, Martin Schulz 0001, Bronis R. de Supinski
ICPP2
2012 iGPU: Exception support and speculative execution on GPUs
abstract
Since the introduction of fully programmable vertex shader hardware, GPU computing has made tremendous advances. Exception support and speculative execution are the next steps to expand the scope and improve the usability of GPUs. However, traditional mechanisms to support exceptions and speculative execution are highly intrusive to GPU hardware design. This paper builds on two related insights to provide a unified lightweight mechanism for supporting exceptions and speculation on GPUs. First, we observe that GPU programs can be broken into code regions that contain little or no live register state at their entry point. We then also recognize that it is simple to generate these regions in such a way that they are idempotent, allowing their entry points to function as program recovery points and enabling support for exception handling, fast context switches, and speculation, all with very low overhead. We call the architecture of GPUs executing these idempotent regions the iGPU architecture. The hardware extensions required are minimal and the construction of idempotent code regions is fully transparent under the typical dynamic compilation framework of GPUs. We demonstrate how iGPU exception support enables virtual memory paging with very low overhead (1% to 4%), and how speculation support enables circuit-speculation techniques that can provide over 25% reduction in energy.
Jai Menon 0003, Marc de Kruijf, Karthikeyan Sankaralingam
ISCA2
2012 Static analysis and compiler design for idempotent processing
abstract
Recovery functionality has many applications in computing systems, from speculation recovery in modern microprocessors to fault recovery in high-reliability systems. Modern systems commonly recover using checkpoints. However, checkpoints introduce overheads, add complexity, and often save more state than necessary.
Marc de Kruijf, Karthikeyan Sankaralingam, Somesh Jha
PLDI1
2011 Sampling + DMR: practical and low-overhead permanent fault detection
abstract
With technology scaling, manufacture-time and in-field permanent faults are becoming a fundamental problem. Multi-core architectures with spares can tolerate them by detecting and isolating faulty cores, but the required fault detection coverage becomes effectively 100% as the number of permanent faults increases. Dual-modular redundancy(DMR) can provide 100% coverage without assuming device-level fault models, but its overhead is excessive.
Shuou Nomura, Matthew D. Sinclair, Chen-Han Ho, Venkatraman Govindaraju, Marc de Kruijf, Karthikeyan Sankaralingam
ISCA5
2011 Idempotent processor architecture
abstract
Improving architectural energy efficiency is important to address diminishing energy efficiency gains from technology scaling. At the same time, limiting hardware complexity is also important. This paper presents a new processor architecture, the idempotent processor architecture, that advances both of these directions by presenting a new execution paradigm that allows speculative execution without the need for hardware checkpoints to recover from mis-speculation, instead using only re-execution to recover. Idempotent processors execute programs as a sequence of compiler-constructed idempotent (re-executable) regions. The nature of these regions allows precise state to be reproduced by re-execution, obviating the need for hardware recovery support. We build upon the insight that programs naturally decompose into a series of idempotent regions and that these regions can be large. The paradigm of executing idempotent regions, which we call idempotent processing, can be used to support various types of speculation, including branch prediction, dependence prediction, or execution in the presence of hardware faults or exceptions.
Marc de Kruijf, Karthikeyan Sankaralingam
MICRO1
2010 Design and implementation of the PLUG architecture for programmable and efficient network lookups
abstract
This paper proposes a new architecture called Pipelined LookUp Grid (PLUG) that can perform data structure lookups in network processing. PLUGs are programmable and through simplicity achieve power efficiency. We draw upon one key insights: data structure lookups have natural structure that can be statically determined and exploited. The PLUG execution model transforms data-structure lookups into pipelined stages of computation and associates small code-blocks with data. The PLUG architecture is a tiled architecture with each tile consisting predominantly of SRAMs, a lightweight no-buffering router, and an array of lightweight computation cores. Using a principle of fixed delays in the execution model, the architecture is contention-free and completely statically scheduled thus achieving high energy efficiency. The architecture enables rapid deployment of new network protocols and generalizes as a data-structure accelerator.
Amit Kumar 0007, Lorenzo De Carli, Marc de Kruijf, Karthikeyan Sankaralingam, Cristian Estan, Somesh Jha
PACT4
2010 A unified model for timing speculation: Evaluating the impact of technology scaling, CMOS design style, and fault recovery mechanism
abstract
Due to fundamental device properties, energy efficiency from CMOS scaling is showing diminishing improvements. To overcome the energy efficiency challenges, timing speculation has been proposed to optimize for common-case timing conditions, with errors occurring under worst-case conditions detected and corrected in hardware. Although various timing speculation techniques have been proposed, no general framework exists for reasoning about the trade-offs and high-level design considerations of timing speculation. This paper develops two models to study the end-to-end behavior of timing speculation: a hardware-level efficiency model that considers the effects of process variations on path delays, and a complementary system-level recovery model. When combined, the models are used to assess the impact of technology scaling, CMOS design style, and fault recovery mechanism on the efficiency of timing speculation. Our results show that (1) efficiency gains from timing speculation do not improve as technology scales, (2) ultra-low power (sub-threshold) CMOS designs benefit most from timing speculation - we report a 47% potential energy-delay reduction, and (3) fine-grained fault recovery is key to significant energy improvements. The combined model uses only high-level inputs to derive quantitative energy efficiency benefits without any need for detailed simulation, making it a potentially useful tool for hardware developers.
Marc de Kruijf, Shuou Nomura, Karthikeyan Sankaralingam
DSN1
2010 Relax: an architectural framework for software recovery of hardware faults
abstract
As technology scales ever further, device unreliability is creating excessive complexity for hardware to maintain the illusion of perfect operation. In this paper, we consider whether exposing hardware fault information to software and allowing software to control fault recovery simplifies hardware design and helps technology scaling.
Marc de Kruijf, Shuou Nomura, Karthikeyan Sankaralingam
ISCA1