VLDB 2026 Research / reviewers in the wild / expert
Nicolai Oswald
dblp:217/6692
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0009-9272-0518ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | vCXLGen: Automated Synthesis and Verification of CXL Bridges for Heterogeneous ArchitecturesabstractCompute Express Link (CXL) offers byte-addressable, cache-coherent remote memory accesses across multiple hosts. Unfortunately, the CXL specification lacks mechanisms to ensure safe interoperability between heterogeneous host architectures with diverse cache coherence (CC) protocols and memory consistency models (MCMs). This semantic gap poses fundamental challenges and a significant barrier to adopting CXL in modern heterogeneous data centers. Anatole Lefort, Julian Pritzi, Nicolò Carpentieri, David Schall, Simon Dittrich, Soham Chakraborty 0001, Nicolai Oswald, Pramod Bhatotia |
ASPLOS (2) | 7 |
| 2026 | C³: CXL Coherence Controllers for Heterogeneous ArchitecturesabstractWe introduce$\mathbf{C}^{\mathbf{3}}$, a systematic methodology for designing Compute Express Link (CXL) coherence controllers, to overcome interoperability challenges that arise from the mismatch of coherence protocols and memory consistency models in heterogeneous CXL-connected systems. Crucially, CXL lacks a unified heterogeneous computing interface, which can lead to unpredictable and inconsistent behavior when multiple heterogeneous devices decide to share cache-coherent CXL memory. C$^{3}$acts as a pivotal interface between diverse heterogeneous compute units, bridging the semantic differences without necessitating disruptive changes to existing system architectures. Our approach hinges on two key principles: delegating memory operations across coherence domains and enforcing atomicity at domain boundaries, thereby preserving the native memory consistency model semantics of each unit. We implement$\mathbf{C}^{\mathbf{3}}$as a generic gem5 model and validate its correctness through exhaustive litmus testing. We also show that$\mathbf{C}^{\mathbf{3}}$incurs minimal performance overhead compared to unified native coherence protocols. Anatole Lefort, David Schall, Nicolò Carpentieri, Julian Pritzi, Soham Chakraborty 0001, Nicolai Oswald, Pramod Bhatotia |
HPCA | 6 |
| 2025 | CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingabstractIncreasingly, multi-processing unit (PU) systems (e.g., CPU-GPU, multi-CPU, multi-GPU, etc.) are embracing cache-coherent shared memory to facilitate inter-PU communication.The coherence protocols in these systems support write-through accesses that place the data directly at the LLC to enable efficient producer-consumer communications pervasive in AI/ML workloads.Moreover, release consistency has emerged as the standard memory model in such systems due to its programming simplicity and ability to support high performance.In today's multi-PU systems, the source processor that issues the writes also orders them to enforce release consistency, even for write-through accesses.Unfortunately, such source ordering of write-through operations results in unnecessary communications between the source processor and the LLC directory, incurring significant performance, interconnect traffic, and energy overheads for multi-PU applications.To eliminate such communication, we present cord 1 , a novel cache coherence protocol that orders write-through accesses directly at the cache directory.cord employs several novel mechanisms to minimize the metadata required for ordering traffic while efficiently scaling to multiple directories.Evaluations atop the gem5 simulator show that compared to source ordering, cord improves application performance by 24% and reduces traffic by 13% on average while incurring < 1% storage, area, and power overheads.Compared to hand-optimized message-passing implementations, cord observes a mere 3% performance overhead and 6% more traffic on average with a significantly simpler programming model. Yanpeng Yu, Nicolai Oswald, Anurag Khandelwal |
ISCA | 2 |
| 2024 | PipeGen: Automated Transformation of a Single-Core Pipeline into a Multicore Pipeline for a Given Memory Consistency ModelabstractDesigning a pipeline for a multicore processor is difficult. One major challenge is designing it such that the pipeline correctly enforces the intended memory consistency model (MCM). We have developed the PipeGen design automation tool to allow architects to start with a single core pipeline that only enforces single-threaded correctness and automatically transform it to enforce a given MCM. Our key innovation is a set of compiler-like transformations that codify three different ways of enforcing memory ordering at the pipeline. We have validated that PipeGen correctly enforces the ARMv8 and x86TSO MCMs on three distinct pipeline implementations, using litmus tests with the Murphi model checker. An Qi Zhang, Andres Goens, Nicolai Oswald, Tobias Grosser, Daniel J. Sorin, Vijay Nagarajan |
PACT | 3 |
| 2024 | Determining the Minimum Number of Virtual Networks for Different Coherence ProtocolsabstractWe revisit the question of how many virtual networks (VNs) are required to provably avoid deadlock in a cache coherence protocol. The textbook way of reasoning about VNs says that the number of VNs depends on the longest chain of message dependencies in the protocol. We show that this conventional wisdom is incorrect and results in a number of virtual networks that is neither necessary nor sufficient for the general system model of an arbitrary interconnection network (ICN) topology and multiple directories. We have created a formalism for modeling coherence protocols and their interactions with ICN queueing. Using that formalism, we have developed an algorithm that (a) determines the minimum number of virtual networks required to avoid deadlock and (b) generates the mappings from message types to virtual networks. Weihang Li, Andres Goens, Nicolai Oswald, Vijay Nagarajan, Daniel J. Sorin |
ISCA | 3 |
| 2023 | Āpta: Fault-tolerant object-granular CXL disaggregated memory for accelerating FaaSabstractAs cloud workloads increasingly adopt the fault-tolerant Function-as-a-Service (FaaS) model, demand for improved performance has increased. Alas, the performance of FaaS applications is heavily bottlenecked by the remote object store in which FaaS objects are maintained. We identify that the upcoming CXL-based cache-coherent disaggregated memory is a promising technology for maintaining FaaS objects. Our analysis indicates that CXL's low-latency, high-bandwidth access characteristics coupled with compute-side caching of objects, provides significant performance potential over an in-memory RDMA-based object store. We observe however that CXL lacks the requisite level of fault-tolerance necessary to operate at an inter-server scale within the datacenter. Furthermore, its cache-line granular accesses impose inefficiencies for object-granular data store accesses. We propose Āpta, a CXL-based object-granular memory interface for maintaining FaaS objects. Āpta's key innovation is a novel fault-tolerant coherence protocol for keeping the cached objects consistent without compromising availability in the face of compute server failures. Our evaluation of Āpta using 6 full FaaS application workflows (totaling 26 functions) indicates that it outperforms a state-of-the-art fault-tolerant object caching protocol on an RDMA-based system by 21-90% and an uncached CXL-based system by 15-42%. Adarsh Patil 0002, Vijay Nagarajan, Nikos Nikoleris, Nicolai Oswald |
DSN | 4 |
| 2023 | Compound Memory ModelsabstractToday's mobile, desktop, and server processors are heterogeneous, consisting not only of CPUs but also GPUs and other accelerators. Such heterogeneous processors are starting to expose a shared memory interface across these devices.Given that each of these individual devices typically supports a distinct instruction set architecture and a distinct memory consistency model, it is not clear what the memory consistency model of the heterogeneous machine should be. In this paper, we answer this question by formalizing "compound" memory models: we present a compositional operational model describing the resulting model when devices with distinct consistency models are fused together. We instantiate our model with the compound x86TSO/PTX model -- a CPU enforcing x86TSO and a GPU enforcing the PTX model. A key result is that the x86TSO/PTX compound model retains compiler mappings from the language-based (scoped) C memory model. This means that threads mapped to the x86TSO device can continue to use the already proven C-to-x86TSO compiler mapping, and the same for PTX. Andres Goens, Soham Chakraborty 0001, Susmit Sarkar, Sukarn Agarwal, Nicolai Oswald, Vijay Nagarajan |
Proc. ACM Program. Lang. | 5 |
| 2022 | HeteroGen: Automatic Synthesis of Heterogeneous Cache Coherence ProtocolsabstractWe solve the two challenges architects face when designing heterogeneous processors with cache coherent shared memory. First, we develop an automated tool, called HeteroGen, for composing clusters of cores, each with its own coherence protocol. Second, we show that the output of HeteroGen adheres to a precisely defined memory consistency model that we call a compound consistency model. For a wide variety of protocols—including the MOESI variants, as well as those that are targeted towards Total Store Order and Release Consistency—we show that HeteroGen can correctly fuse them. To validate HeteroGen, we develop the first litmus tests for verifying that heterogeneous protocols satisfy compound consistency models. To understand the possible performance implications of automatic protocol generation, we compared against a publicly available manually-generated heterogeneous protocol. Our results show that performance is comparable. Nicolai Oswald, Vijay Nagarajan, Daniel J. Sorin, Vasilis Gavrielatos, Theo X. Olausson, Reece Carr |
HPCA | 1 |
| 2021 | Dvé: Improving DRAM Reliability and Performance On-Demand via Coherent ReplicationabstractAs technologies continue to shrink, memory system failure rates have increased, demanding support for stronger forms of reliability. In this work, we take inspiration from the two-tier approach that decouples correction from detection and explore a novel extrapolation. We propose Dvé, a hardware-driven replication mechanism where data blocks are replicated in 2 different sockets across a cache-coherent NUMA system. Each data block is also accompanied by a code with strong error detection capabilities so that when an error is detected, correction is performed using the replica. Such an organization has the advantage of offering two independent points of access to data which enables: (a) strong error correction that can recover from a range of faults affecting any of the components in the memory, upto and including the memory controller, and (b) higher performance by providing another nearer point of memory access. Dvé realizes both of these benefits via Coherent Replication, a technique that builds on top of existing cache coherence protocols for not only keeping the replicas in sync for reliability, but also to provide coherent access to the replicas during fault-free operation for performance. Dvé can flexibly provide these benefits on-demand by simply using the provisioned memory capacity which, as reported in recent studies, is often underutilized in today’s systems. Thus, Dvé introduces a unique design point that offers higher reliability and performance for workloads that do not require the entire memory capacity. Adarsh Patil 0002, Vijay Nagarajan, Rajeev Balasubramonian, Nicolai Oswald |
ISCA | 4 |
| 2020 | HieraGen: Automated Generation of Concurrent, Hierarchical Cache Coherence ProtocolsabstractWe present HieraGen, a new tool for automatically generating hierarchical cache coherence protocols. HieraGen's inputs are the simple, atomic, stable state protocols for each level of the hierarchy. HieraGen's output is a highly concurrent hierarchical protocol, in the form of the finite state machines for all of the cache and directory controllers. HieraGen thus reduces the complexity that architects face, by offloading the challenging tasks of composing protocols and managing concurrency. Experiments show that HieraGen can automatically generate correct-by-construction MOESI family of hierarchical protocols with dozens of states and hundreds of transitions. We have verified all of the generated protocols for safety and deadlock freedom using a model checker. Nicolai Oswald, Vijay Nagarajan, Daniel J. Sorin |
ISCA | 1 |
| 2018 | Scale-out ccNUMA: exploiting skew with strongly consistent cachingabstractToday's cloud based online services are underpinned by distributed key-value stores (KVS). Such KVS typically use a scale-out architecture, whereby the dataset is partitioned across a pool of servers, each holding a chunk of the dataset in memory and being responsible for serving queries against the chunk. One important performance bottleneck that a KVS design must address is the load imbalance caused by skewed popularity distributions. Despite recent work on skew mitigation, existing approaches offer only limited benefit for high-throughput in-memory KVS deployments. Vasilis Gavrielatos, Antonios Katsarakis, Arpit Joshi, Nicolai Oswald, Boris Grot, Vijay Nagarajan |
EuroSys | 4 |
| 2018 | ProtoGen: Automatically Generating Directory Cache Coherence Protocols from Atomic SpecificationsabstractDesigning directory cache coherence protocols is complicated because coherence transactions are not atomic in modern multicore processors. A coherence transaction comprises multiple messages, and these messages can interleave with other conflicting coherence transactions initiated by other cores. To overcome this architectural challenge, we present ProtoGen, an automated tool for taking the description of a directory protocol with atomic transactions (i.e., no concurrency) and generating the corresponding protocol for a multicore with non-atomic transactions. ProtoGen outputs the finite state machines for the cache and directory controllers, including all of the transient states that are possible with concurrent transactions. We have used ProtoGen to generate complete MSI, MESI, and MOSI protocols given their stable state protocol specifications. We have verified the generated protocols for safety and deadlock freedom using the Murφ model checker. Our generated protocols are identical to or better than manually generated protocols, at times even discovering opportunities to reduce stalling. Nicolai Oswald, Vijay Nagarajan, Daniel J. Sorin |
ISCA | 1 |