EDBT 2026 Demo / reviewers in the wild / expert
Nils Asmussen
dblp:162/7710
· DBLP profile ↗
11ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-4232-4519ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TEEM³: Core-Independent and Cooperating Trusted Execution EnvironmentsabstractTrusted Execution Environments (TEEs) enable secure code execution on machines that are not fully trusted by the user who runs the workload. However, existing TEE solutions mostly target CPUs and are typically tied to one specific instruction set architecture. Although some accelerators also provide support for TEEs, this leads to multiple, different TEE implementations on the same system, increasing its complexity and trusted computing base (TCB). This challenge becomes particularly apparent when workloads span heterogeneous processing units, because the diversity of TEE implementations complicates the creation of secure communication channels between the individual TEEs. Nils Asmussen, Sebastian Haas, Carsten Weinhold, Nicholas Gordon, Stephan Gerhold, Friedrich Pauls, Nilanjana Das, Michael Roitzsch |
ASPLOS (2) | 1 |
| 2025 | Distrusting cores by separating computation from isolationabstractSecurity mechanisms such as address spaces rely on the assumption that processor cores can be fully trusted. But the steady influx of side-channel vulnerabilities in processors is challenging this assumption. To minimize the impact of security vulnerabilities in processors, we need a system architecture that can tolerate potentially exploitable cores. In this paper, we propose the untrusted core isolation model to protect critical computation on trusted cores from untrusted and potentially buggy cores. We survey how current architectural building blocks such as MMUs fall short of this goal and derive requirements for untrusted core isolation. To demonstrate its feasibility, we discuss both changes to commodity platforms and show how research works such as fulfill the requirements. We evaluate the security benefits via a qualitative comparison of current architectures in both industry and academia and study its costs by a quantitative comparison of the most promising approaches on off-the-shelf and FPGA-based platforms. Nils Asmussen, Till Miemietz, Sebastian Haas, Michael Roitzsch |
J. Syst. Archit. | 1 |
| 2024 | Core-Local Reasoning and Predictable Cross-Core Communication with M3abstractModern cyber-physical systems often require security, heterogeneity, and real-time operation from their hardware platform and operating system. However, highly predictable real-time operating systems such as FreeRTOS do not employ strong component isolation required for platform security. Microkernels implement such isolation using virtual memory and code running in the privileged CPU mode, complicating real-time analysis. In this work, we start with a different architectural approach: M3 is an existing hardware/software co-design for heterogeneous systems that features strong isolation between cores. However, the real-time properties of this platform have not been investigated. We first survey M3{\prime}s current state for real-time applicability and study both the communication latencies in comparison to other systems and M3's different approach to task priorities. Furthermore we improve M3's real-time applicability by adding a network-on-chip traffic regulation and enabling the enforcement of resource limits. With these additions, M3 enables local reasoning about application execution. We perform the evaluation with an FPGA-based hardware prototype and in simulation based on gem5. Nils Asmussen, Sebastian Haas, Adam Lackorzynski, Michael Roitzsch |
RTAS | 1 |
| 2023 | Software-Defined CPU ModesabstractOur CPUs contain a compute instruction set, which regular applications use. But they also feature an intricate underworld of different CPU modes, combined with trap and exception handling to transition between these modes. These mechanisms are manifold and complex, yet the layering and functionality offered by the CPU modes is fixed. We have to take what CPU vendors provide, including potential security problems from unneeded modes. This paper explores the question, whether CPU modes could instead be defined entirely by software. We show how such a design would function and explore the advantages it enables. We believe that pushing all existing modes under a common design umbrella would enforce a cleaner structure and more control over exposed functionality. At the same time, the flexibility of software-defined modes enables interesting new use cases. Michael Roitzsch, Till Miemietz, Christian von Elm, Nils Asmussen |
HotOS | 4 |
| 2022 | Efficient and scalable core multiplexing with M³vabstractThe M³ system (ASPLOS ’16) proposed a hardware/software co-design that simplifies integration between general-purpose cores and special-purpose accelerators, allowing users to easily utilize them in a unified manner. M³ is a tiled architecture, whose tiles (cores and accelerators) are partitioned between applications, such that each tile is dedicated to its own application. The M³x system (ATC ’19) extended M³ by trading off some isolation to enable coarse-grained multiplexing of tiles among multiple applications. With M³x, if source tile t₁ runs code of application p and sends a message m to destination tile t₂ while t₂ is currently not associated with p, then m is forwarded to the right place through a “slow path”, via some special OS tile. In this paper, we present M³v, which extends M³x by further trading off some isolation between applications to support “fast path” communication that does not require the said OS tile’s involvement. Thus, with M³v, a tile can be efficiently multiplexed between applications provided it is a general-purpose core. M³v achieves this goal by 1) adding a local multiplexer to each such core, and by 2) virtualizing the core’s hardware component responsible for cross-tile communications. We prototype M³v using RISC-V cores on an FPGA platform and show that it significantly outperforms M³x and may achieve competitive performance to Linux. Nils Asmussen, Sebastian Haas, Carsten Weinhold, Till Miemietz, Michael Roitzsch |
ASPLOS | 1 |
| 2022 | Slashing the disaggregation tax in heterogeneous data centers with FractOSabstractDisaggregated heterogeneous data centers promise higher efficiency, lower total costs of ownership, and more flexibility for data-center operators. However, current software stacks can levy a high tax on application performance. Applications and OSes are designed for systems where local PCIe-connected devices are centrally managed by CPUs, but this centralization introduces unnecessary messages through the shared data-center network in a disaggregated system. Lluís Vilanova, Lina Maudlej, Shai Bergman, Till Miemietz, Matthias Hille, Nils Asmussen, Michael Roitzsch, Hermann Härtig, Mark Silberstein |
EuroSys | 6 |
| 2019 | M³x: Autonomous Accelerators via Context-Enabled Fast-Path Communication
Nils Asmussen, Michael Roitzsch, Hermann Härtig |
USENIX ATC | 1 |
| 2019 | SemperOS: A Distributed Capability System
Matthias Hille, Nils Asmussen, Pramod Bhatotia, Hermann Härtig |
USENIX ATC | 2 |
| 2016 | M3: A Hardware/Operating-System Co-Design to Tame Heterogeneous ManycoresabstractIn the last decade, the number of available cores increased and heterogeneity grew. In this work, we ask the question whether the design of the current operating systems (OSes) is still appropriate if these trends continue and lead to abundantly available but heterogeneous cores, or whether it forces a fundamental rethinking of how systems are designed. We argue that: 1. hiding heterogeneity behind a common hardware interface unifies, to a large extent, the control and coordination of cores and accelerators in the OS, 2. isolating at the network-on-chip rather than with processor features (like privileged mode, memory management unit, ...), allows running untrusted code on arbitrary cores, and 3. providing OS services via protocols over the network-on-chip, instead of via system calls, makes them accessible to arbitrary types of cores as well. Nils Asmussen, Marcus Völp, Benedikt Noethen, Hermann Härtig, Gerhard P. Fettweis |
ASPLOS | 1 |
| 2015 | Towards dependable CPS infrastructures: Architectural and operating-system challengesabstractCyber-physical systems (CPSs), due to their direct influence on the physical world, have to meet extended security and dependability requirements. This is particularly true for CPS that operate in close proximity to humans or that control resources that, when tampered with, put all our lives at stake. In this paper, we review the challenges and some early solutions that arise at the architectural and operating-system level when we require cyber-physical systems and CPS infrastructure to withstand advanced and persistent threats. We found that although some of the challenges we identified are already matched by rudimentary solutions, further research is required to ensure sustainable and dependable operation of physically exposed CPS infrastructure and, more importantly, to guarantee graceful degradation in case of malfunction or attack. Marcus Völp, Nils Asmussen, Hermann Härtig, Benedikt Noethen, Gerhard P. Fettweis |
ETFA | 2 |
| 2015 | Demo abstract: Taming many heterogeneous coresabstractMany-core systems are increasingly used in real-time settings to meet the performance requirements of advanced applications such as the classification and tracking of dynamic objects for autonomous driving [1] or the generation of safe trajectories through rough terrain [2]. Task sets of these applications are often mixtures of short running, low latency tasks, such as the various filtering steps required for signal or image processing, and long running tasks, such as route planning, which occupy their assigned core for extended periods of time. Short running tasks often follow a data flow programming paradigm and are organized into directed acyclic graphs (DAG) based on their input-/output-dependencies. Once these dependencies are met, they execute without further task interactions until they complete producing outputs for subsequent tasks. Long running tasks on the other hand interact frequently with other tasks, accessing data located in the memories of remote cores or interacting with operating-system services. This demonstrator shows how both types of applications can be integrated into a single many-core architecture. Nils Asmussen, Marcus Völp, Benedikt Noethen, Annett Ungethüm |
RTAS | 1 |