Michael Roitzsch

dblp:18/2206 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-2416-6537ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 TEEM³: Core-Independent and Cooperating Trusted Execution Environments
abstract
Trusted Execution Environments (TEEs) enable secure code execution on machines that are not fully trusted by the user who runs the workload. However, existing TEE solutions mostly target CPUs and are typically tied to one specific instruction set architecture. Although some accelerators also provide support for TEEs, this leads to multiple, different TEE implementations on the same system, increasing its complexity and trusted computing base (TCB). This challenge becomes particularly apparent when workloads span heterogeneous processing units, because the diversity of TEE implementations complicates the creation of secure communication channels between the individual TEEs.
Nils Asmussen, Sebastian Haas, Carsten Weinhold, Nicholas Gordon, Stephan Gerhold, Friedrich Pauls, Nilanjana Das, Michael Roitzsch
ASPLOS (2)8
2026 BRUMM: A Case for Predictable Memory Reclamation
abstract
Edge data centers process latency-sensitive workloads of nearby Internet-of-Things devices. These security-critical, multi-tenant environments are equipped with comparatively limited compute resources. Consequently, resource management must be fast and predictable even in the presence of malicious tenants because there is no surplus of resources to compensate for performance attacks. In particular, there is a need to constrain the time it takes to reclaim memory from applications. Existing accounting mechanisms in operating systems focus on limiting memory or scheduling-time usage; they provide no guarantees about the latency of resource reclamation, which can vary greatly and is a potential vector for performance attacks. To solve this standing issue, we introduce BRUMM (Bounded Reclamation of User-space Memory Mappings): This accounting-driven mechanism predicts and tracks how long it will take to reclaim memory allocated to applications, enforcing an upper limit via a configurable latency budget. As a case study, we extended the L4Re microkernel to add a quota object for reclamation latency and enforce its limit. Our evaluation demonstrates that this implementation of BRUMM achieves a consistent overestimation of reclamation latency, staying within the same order of magnitude to real, measured latencies. The implementation only shows modest performance overhead on kernel operations, ranging from 2.4% overhead for simple system calls to 28% in synthetic worst-case scenarios. BRUMM makes reclamation latency a first-class resource that can be accounted for, thereby improving isolation and reliability in edge clouds.
Viktor Reusch, Michael Roitzsch, Horst Schirmeier
ECRTS2
2026 Scheduling Constraints: A Universal OS Mechanism for Managing Shared Resources
Moritz Lumme, Michael Roitzsch, Adam Lackorzynski
RTAS2
2025 CoRD: Converged RDMA Dataplane
abstract
HPC networking is often characterized by kernel bypass, which is considered mandatory for large parallel and distributed applications. However, kernel bypass comes at a price because it breaks the traditional OS architecture, requiring applications to use special APIs and limiting the OS's control over existing network connections. We make the case that kernel bypass is not mandatory. Rather, high-performance networking relies on multiple performance-improving techniques, with kernel bypass even being detrimental to performance under specific conditions. CoRD removes kernel bypass from RDMA networks, primarily to enable efficient OS-level control over the RDMA dataplane. This control can be used to enhance security or resource allocation policies, and, as we demonstrate in one of the use cases, can improve end-to-end application performance by up to 10%. This architecture can enable Cloud-based distributed RDMA applications and facilitate deployment of coupled HPC applications.
Maksym Planeta, Jan Bierbaum, Michael Roitzsch, Hermann Härtig
IPDPS3
2025 MettEagle: Costs and Benefits of Implementing Containers on Microkernels
Till Miemietz, Viktor Reusch, Matthias Hille, Lars Wrenger, Jana Eisoldt, Jan Klötzke, Max Kurze, Adam Lackorzynski, Michael Roitzsch, Hermann Härtig
OSDI9
2025 Separate but Together: Integrating Remote Attestation into TLS
Carsten Weinhold, Muhammad Usama Sardar, Ionut Mihalcea, Yogesh Deshpande, Hannes Tschofenig, Yaron Sheffer, Thomas Fossati, Michael Roitzsch
USENIX ATC8
2025 Distrusting cores by separating computation from isolation
abstract
Security mechanisms such as address spaces rely on the assumption that processor cores can be fully trusted. But the steady influx of side-channel vulnerabilities in processors is challenging this assumption. To minimize the impact of security vulnerabilities in processors, we need a system architecture that can tolerate potentially exploitable cores. In this paper, we propose the untrusted core isolation model to protect critical computation on trusted cores from untrusted and potentially buggy cores. We survey how current architectural building blocks such as MMUs fall short of this goal and derive requirements for untrusted core isolation. To demonstrate its feasibility, we discuss both changes to commodity platforms and show how research works such as fulfill the requirements. We evaluate the security benefits via a qualitative comparison of current architectures in both industry and academia and study its costs by a quantitative comparison of the most promising approaches on off-the-shelf and FPGA-based platforms.
Nils Asmussen, Till Miemietz, Sebastian Haas, Michael Roitzsch
J. Syst. Archit.4
2024 Core-Local Reasoning and Predictable Cross-Core Communication with M3
abstract
Modern cyber-physical systems often require security, heterogeneity, and real-time operation from their hardware platform and operating system. However, highly predictable real-time operating systems such as FreeRTOS do not employ strong component isolation required for platform security. Microkernels implement such isolation using virtual memory and code running in the privileged CPU mode, complicating real-time analysis. In this work, we start with a different architectural approach: M3 is an existing hardware/software co-design for heterogeneous systems that features strong isolation between cores. However, the real-time properties of this platform have not been investigated. We first survey M3{\prime}s current state for real-time applicability and study both the communication latencies in comparison to other systems and M3's different approach to task priorities. Furthermore we improve M3's real-time applicability by adding a network-on-chip traffic regulation and enabling the enforcement of resource limits. With these additions, M3 enables local reasoning about application execution. We perform the evaluation with an FPGA-based hardware prototype and in simulation based on gem5.
Nils Asmussen, Sebastian Haas, Adam Lackorzynski, Michael Roitzsch
RTAS4
2023 Software-Defined CPU Modes
abstract
Our CPUs contain a compute instruction set, which regular applications use. But they also feature an intricate underworld of different CPU modes, combined with trap and exception handling to transition between these modes. These mechanisms are manifold and complex, yet the layering and functionality offered by the CPU modes is fixed. We have to take what CPU vendors provide, including potential security problems from unneeded modes. This paper explores the question, whether CPU modes could instead be defined entirely by software. We show how such a design would function and explore the advantages it enables. We believe that pushing all existing modes under a common design umbrella would enforce a cleaner structure and more control over exposed functionality. At the same time, the flexibility of software-defined modes enables interesting new use cases.
Michael Roitzsch, Till Miemietz, Christian von Elm, Nils Asmussen
HotOS1
2022 Efficient and scalable core multiplexing with M³v
abstract
The M³ system (ASPLOS ’16) proposed a hardware/software co-design that simplifies integration between general-purpose cores and special-purpose accelerators, allowing users to easily utilize them in a unified manner. M³ is a tiled architecture, whose tiles (cores and accelerators) are partitioned between applications, such that each tile is dedicated to its own application. The M³x system (ATC ’19) extended M³ by trading off some isolation to enable coarse-grained multiplexing of tiles among multiple applications. With M³x, if source tile t₁ runs code of application p and sends a message m to destination tile t₂ while t₂ is currently not associated with p, then m is forwarded to the right place through a “slow path”, via some special OS tile. In this paper, we present M³v, which extends M³x by further trading off some isolation between applications to support “fast path” communication that does not require the said OS tile’s involvement. Thus, with M³v, a tile can be efficiently multiplexed between applications provided it is a general-purpose core. M³v achieves this goal by 1) adding a local multiplexer to each such core, and by 2) virtualizing the core’s hardware component responsible for cross-tile communications. We prototype M³v using RISC-V cores on an FPGA platform and show that it significantly outperforms M³x and may achieve competitive performance to Linux.
Nils Asmussen, Sebastian Haas, Carsten Weinhold, Till Miemietz, Michael Roitzsch
ASPLOS5
2022 Slashing the disaggregation tax in heterogeneous data centers with FractOS
abstract
Disaggregated heterogeneous data centers promise higher efficiency, lower total costs of ownership, and more flexibility for data-center operators. However, current software stacks can levy a high tax on application performance. Applications and OSes are designed for systems where local PCIe-connected devices are centrally managed by CPUs, but this centralization introduces unnecessary messages through the shared data-center network in a disaggregated system.
Lluís Vilanova, Lina Maudlej, Shai Bergman, Till Miemietz, Matthias Hille, Nils Asmussen, Michael Roitzsch, Hermann Härtig, Mark Silberstein
EuroSys7
2019 K2: Work-Constraining Scheduling of NVMe-Attached Storage
abstract
For data-driven cyber-physical systems, timely access to storage is an important building block of real-time guarantees. At the same time, storage technology undergoes continued technological advancements. The introduction of NVMe fundamentally changes the interface to the drive by exposing request parallelism available at the flash package level to the storage stack, allowing to extract higher throughput and lower latencies from the drive. The resulting architectural changes within the operating system render many historical designs and results obsolete, requiring a fresh look at the I/O scheduling landscape. In this paper, we conduct a comprehensive survey of the existing NVMe-compatible I/O schedulers in Linux regarding their suitability for real-time applications. We find all schedulers severely lacking in terms of performance isolation and tail latencies. Therefore, we propose K2, a new I/O scheduler specifically designed to reduce latency at the 99.9th percentile, while maintaining the throughput gains promised by NVMe. By limiting the length of NVMe device queues, K2 reduces read latencies up to 10× and write latencies up to 6.8×, while penalizing throughput for non-real-time background load by at most 2.7×.
Till Miemietz, Hannes Weisbach, Michael Roitzsch, Hermann Härtig
RTSS3
2019 M³x: Autonomous Accelerators via Context-Enabled Fast-Path Communication
Nils Asmussen, Michael Roitzsch, Hermann Härtig
USENIX ATC2
2017 Lateral Thinking for Trustworthy Apps
abstract
The growing computerization of critical infrastructure as well as the pervasiveness of computing in everyday life has led to increased interest in secure application development. We observe a flurry of new security technologies like ARM TrustZone and Intel SGX, but a lack of a corresponding architectural vision. We are convinced that point solutions are not sufficient to address the overall challenge of secure system design. In this paper, we outline our take on a trusted component ecosystem of small individual building blocks with strong isolation. In our view, applications should no longer be designed as massive stacks of vertically layered frameworks, but instead as horizontal aggregates of mutually isolated components that collaborate across machine boundaries to provide a service. Lateral thinking is needed to make secure systems going forward.
Hermann Härtig, Michael Roitzsch, Carsten Weinhold, Adam Lackorzynski
ICDCS2
2017 E-Team: Practical Energy Accounting for Multi-Core Systems
Till Smejkal, Marcus Hähnel, Thomas Ilsche, Michael Roitzsch, Wolfgang E. Nagel, Hermann Härtig
USENIX ATC4
2013 Atlas: Look-ahead scheduling using workload metrics
abstract
From video and music to user interface animations, a lot of real-time workloads run on today's desktops and mobile devices, yet commodity operating systems offer scheduling interfaces like nice-levels, priorities or shares that do not adequately convey timing requirements. Real-time research offers many solutions with strong timeliness guarantees, but they often require a periodic task model and ask the developer for information that is hard to obtain like execution times or reservation budgets. Within this design space of easy programming, but weak guarantees on one hand and strong guarantees, but harder development on the other, we propose Atlas, the Auto-Training Look-Ahead Scheduler. With a simple yet powerful interface it relies exclusively on data from the application domain: It uses deadlines to express timing requirements and workload metrics to express resource requirements. It replaces implicit knowledge of future job releases as provided by periodic tasks with explicit job submission to enable look-ahead scheduling. Using video playback as a dynamic high-throughput load, we show that the proposed workload metrics are sufficient for Atlas to know an application's execution time behavior ahead of time. Atlas' predictions have a typical relative error below 10%.
Michael Roitzsch, Stefan Wachtler, Hermann Härtig
IEEE Real-Time and Embedded Technology and Applications Symposium1
2010 Capability wrangling made easy: debugging on a microkernel with valgrind
abstract
Not all operating systems are created equal. Contrasting traditional monolithic kernels, there is a class of systems called microkernels more prevalent in embedded systems like cellphones, chip cards or real-time controllers. These kernels offer an abstraction very different from the classical POSIX interface. The resulting unfamiliarity for programmers complicates development and debugging. Valgrind is a well-known debugging tool that virtualizes execution to perform dynamic binary analysis. However, it assumes to run on a POSIX-like kernel and closely interacts with the system to control execution. In this paper we analyze how to adapt Valgrind to a non-POSIX environment and describe our port to the Fiasco.OC microkernel. Additionally, we analyze bug classes that are indigenous to capability systems and show how Valgrind's flexibility can be leveraged to create custom debugging tools detecting these errors.
Aaron Pohle, Björn Döbel, Michael Roitzsch, Hermann Härtig
VEE3
2008 Video quality and system resources: Scheduling two opponents
Michael Roitzsch, Martin Pohlack
J. Vis. Commun. Image Represent.1
2007 Probabilistic Admission Control to Govern Real-Time Systems under Overload
abstract
Existing real-time research focuses on how to formulate. model and enforce timeliness guarantees for task sets whose correctness has a temporal aspect. However; the resulting systems often exhibit poor resource utilization due to the resource scheduler reserving more resources than required in order to ensure that admitted schedules can be satisfied under worst case conditions. Weakening the guarantees leads to the known concepts of firm and soft real-time tasks, butt we think the paradigm needs to be shifted further,: reifying efficient utilization. With Quality-Assuring Scheduling (QAS) we presented such an algorithm. However: its practical applicability is restricted to uniform and harmonic periods, due to its complexity for arbitrary periods. To overcome this limitation, we introduce Quality-Rate-Monotonic Scheduling (QRMS), which, although slightly more pessimistic, is less complex compared to QAS. Thee admission control is again based on a probabilistic model to ensure that a requested fraction of jobs is successfully executed. Thus the amount of missed deadlines can be externally controlled, even in sustained overload situations.
Claude-Joachim Hamann, Michael Roitzsch, Lars Reuther, Jean Wolter, Hermann Härtig
ECRTS2
2007 Slice-balancing H.264 video encoding for improved scalability of multicore decoding
abstract
With multicore architectures being introduced to the market, the research community is revisiting problems to evaluate them under the new preconditions set by those new systems. Algorithms need to be implemented with scalability in mind. One problem that is known to be computationally demanding is video decoding. In this paper, we will present a technique that increases the scalability of H.264 video decoding by modifying only the encoder stage. In embedded scenarios, increased scalability can also enable reduced clock speeds of the individual cores, thus lowering overall power consumption.
Michael Roitzsch
EMSOFT1
2006 Principles for the Prediction of Video Decoding Times Applied to MPEG-1/2 and MPEG-4 Part 2 Video
abstract
In this paper, we present a method to predict per-frame decoding times of modern video decoder algorithms. By examining especially the MPEG-1, MPEG-2, and MPEG-4 pt. 2 algorithms, we developed a generic model for these decoders, which also applies to a wide range of other decoders. From this model, we derived a method to predict decoding times with an up-to-now unmatched accuracy while keeping the overhead low. We show the effectiveness of this method with an example implementation and compare the resulting predictions with the actual decoding times using video material from commercial DVDs
Michael Roitzsch, Martin Pohlack
RTSS1