EDBT 2026 Demo / reviewers in the wild / expert
Stefan Lankes
dblp:57/4163
· DBLP profile ↗
28ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-4718-2238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Browser to Kernel: Exploring a Lightweight Sandboxed Approach for Unikernel ExtensionsabstractLibrary Operating Systems (libOS) are highly efficient because the entire software stack, from the kernel to the application, is compiled, optimized, and linked together. However, in certain scenarios, such as code injection for network packet analysis or adding custom drivers, it is necessary to extend the kernel as needed. The traditional approach of modifying and recompiling the kernel source code can be time-consuming and error-prone. Martin Kröning, Stefan Lankes, Jonathan Klimt, Antonello Monti |
PLOS@SOSP | 2 |
| 2024 | HEXO: Offloading Long-Running Compute- and Memory-Intensive Workloads on Low-Cost, Low-Power Embedded SystemsabstractOS-capable embedded systems exhibiting a very low power consumption are available at an extremely low price point. It makes them highly compelling in a datacenter context. We show that sharing long-running, compute-intensive datacenter workloads between a server machine and one or a few connected embedded boards of negligible cost and power consumption can yield significant performance and energy benefits. Our approach, named Heterogeneous EXecution Offloading (HEXO), selectively offloads Virtual Machines (VMs) from server-class machines to embedded boards. Our design tackles several challenges. We address the Instruction Set Architecture (ISA) difference between typical servers (x86) and embedded systems (ARM) through hypervisor and guest OS-level support for heterogeneous-ISA runtime VM migration. We cope with the low amount of resources in embedded systems by using lightweight VMs – unikernels – and by using the server's free RAM as remote memory for embedded boards through a transparent lightweight memory disaggregation mechanism for heterogeneous server-embedded clusters, called Netswap. VMs are offloaded based on an estimation of the slowdown expected from running on a given board. We build a prototype of HEXO and demonstrate significant increases in throughput (up to 67%) and energy efficiency (up to 56%) using benchmarks representative of compute-intensive long-running workloads. Pierre Olivier, A. K. M. Fazla Mehrab, Sandeep Errabelly, Stefan Lankes, Mohamed Lamine Karaoui, Robert Lyerly, Sang-Hoon Kim, Antonio Barbalace, Binoy Ravindran |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | On the Challenge of Sound Code for Operating SystemsabstractThe memory-safe systems programming language Rust is gaining more and more attention in the operating system development communities, as it provides memory safety without sacrificing performance or control. However, these safety guarantees only apply to the safe subset of Rust, while bare-metal programming requires some parts of the program to be written in unsafe Rust. Writing abstractions for these parts of the software that are sound, meaning that they guarantee the absence of undefined behavior and thus uphold the invariants of safe Rust, can be challenging. Producing sound code, however, is essential to avoid breakage when the code is used in new ways or the compiler behavior changes. Jonathan Klimt, Martin Kröning, Stefan Lankes, Antonello Monti |
PLOS@SOSP | 3 |
| 2022 | Cricket: A virtualization layer for distributed execution of CUDA applications with checkpoint/restart supportabstractAbstract In high‐performance computing and cloud computing the introduction of heterogeneous computing resources, such as GPU accelerator have led to a dramatic increase in performance and efficiency. While the benefits of virtualization features in these environments are well researched, GPUs do not offer virtualization support that enables fine‐grained control, increased flexibility, and fault tolerance. In this article, we present Cricket: A transparent and low‐overhead solution to GPU virtualization that enables future research into other virtualization techniques, due to its open‐source nature. Cricket supports remote execution and checkpoint/restart of CUDA applications. Both features enable the distribution of GPU tasks dynamically and flexibly across computing nodes and the multitenant usage of GPU resources, thereby improving flexibility and utilization for high‐performance and cloud computing. Niklas Eiling, Jonas Baude, Stefan Lankes, Antonello Monti |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | A Syscall-Level Binary-Compatible UnikernelabstractUnikernels are minimal single-purpose virtual machines. They are highly popular in the research domain due to the benefits they provide. A barrier to their widespread adoption is the difficulty/impossibility to port existing applications to current unikernels. HermiTux is the first unikernel providing system call-level binary compatibility with Linux applications. It is composed of a hypervisor and a lightweight kernel layer emulating the load- and runtime Linux ABI. HermiTux relieves application developers from the burden of porting software, while providing unikernel benefits such as security through hardware-assisted virtualized isolation, swift boot time, and low disk/memory footprint. Fast system calls and kernel modularity are enabled through binary rewriting and analysis techniques, as well as shared library substitution. HermiTuxs design principles are architecture-independent and we present a prototype on both the x86-64 and ARM aarch64 ISAs, targeting various cloud as well as edge/embedded deployments. We demonstrate HermiTuxs compatibility over a range of native C/C++/Fortran/Python Linux applications. We also show that it offers a similar degree of lightweightness compared to other unikernels, and that it performs similarly to Linux in many cases: its performance overhead averages 3% in memory- and compute-bound scenarios, and its I/O performance is acceptable. Pierre Olivier, Hugo Lefeuvre, Daniel Chiba, Stefan Lankes, Changwoo Min, Binoy Ravindran |
IEEE Trans. Computers | 4 |
| 2020 | Scaling Shared Memory Multiprocessing Applications in Non-cache-coherent DomainsabstractDue to the slowdown of Moore's Law, systems designers have begun integrating non-cache-coherent heterogeneous computing elements in order to continue scaling performance. Programming such systems has traditionally been difficult - developers were forced to use programming models that exposed multiple memory regions, requiring developers to manually maintain memory consistency. Previous works proposed distributed shared memory (DSM) as a way to achieve high programmability in such systems. However, past DSM systems were plagued by low-bandwidth networking and utilized complex memory consistency protocols, which limited their adoption. Recently, new networking technologies have begun to change the assumptions about which components are bottlenecks in the system. Additionally, many popular shared-memory programming models utilize memory consistency semantics similar to those proposed for DSM, leading to widespread adoption in mainstream programming. Ho-Ren Chuang, Robert Lyerly, Stefan Lankes, Binoy Ravindran |
SYSTOR | 3 |
| 2020 | Intra-unikernel isolation with Intel memory protection keysabstractUnikernels are minimal, single-purpose virtual machines. This new operating system model promises numerous benefits within many application domains in terms of lightweightness, performance, and security. Although the isolation between unikernels is generally recognized as strong, there is no isolation within a unikernel itself. This is due to the use of a single, unprotected address space, a basic principle of unikernels that provide their lightweightness and performance benefits. In this paper, we propose a new design that brings memory isolation inside a unikernel instance while keeping a single address space. We leverage Intel's Memory Protection Key to do so without impacting the lightweightness and performance benefits of unikernels. We implement our isolation scheme within an existing unikernel written in Rust and use it to provide isolation between trusted and untrusted components: we isolate (1) safe kernel code from unsafe kernel code and (2) kernel code from user code. Evaluation shows that our system provides such isolation with very low performance overhead. Notably, the unikernel with our isolation exhibits only 0.6% slowdown on a set of macro-benchmarks. Mincheol Sung, Pierre Olivier, Stefan Lankes, Binoy Ravindran |
VEE | 3 |
| 2019 | HEXO: Offloading HPC Compute-Intensive Workloads on Low-Cost, Low-Power Embedded SystemsabstractOS-capable embedded systems exhibiting a very low power consumption are available at an extremely low price point. It makes them highly compelling in a datacenter context. In this paper we show that sharing long-running, compute-intensive datacenter HPC workloads between a server machine and one or a few connected embedded boards of negligible cost and power consumption can bring significant benefits in terms of consolidation. Our approach, named Heterogeneous EXecution Offloading (HEXO), selectively offloads Virtual Machines (VMs) from server class machines to embedded boards. Our design tackles several challenges. We address the Instruction Set Architecture (ISA) difference between typical servers (x86) and embedded systems (ARM) through hypervisor and guest OS-level support for heterogeneous-ISA runtime VM migration. We cope with the low amount of resources in embedded systems by using lightweight VMs: unikernels. VMs are offloaded based on an estimation of the slowdown expected from running on a given board. We build a prototype of HEXO and demonstrate significant increase in throughput (up to 67%) and energy efficiency (up to 56%) over a set of macro-benchmarks running datacenter compute-intensive jobs. Pierre Olivier, A. K. M. Fazla Mehrab, Stefan Lankes, Mohamed Lamine Karaoui, Robert Lyerly, Binoy Ravindran |
HPDC | 3 |
| 2019 | Exploring Rust for Unikernel DevelopmentabstractSystem-level development has been dominated by programming languages like C/C++ for decades. These languages are inherently unsafe, error-prone, and a major reason for vulnerabilities. High-level programming languages with a secure memory model and strong type system are able to improve the quality of the system software. In this paper, we explore the programming language Rust for kernel development and present RustyHermit, which is a unikernel completely written in Rust without any C/C++. We show that the support for RustyHermit can be transparently integratable in the Rust toolchain and common Rust applications are build-able on top of RustyHermit. Previously, we developed the C-based unikernel HermitCore with a similar design to RustyHermit and we are able to compare both kernels. We show that the performance of both kernels is similar and only ~3.27 % of RustyHermit relies on unsafe code, that cannot be checked by the compiler in detail. Stefan Lankes, Jens Breitbart, Simon Pickartz |
PLOS@SOSP | 1 |
| 2019 | A binary-compatible unikernelabstractUnikernels are minimal single-purpose virtual machines. They are highly popular in the research domain due to the benefits they provide. A barrier to their widespread adoption is the difficulty/impossibility to port existing applications to current unikernels. HermiTux is the first unikernel providing binary-compatibility with Linux applications. It is composed of a hypervisor and lightweight kernel layer emulating OS interfaces at load- and runtime in accordance with the Linux ABI. HermiTux relieves application developers from the burden of porting software, while providing unikernel benefits such as security through hardware-assisted virtualized isolation, swift boot time, and low disk/memory footprint. Fast system calls and kernel modularity are enabled through binary rewriting and analysis techniques, as well as shared library substitution. Compared to other unikernels, HermiTux boots faster and has a lower memory/disk footprint. We demonstrate that over a range of native C/C++/Fortran/Python Linux applications, HermiTux performs similarly to Linux in most cases: its performance overhead averages 3% in memory- and compute-bound scenarios. Pierre Olivier, Daniel Chiba, Stefan Lankes, Changwoo Min, Binoy Ravindran |
VEE | 3 |
| 2018 | Prospects and challenges of virtual machine migration in HPCabstractSummary The continuous growth of supercomputers is accompanied by increased complexity of the intra‐node level and the interconnection topology. Consequently, the whole software stack ranging from the system software to the applications has to evolve, eg, by means of fault tolerance and support for the rising intra‐node parallelism. Migration techniques are one means to address these challenges. On the one hand, they facilitate the maintenance process by enabling the evacuation of individual nodes during runtime, ie, the implementation of fault avoidance. On the other hand, they enable dynamic load balancing for an improvement of the system's efficiency. However, these prospects come along with certain challenges. On the process level, migration mechanisms have to resolve so‐called residual dependencies to the source node, eg, the communication hardware. On the job level, migrations affect the communication topology, which should be addressed by the communication stack, ie, the optimal communication path between a pair of processes might change after a migration. In this article, we explore migration mechanisms for HPC and discuss their prospects as well as the challenges. Furthermore, we present solutions enabling their efficient usage in this domain. Finally, we evaluate our prototype co‐scheduler leveraging migration for workload optimization. Simon Pickartz, Carsten Clauss, Jens Breitbart, Stefan Lankes, Antonello Monti |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Revisiting locality-awareness in view of dynamically changing topologiesabstractAs a general rule, when writing parallel applications according to the MPI standard, the programmer does not need to worry about the underlying hardware topology. This is because the MPI standard intentionally hides the actual hardware topology from the application programmer for the seizure of portability, while at the same time burdening the MPI implementation to handle hardware-related peculiarities as optimal as possible. So, for instance, with the emergence of SMP systems, locality-awareness in terms of the recognition of accelerated node-internal communication found its way into all major MPI libraries in the early 2000s. However, the actually implemented degree of such a locality-awareness can vary: From the simple usage of point-to-point communication over shared-memory, via the smart adaptation of collective communication patterns, through to the exploitation of direct accessible address spaces for one-sided communication. Until now, all these locality-related optimizations basically assume a static hierarchical topology in the course of a parallel program. In contrast, this article strives for a discussion of how dynamically changing topologies, as they may result from process migrations, can be considered during runtime for locality-awareness. In doing so, the article focuses on collective communication, but also discusses the challenges for point-to-point and one-sided communication. Simon Pickartz, Carsten Clauss, Stefan Lankes, Antonello Monti |
Parallel Comput. | 3 |
| 2018 | Zeroing memory deallocator to reduce checkpoint sizes in virtualized HPC environments
Ramy Gad, Simon Pickartz, Tim Süß, Lars Nagel 0001, Stefan Lankes, Antonello Monti, André Brinkmann |
J. Supercomput. | 5 |
| 2017 | Dynamic Co-Scheduling Driven by Main Memory Bandwidth UtilizationabstractMost applications running on supercomputers achieve only a fraction of a system's peak performance. It has been demonstrated that the co-scheduling of applications can improve the overall system utilization. However, following this approach, applications need to fulfill certain criteria such that the mutual slowdown is kept at a minimum. In this paper, we present an HPC scheduler that applies co-scheduling and utilizes virtual machine migration for a re-orchestration of applications at runtime based on their main memory bandwidth requirements. Given a job queue consisting of main memory-bound applications and compute-bound applications, we can see a throughput increase of up to 35% while at the same time reducing energy consumption by around 30%. Jens Breitbart, Simon Pickartz, Stefan Lankes, Josef Weidendorfer, Antonello Monti |
CLUSTER | 3 |
| 2017 | Data Mining-Based Analysis of HPC Center OperationsabstractSize and complexity of contemporary High Performance Computing (HPC) systems increases permanently. While the reliability of a single component and compute node is high, the huge amount of components comprising these systems results in the fact that defects happen regularly. This drives the need to manage failure situations. Common issues are component failures or node soft lock-ups that typically lead to crashes of the user jobs that are scheduled on the affected node, and may cause undesired downtime. One approach to mitigate the impact of such problems is to predict node failures with a sufficient lead time in order to take proactive measures. However, accurate prediction is a challenging task.The literature describes several approaches that focus on gathering and analyzing system event logs in order to create prediction models. In this paper, we present a different approach by using descriptive statistics and supervised machine learning to create a prediction model from monitoring data. Our approach is based on the assumption, that features of a certain time frame before a critical event (i. e., a failure or soft lock-up) can serve as an indicator. Consequently, our model is trained with monitoring data from critical and healthy time frames. The evaluation with standard monitoring data collected from the HPC systems at RWTH Aachen University shows that our classifier is able to locate potentially failing nodes with a 10-fold cross precision of 98% and recall of 91 %. Jannis Klinkenberg, Christian Terboven, Stefan Lankes, Matthias S. Müller |
CLUSTER | 3 |
| 2016 | VarySched: A Framework for Variable Scheduling in Heterogeneous EnvironmentsabstractDespite many efforts to better utilize the potential of GPUs and CPUs, it is far from being fully exploited. Although many tasks can be easily sped up by using accelerators, most of the existing schedulers are not flexible enough to really optimize the resource usage of the complete system. The main reasons are (i) that each processing unit requires a specific program code and that this code is often not provided for every task, and (ii) that schedulers may follow the run-until-completion model and, hence, disallow resource changes during runtime. In this paper, we present VarySched, a configurable task scheduler framework tailored to efficiently utilize all available computing resources in a system. VarySched allows a more fine-grained task-to-resource placement which is even further enhanced by allowing the tasks to migrate to another resource during their runtime. In addition, VarySched can manage multiple scheduling strategies - optimizing, for instance, throughput or energy efficiency - and switch between them at any time. Tim Süß, Nils Döring, Ramy Gad, Lars Nagel 0001, André Brinkmann, Dustin Feld, Thomas Soddemann, Stefan Lankes |
CLUSTER | 8 |
| 2015 | New system software for parallel programming models on the Intel SCC many-core processorabstractSummary Since the beginning of the multicore era, parallel processing has become prevalent across the board. On a traditional multicore system, a single operating system manages all cores and schedules threads and processes among them, inherently supported by hardware‐implemented cache coherence protocols. However, a further growth of the number of cores per system implies an increasing chip complexity, especially with respect to the cache coherence protocols. Therefore, a very attractive alternative for future many‐core systems is to waive the hardware‐based cache coherency and to introduce a software‐oriented message‐passing based architecture instead: a so‐called Cluster‐on‐Chip architecture. Intel's Single‐chip Cloud Computer (SCC), a many‐core research processor with 48 non‐coherent memory‐coupled cores, is a very recent example for such a cluster‐on‐chip architecture. The SCC can be configured to run one operating system instance per core by partitioning the shared main memory in a strict manner. However, it is also possible to access the shared main memory in an unsplit and concurrent manner, provided that either the caches are disabled or the cache coherency is then ensured by software. In this article, we detail our experiences gained while developing low‐level software for message‐passing and shared‐memory programming on the SCC. We present an SCC‐customized MPI library (called SCC‐MPICH) as well as a shared virtual memory system (called MetalSVM) for the SCC. In doing so, we evaluate the potential of both programming models and we show how these models can be improved especially with respect to the SCC's many‐core architecture. Copyright © 2013 John Wiley & Sons, Ltd. Carsten Clauss, Stefan Lankes, Pablo Reble, Thomas Bemmerl |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Using a Multitasking GPU Environment for Content-Based Similarity Measures of Big Data
Ayman Tarakji, Marwan Hassani, Stefan Lankes, Thomas Seidl 0001 |
ICCSA (5) | 3 |
| 2012 | Towards a Multicore Communications API Implementation (MCAPI) for the Intel Single-Chip Cloud Computer (SCC)abstractIn this paper, we present a prototype implementation of the Multicore Communications API (MCAPI) for the Intel Single-Chip Cloud Computer (SCC). The SCC is a 48 core concept vehicle for future many-core systems that exhibit message-passing oriented architectures. The MCAPI specification, recently developed by the Multicore Association, resembles a lightweight interface for message-passing in today's multicore systems. The presented prototype implementation should be used to evaluate the MCAPI's capability and feasibility for its employment also in future many-core systems. Carsten Clauss, Simon Pickartz, Stefan Lankes, Thomas Bemmerl |
ISPDC | 3 |
| 2011 | Performance Tuning of SCC-MPICH by Means of the Proposed MPI-3.0 Tool Interface
Carsten Clauss, Stefan Lankes, Thomas Bemmerl |
EuroMPI | 2 |
| 2010 | Use Case Evaluation of the Proposed MPIT Configuration and Performance Interface
Carsten Clauss, Stefan Lankes, Thomas Bemmerl |
EuroMPI | 2 |
| 2009 | MPIXternal: A library for a portable adjustment of parallel MPI applications to heterogeneous environmentsabstractNowadays, common systems in the area of high performance computing exhibit highly hierarchical architectures. As a result, achieving satisfactory application performance demands an adaptation of the respective parallel algorithm to such systems. This, in turn, requires knowledge about the actual hardware structure even at the application level. However, the prevalent Message Passing Interface (MPI) standard (at least in its current version 2.1) intentionally hides heterogeneity from the application programmer in order to assure portability. In this paper, we introduce the MPIXternal library which tries to circumvent this obvious semantic gap within the current MPI standard. For this purpose, the library offers the programmer additional features that should help to adapt applications to today's hierarchical systems in a convenient and portable way. Carsten Clauss, Stefan Lankes, Thomas Bemmerl |
IPDPS | 2 |
| 2008 | Design and Implementation of a Service-integrated Session Layer for Efficient Message Passing in Grid Computing EnvironmentsabstractWhen running large parallel applications with demands for resources that exceed the capacity the local computing site offers, the deployment in a distributed grid environment may help to satisfy these demands. However, since such an environment is a heterogeneous system by nature, there are some drawbacks that, if not taken into account, are limiting its applicability.First of all, one has to apply a meta-computing or Grid-enabled message-passing library in order to have the ability to route messages to remote sites as well as still being able to exploit fast site-local network facilities.Then, because the inter-site communication usually constitutes the system's bottleneck, appropriate quality of service parameters should be provided and policed for those connections during the application's execution. And finally, the parallel runtime environment of the distributed application should offer service interfaces in order to enable its interaction with Grid middleware. In this paper, we present a new library called ISI whose functionalities meet those requirements in terms of a session layer to be integrated into grid-enabled message-passing implementations. Carsten Clauss, Stefan Lankes, Thomas Bemmerl |
ISPDC | 2 |
| 2007 | A Fair Benchmark for Evaluating the Latent Potential of Heterogeneous Coupled ClustersabstractCoupled clusters usually exhibit a heterogeneous but also hierarchical structure in terms of communication and computation. Therefore, it is inevitable to adapt parallel applications to such systems in order to gain reasonable performance results. Moreover, also regular benchmark tools are not capable of exposing the latent potential of such coupled cluster systems. Though without adapted (or better self-adapting) benchmark tools for such systems, it is almost not possible to forecast the scalability of well-adapted applications and one is not able to compare the possibly achievable performance in an application independent manner. In this paper we present such a fair, self-adapting and meaningful benchmark tool for heterogeneous coupled cluster systems, following the MPI standard. Carsten Clauss, Stephan Gsell, Stefan Lankes, Thomas Bemmerl |
ISPDC | 3 |
| 2006 | Parallelisation of a simulation tool for casting and solidification processes on Windows platformsabstractSince the beginning of computational engineering, the numerical simulation of physical processes is an essential element in the area of high performance computing. Thus, also the domain of metal foundry demands the computational simulation of casting and solidification processes. A popular software tool for this purpose has been developed by the RWP GmbH in Roetgen, Germany. This tool, named WinCast, is a complete software suite, which contains modules for pre-, main- and post-processing of simulation data sets. A core module of WinCast is TFB, which determines the chronological temperature distribution of a casting process based on a finite-element-method and a Gauss-Seidel solver. With the increasing demand for even higher precision of the simulation results on one hand, and a growing need for even larger data sets on the other hand, the parallelisation of this module became inevitable. In this paper, we present our work accomplished to parallelise the solving algorithm of this module. We have chosen an MPI based master-slave approach for compute clusters by using a self-developed MPI library for Windows platforms Carsten Clauss, Silke Schuch, Rainer Finocchiaro, Stefan Lankes, Thomas Bemmerl |
IPDPS | 4 |
| 2005 | Design and performance of a CAN-based connection-oriented protocol for Real-Time CORBA
Stefan Lankes, Andreas Jabs, Thomas Bemmerl |
J. Syst. Softw. | 1 |
| 2004 | Design of a Real-Time CORBA Event Service Customised for the CAN BusabstractSummary form only given. Real-time CORBA and minimum CORBA are the foundations that many so called distributed real-time embedded (DRE) systems are built upon. These specifications describe middleware suitable for connecting different parts of a complex embedded system. Efficient group based communication in such a system can be achieved by using the event service. We focus on the design of such a service complying with the OMG event service standard. It is optimised for the CAN bus, a widely used interconnect, where real-time characteristics are a requirement. A new protocol for the efficient distribution of events in a CAN-based distributed control system is presented, a protocol which is tailored to the CAN bus and produces very low overhead by utilising CAN-specific features. Rainer Finocchiaro, Stefan Lankes, Andreas Jabs |
IPDPS | 2 |
| 2001 | Design and Implementation of a SCI-Based Real-Time CORBAabstractThe Real-Time CORBA and minimumCORBA specifications in the forthcoming CORBA 3.0 standard are important steps towards defining standard-based middleware which can satisfy real time requirements in an embedded system. The article describes these new specifications and an implementation called ROFES. ROFES supports different network architectures, for example the Scalable Coherent Interface (SCI). Furthermore, the article examines the SCI-network and whether it possesses real time characteristics. Stefan Lankes, Michael Pfeiffer 0002, Thomas Bemmerl |
ISORC | 1 |