VLDB 2026 Research / reviewers in the wild / expert
Stefan Wallentowitz
dblp:34/1253
· DBLP profile ↗
11ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-3182-4929ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | parSAT: Parallel Solving of Floating-Point SatisfiabilityabstractSatisfiability-based verification techniques, leveraging modern Boolean satisfiability (SAT) and Satisfiability Modulo Theories (SMT) solvers, have demonstrated efficacy in addressing practical problem instances within program analysis.However, current SMT solver implementations often encounter limitations when addressing non-linear arithmetic problems, particularly those involving floating point (FP) operations.This poses a significant challenge for safety critical applications, where accurate and reliable calculations based on FP numbers and elementary mathematical functions are essential.This paper shows how an alternative formulation of the satisfiability problem for FP calculations allows for exploiting parallelism for FP constraint solving.By combining global optimization approaches with parallel execution on modern multi-core CPUs, we construct a portfolio-based semidecision procedure specifically tailored to handle FP arithmetic.We demonstrate the potential of this approach to complement conventional methods through the evaluation of various benchmarks. Markus Krahl, Matthias Güdemann, Stefan Wallentowitz |
J. Log. Algebraic Methods Program. | 3 |
| 2025 | Benchmarking WebAssembly for Embedded SystemsabstractWebAssembly is a modern, low-level virtual machine with designed for improved application performance in web browsers. Recently, WebAssembly gained interest for its use outside the web, for example as a replacement for serverless container runtimes. A number of non-web WebAssembly implementations are actively supported, some of which target microcontrollers, IoT devices, and embedded systems. Such hardware platforms have strict resource constraints which may render the usage of WebAssembly impossible or too costly, for example, due to its performance overhead and memory requirements. However, it is currently unclear what performance to expect of WebAssembly on low-resource microcontrollers compared with machine code and alternative application virtual machines. To answer this question, we evaluated the processing overhead and memory characteristics of WebAssembly application virtual machines on microcontrollers, and compared it to native execution, and the established application virtual machines: MicroPython and Lua. Furthermore, we analyzed the feature-set and architecture of the WebAssembly implementations in more detail, and measured the performance impact different runtime features have. We found that WebAssembly, despite its high extensibility and versatility in supported source languages, application paradigms, and target hardware, delivers very competitive performance. We conclude that WebAssembly can find wider industry usage for embedded systems and could replace other more costly or less flexible virtualization techniques, such as Java. Konrad Moron, Stefan Wallentowitz |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Fast Behavioural RTL Simulation of 10B Transistor SoC Designs with Metro-MpiabstractChips with tens of billions of transistors have become today's norm. These designs are straining our electronic design automation tools throughout the design process, requiring ever more computational resources. In many tools, parallelisation has improved both latency and throughput for the designer's benefit. However, tools largely remain restricted to a single machine and in the case of RTL simulation, we believe that this leaves much potential performance on the table. We introduce Metro-MPI to improve RTL simulation for modern 10 billion transistor-scale chips. Metro-MPI exploits the natural boundaries present in chip designs to partition RTL simulations and leverage High Performance Computing (HPC) techniques to extract parallelism. For chip designs that scale in size by exploiting latency-insensitive interfaces like networks-on-chip and AXI, Metro-MPI offers a new paradigm for RTL simulation scalability. Our implementation of Metro-MPI in Open-Piton+Ariane delivers 2.7 MIPS of RTL simulation throughput for the first time on a design with more than 10 billion transistors and 1,024 Linux-capable cores, opening new avenues for distributed RTL simulation of emerging system-on-chip designs. Compared to sequential and multithreaded RTL simulations of smaller designs, Metro-MPI achieves up to$135.98\times$and$9.29\times$speedups. Similarly, for a representative regression run, Metro-Mpireduces energy consumption by up to$2.53\times$and$2.91\times$. Guillem López-Paradís, Brian Li, Adrià Armejach, Stefan Wallentowitz, Miquel Moretó, Jonathan Balkind |
DATE | 4 |
| 2023 | AutoNLP: A System for Automated Market Research Using Natural Language Processing and Flow-based Programming
Florian Würmseer, Stefan Wallentowitz, Markus Friedrich 0001 |
I4CS | 2 |
| 2015 | A Hardware/Software Approach for Mitigating Performance Interference Effects in Virtualized Environments Using SR-IOVabstractSingle Root I/O Virtualization (SR-IOV) is an extension to the PCI Express (PCIe) standard that allows virtual machines (VMs) to directly access shared I/O devices without host involvement. This enabled SR-IOV to become the best-performing solution for virtual I/O to date, which lead to its commercial adoption, e.g., In the Amazon EC2. On the downside, a malicious VM can exploit the direct access to an SR-IOV device by flooding it with PCIe packets. This results in a congestion on the PCIe interconnect, which leads to performance interference effects between the malicious VM, concurrent VMs and even the host. In this paper, we present a hardware/software approach that detects and mitigates such Denial-of-Service (DoS) attacks. On the hardware side, we propose monitoring extensions within SR-IOV devices that distinguish legal device use from malicious device use by observing the rate of incoming PCIe transactions at VM granularity. Malicious VMs are reported to the host via interrupts. On the software side, performance interference effects can then be mitigated by dynamically adjusting the host's scheduling of the malicious VM or even shutting it down. We implement a prototype with a commercial off-the-shelf SR-IOV Ethernet controller and an FPGA board. On it, we demonstrate that appropriate scheduling of malicious VMs successfully mitigates interference effects for three cloud-relevant benchmarks. For example, Memcached is restored to 99.4% of baseline performance (compared to 61.8% without our extensions). In contrast to QoS features proposed in the PCIe 3.0 standard, our solution is more flexible. Additionally, it can be realized as an add-on to existing misuse detection hardware like the Intel Malicious Driver Detection (MDD). Andre Oliver Richter, Christian Herber, Stefan Wallentowitz, Thomas Wild, Andreas Herkersdorf |
CLOUD | 3 |
| 2015 | An Analytic Approach on End-to-End Packet Error Rate Estimation for Network-on-ChipabstractNetwork-on-Chip (NoC) are well-established for scalable on-chip communication, but technology generations of 22~nm and below, as well as aggressive voltage scaling to reduce NoC power consumption, introduce new variability challenges resulting in errors on wires and registers. Based on the probabilities of single bit flips, this paper focuses on the expected end-to-end packet error probabilities in NoC. We investigate the influence of individual bit error probabilities, the number of hops between communication partners, as well as the packet size. To evaluate these parameters, we propose an analytic approach which abstracts technology details of NoC data transport entities, such as links and buffers, and models each entity as a binary symmetric channel (BSC). The proposed probabilistic approach obtains equations for system-level NoC reliability estimates which allow an evaluation without the necessity to deploy time-consuming simulations. Michael Vonbun, Stefan Wallentowitz, Andreas Oeldemann, Andreas Herkersdorf |
DSD | 2 |
| 2014 | Distributed cooperative shared last-level caching in tiled multiprocessor system on chipabstractIn a shared-memory based tiled many-core system-on-chip architecture, memory accesses present a huge performance bottleneck in terms of access latency as well as bandwidth requirements. The best practice approach to address this issue is to provide a multi-level cache hierarchy and a suitable cache-coherency mechanism. This paper presents a method to increase the memory access performance in distributed-directory-coherency-protocol based tiled many-core systems. The proposed method introduces an alternate design for the system-wide shared last-level caches (LLC) placed between the memory and the node private caches (NPC). The proposed system-wide shared LLC layer is distributed over the entire network and it interacts with the home directories of specific cache lines. Results from simulating SPEC2000 benchmark applications executed on a SystemC model of the proposed design show a minimum performance improvement of 20-25% when compared to a model without the shared cache layer at the expense of an additional 2% of the total cache memory space (NPC + LLC memory). In addition, the proposed design shows a minimum 7-15% and an average 14-15% improvement in performance in comparison to centralized system-wide shared LLC of equivalent size and dynamic mapped distributed LLC of equivalent size respectively. Preethi P. Damodaran, Stefan Wallentowitz, Andreas Herkersdorf |
DATE | 2 |
| 2014 | Dependable task and communication migration in tiled manycore system-on-chipabstractPower densities and thermal hotspots are a major concern for the dependability of future multi-processor systemon- chip. They can lead to transient faults affecting the functionality in the short term and can cause permanent damage of a device. The dependability problem can be tackled on different layers such as technology hardening or application awareness. This work is based on an approach that addresses the issue for tile-based manycore system-on-chip on software and architecture layer. An agent-based system management employs task migration to react to thermal hotspots and pro-actively avoid them. The inter-task communication plays an important role as communication channels need to be migrated accordingly. The presented work focuses on the issue of communication migration and is based on the idea of handling it transparently to the task migration. Network-on-chip protection switching techniques have been introduced before and in this paper we evaluate the potential and bottlenecks of such methods in a realistic platform. Stefan Wallentowitz, Stefan Rosch, Thomas Wild, Andreas Herkersdorf, Volker Wenzel, Jörg Henkel |
FDL | 1 |
| 2012 | A framework for Open Tiled Manycore System-On-ChipabstractTiled manycore architectures have become dominant for the integration of tens or even a hundred processor cores on a chip. While commercial products are increasingly available, research on the hardware of such platforms and especially prototyping often rely on building such a platform from scratch or is bound to abstract simulation. In this paper we present the Open Tiled Manycore System-on-Chip (Op-TiMSoC) which is a library-based tool flow that helps generating a tiled manycore platform based on a library of open standard components. OpTiMSoC allows for research and prototyping of both shared memory and distributed memory platforms. It includes LISNoC which is a flexible NoC implementation. An OpTiMSoC system can easily be generated based on the publicly available repository and prototyped on an FPGA. As exemplary targets we evaluated the usage of different FPGA boards and an emulation platform. Stefan Wallentowitz, Andreas Lankes, Aurang Zaib, Thomas Wild, Andreas Herkersdorf |
FPL | 1 |
| 2012 | Benefits of selective packet discard in networks-on-chipabstractToday, Network on Chip concepts principally assume inherent lossless operation. Considering that future nanometer CMOS technologies will witness increased sensitivity to all forms of manufacturing and environmental variations (e.g., IR drop, soft errors due to radiation, transient temperature induced timing problems, device aging), efforts to cope with data corruption or packet loss will be unavoidable. Possible counter measures against packet loss are the extension of flits with ECC or the introduction of error detection with retransmission. We propose to make use of the perceived deficiency of packet loss as a feature. By selectively discarding stuck packets in the NoC, a proven practice in computer networks, all types of deadlocks can be resolved. This is especially advantageous for solving the problem of message-dependent deadlocks, which otherwise leads to high costs either in terms of throughput or chip area. Strict ordering, the most popular approach to this problem, results in a significant buffer overhead and a more complex router architecture. In addition, we will show that eliminating local network congestions by selectively discarding individual packets also can improve the effective throughput of the network. The end-to-end retransmission mechanism required for the reliable communication, then also provides lossless communication for the cores. Andreas Lankes, Thomas Wild, Stefan Wallentowitz, Andreas Herkersdorf |
ACM Trans. Archit. Code Optim. | 3 |
| 2006 | A SW performance estimation framework for early system-level-design using fine-grained instrumentationabstractThe increasing demands of high-performance in embedded applications under shortening time-to-market has prompted system architects in recent time to opt for multi-processor systems-on-chip (MP-SoCs) employing several programmable devices. The programmable cores provide a high amount of flexibility and reusability, and can be optimized to the requirements of the application to deliver high-performance as well. Since application software forms the basis of such designs, the need to tune the underlying SoC architecture for extracting maximum performance from the software code has become imperative. In this paper, we propose a framework that enables software development, verification and evaluation from the very beginning of MP-SoC design cycle. Unlike traditional SoC design flows where software design starts only after the initial SoC architecture is ready, our framework allows a co-development of the hardware and the software components in a tightly coupled loop where the hardware can be refined by considering the requirements of the software in a stepwise manner. The key element of this framework is the integration of a fine-grained software instrumentation tool into a system-level-design (SLD) environment to obtain accurate software performance and memory access statistics. The accuracy of such statistics is comparable to that obtained through instruction set simulation (ISS), while the execution speed of the instrumented software is almost an order of magnitude faster than ISS. Such a combined design approach assists system architects to optimize both the hardware and the software through fast exploration cycles, and can result in far shorter design cycles and high productivity. We demonstrate the generality and the efficiency of our methodology with two case studies selected from two most prominent and computationally intensive embedded application domains. Torsten Kempf, Kingshuk Karuri, Stefan Wallentowitz, Gerd Ascheid, Rainer Leupers, Heinrich Meyr |
DATE | 3 |