Cornelia Wulf

dblp:279/9278 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-6100-955XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCISSORS: System Level Error Detection for Enabling Near-Threshold Operating Systolic Arrays
abstract
Since dynamic power has a quadratic relationship with voltage, reducing voltage is an effective way to lower power consumption in digital circuits. However, maintaining stable operation at lower voltages is challenging due to increased sensitivity to Process, Voltage, and Temperature (PVT) variations, making it difficult to determine optimal operating points using Static Timing Analysis (STA). While circuit-and device-level solutions like Timing Error Detection (TED) systems can enable lower voltage operation, they introduce significant overhead and design complexity. In this paper, we integrate an Algorithm-Based Fault Detection (ABFT) method into the structure of systolic arrays to capture timing errors when voltage is scaled down, ensuring safe and optimized low-voltage operation. Our proposed approach, SCISSORS, demonstrates how extra voltage margins in systolic arrays used for matrix arithmetic can be trimmed by integrating a simple algorithmic technique into the structure of the array. This solution not only detects errors in the accelerator but also those caused by voltage reduction in on-chip memory and auxiliary circuits. It is fully implementable through HDL without requiring transistor-or circuit-level modifications to the netlist. Implementation on a Zynq System-on-Chip (SoC) shows that SCISSORS introduces only a tolerable overhead of 11% and 8% for 32×32 and 64×64 systolic arrays, respectively, while achieving nearly a 2× improvement in energy efficiency. Experimental results further demonstrate that SCISSORS adaptively adjusts voltage in response to the voltage-temperature coupling behavior of digital circuits at runtime, specifically addressing Inverse Temperature Dependence (ITD).
Ensieh Aliagha, Mehdi Safarpour, Cornelia Wulf, Olli Silvén, Diana Göhringer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Hardware-level Access Control and Scheduling of Shared Hardware Accelerators
abstract
With the trend to consolidate hardware on a single platform, FPGA virtualization plays an increasingly important role in the embedded domain. FPGA virtualization allows multiple software tasks or even guest operating systems to share recon- figurable resources. However, state-of-the-art approaches assign each hardware accelerator to a single software task for a fixed duration. This becomes a problem when the number of hardware accelerators required by software tasks concurrently exceeds the FPGA area. If several software tasks request to accelerate the same functionality, accelerators can be shared. Embedded reconfigurable systems face the challenge of a uniform address space. When several tasks use a memory-mapped communication interface that allows to directly access the accelerator's address space, access control and the protection from unauthorized access must be ensured. Existing software-based approaches lead to high latencies. Thus, we propose a hardware-level scheduler that schedules hardware tasks in spatial and temporal respect. The allocation to a hardware accelerator is combined with the assignment of access rights. Any unauthorized access leads to a page fault. When hardware tasks share an accelerator, they are scheduled according to the Earliest Deadline First (EDF) policy. Buffers ensure data isolation. Compared to hardware task scheduling in software, a performance increase of 7.02 times is reached.
Cornelia Wulf, Sergio A. Pertuz 0001, Diana Göhringer
DSD1
2024 Energy-Aware Synchronization of Hardware Tasks in Virtualized Embedded Systems
abstract
Dynamic Voltage and Frequency Scaling (DVFS) is an effective means to reduce the energy dissipation of digital designs. While on most commodity FPGAs, memory and processor have separately controlled voltages, the programmable logic section relies on a single voltage rail and thus imposes the same voltage for all hardware accelerators that operate concurrently. Finding time slots eligible for voltage scaling gets difficult in virtualized systems, where the FPGA is shared by tasks executed in multiple guest operating systems. The situation gets even more complicated, when error-tolerant tasks are considered that allow the voltage to be reduced below its nominal value, which could provoke a certain rate of faulty hardware accelerator runs. As a solution, we propose a strategy that synchronizes concurrently executed periodic hardware tasks under consideration of their reliability as well as their real-time requirements so that the supply voltage is controlled accordingly. The proposed strategy can be combined with further mechanisms for saving energy. Our run-time module performs clock gating and adjusts the voltage to the requirements of aperiodic tasks. For fault-tolerant tasks, we monitor the error rate using Algorithm Based Fault Tolerance (ABFT) that can detect and characterize errors with an accuracy close to $100 \%$. Compared to a strategy that scales voltage without synchronizing hardware tasks, we achieve in the best case a power saving by $29.4 \%$ and an average saving by $7 \%$.
Cornelia Wulf, Gökhan Akgün, Mehdi Safarpour, Anastacia Grishchenko, Diana Göhringer
FPL1
2023 EuFRATE: European FPGA Radiation-hardened Architecture for Telecommunications
abstract
The EuFRATE project aims to research, develop and test radiation-hardening methods for telecommunication payloads deployed for Geostationary-Earth Orbit (GEO) using Commercial-Off- The-Shelf Field Programmable Gate Arrays (FPGAs). This project is conducted by Argotec Group (Italy) with the collaboration of two partners: Politecnico di Torino (Italy) and Technische Universität Dresden (Germany). The idea of the project focuses on high-performance telecommunication algorithms and the design and implementation strategies for connecting an FPGA device into a robust and efficient cluster of multi-FPGA systems. The radiation-hardening techniques currently under development are addressing both device and cluster levels, with redundant datapaths on multiple devices, comparing the results and isolating fatal errors. This paper introduces the current state of the project's hardware design description, the composition of the FPGA cluster node, the proposed cluster topology, and the radiation hardening techniques. Intermediate stage experimental results of the FPGA communication layer performance and fault detection techniques are presented. Finally, a wide summary of the project's impact on the scientific community is provided.1
Ludovica Bozzoli, Antonino Catanese, Emilio Fazzoletto, Eugenio Scarpa, Diana Göhringer, Sergio A. Pertuz 0001, Lester Kalms, Cornelia Wulf, Najdet Charaf, Luca Sterpone, Sarah Azimi, Daniele Rizzieri, Salvatore Gabriele La Greca, David Merodio Codinachs
DATE8
2023 Virtualization of Hardware Accelerators in a Network-on-Chip
abstract
Networks-on-Chip (NoCs) are beneficial for reconfigurable systems that require a high degree of parallel and scalable communication. NoCs are reusable as hardware accelerators can be exchanged via dynamic partial reconfiguration. Nevertheless, NoCs are not conceptualized for the use in a virtualized environment where applications from multiple virtual machines have to share reconfigurable resources. Many state-of-the-art works assign hardware accelerators exclusively to a single virtual machine, which limits the number of processed hardware tasks and leads to underutilization of FPGA area. Therefore, we provide a NoC virtualization layer that allows the execution of several pipelined hardware tasks agnostic of the location of the required hardware accelerators. The allocation of tasks to processing elements can be adapted to dynamically changing requirements, while unauthorized access is prohibited. Further, we provide a scheduler that schedules hardware tasks in spatial and temporal respect to processing elements in the NoC. The proposed heuristic considers task priorities, a possible reuse of accelerators and hop counts. In over-load conditions, the tasks with the lowest priorities are postponed. Our virtualization layer increases the number of tasks processed by 22.6% compared to an approach that grants exclusive access.
Cornelia Wulf, Julian Haase, Matthias Nickel, Diana Göhringer
DSD1
2022 Scheduling of Hardware Tasks in Reconfigurable Mixed-Criticality Systems
abstract
FPGA virtualization allows the shared usage of an FPGA by several operating systems with different criticality levels. To avoid mutual interference, most state-of-the-art systems strictly isolate subsystems in spatial respect at the expense of lower resource utilization. We present an allocation and scheduling strategy for hardware tasks that improves resource utilization while respecting different real-time levels (hard, soft, and no real-time) of guest operating systems. To not jeopardize deadlines, Dynamic Partial Reconfiguration (DPR) latencies are reduced by reusing, prefetching and reserving of hardware accelerators. Compared with an existing scheduler for hardware tasks, we could increase the resource usage by 156% while deadline misses were reduced by 6%.
Cornelia Wulf, Najdet Charaf, Diana Göhringer
FCCM1
2022 Virtualization of Reconfigurable Mixed-Criticality Systems
abstract
The increasing complexity of reconfigurable embedded systems often requires the integration of multiple applications with potentially different levels of criticality on the same hardware platform. As the deployment scales, there is a need for resource management, isolation, and performance that makes FPGA virtualization techniques a key consideration. FPGA virtualization enables multiple guest operating systems to run with different requirements, such as real-time, safety, or security. Most state-of-the-art systems incorporate mechanisms to strictly isolate subsystems in spatial respect at the expense of lower resource utilization. In this work, we present L4ReC, a microkernel-based virtualization layer that enables the sharing of reconfigurable resources among multiple virtual machines. The mapping and scheduling strategy for hardware threads considers not only deadlines, but also the real-time levels of guest operating systems. A POSIX thread-based interface facilitates the access to hardware accelerators. Compared with an existing scheduler for hardware threads, the average utilization factor - indicating the FPGA resource usage - is 1,9 times higher when threads are mapped and scheduled with L4ReC. Deadline misses are reduced by 3%.
Cornelia Wulf, Najdet Charaf, Diana Göhringer
FPL1
2022 Virtualization of Embedded Reconfigurable Systems
abstract
With the trend to consolidate multiple systems onto the same hardware platform, which can be observed for example in the automotive industry, virtualization of embedded systems becomes increasingly important. Often small and efficient real time operating systems (RTOS) run besides general purpose operating systems (GPOS) with a convenient, high level application interface. When virtualizing embedded reconfigurable systems, FPGA characteristics have to be considered like limited FPGA area and high reconfiguration latencies. We present the FPGA virtualization layer L4ReC that enables the shared usage of reconfigurable resources by several guest operating systems under consideration of the constraints given by embedded reconfigurable systems. First results target isolation, energy efficiency, and FPGA resource management considering special requirements of guest operating systems.
Cornelia Wulf, Diana Göhringer
FPL1
2021 A Survey on Hypervisor-based Virtualization of Embedded Reconfigurable Systems
abstract
The increase of size, capabilities, and speed of FPGAs enables the shared usage of reconfigurable resources by multiple applications and even operating systems. While research on FPGA virtualization in HPC-datacenters and cloud is already well advanced, it is a rather new concept for embedded systems. The necessity for FPGA virtualization of embedded systems results from the trend to integrate multiple environments into the same hardware platform. As multiple guest operating systems with different requirements, e.g., regarding real-time, security, safety, or reliability share the same resources, the focus of research lies on isolation under the constraint of having minimal impact on the overall system. Drivers for this development are, e.g., computation intensive AI-based applications in the automotive or medical field, embedded 5G edge computing systems, or the consolidation of electronic control units (ECUs) on a centralized MPSoC with the goal to increase reliability by reducing complexity. This survey outlines key concepts of hypervisor-based virtualization of embedded reconfigurable systems. Hypervisor approaches are compared and classified into FPGA-based hypervisors, MPSoC-based hypervisors and hypervisors for distributed embedded reconfigurable systems. Strong points and limitations are pointed out and future trends for virtualization of embedded reconfigurable systems are identified.
Cornelia Wulf, Michael Willig, Diana Göhringer
FPL1