Felix Staudigl

dblp:308/0616 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-9673-3070ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 It's Getting Hot in Here: Hardware Security Implications of Thermal Crosstalk on ReRAMs
abstract
Emerging non-volatile memories (eNVM) promise to solve the imminent von Neumann bottleneck by enabling future computing systems to utilize the computing-in-memory (CIM) paradigm offering exceptional energy efficiency and performance advantages. As Moore's law becomes obsolete, CIM architectures are prominent candidates to push the boundaries of existing computing systems and usher in a new generation of computing models, such as neuromorphic systems. Furthermore, conventional systems face another significant problem in addition to the von Neumann bottleneck. Hardware security threats (e.g., Rowhammer) have gained momentum and can expose an entirely pristine attack surface for adversaries. These vulnerabilities distinguish themselves by being particularly challenging to patch because their origin lies in the rigid hardware layout. Unfortunately, neuromorphic systems are no exception. We presented NeuroHammer as one of the first unique hardware security attacks on eNVMs, enabling an attacker to intentionally flip bits in memristive crossbar arrays. This article extends our previous results by thoroughly examining the underlying concepts leading to the NeuroHammer attack. First, we investigate memory access patterns to gain insight into the tangible impact of NeuroHammer. Second, we extend our simulation methodology to accommodate transistor/one resistive (1T1R) crossbar structures and prove the prevalence of the NeuroHammer attack. Finally, we discuss the real-world implications of NeuroHammer on CIM architectures.
Felix Staudigl, Hazem Al Indari, Daniel Schön, Dominik Germek, Jan Moritz Joseph, Vikas Rana, Stephan Menzel, Amelie Hagelauer, Rainer Leupers
IEEE Trans. Reliab.1
2024 CLSA-CIM: A Cross-Layer Scheduling Approach for Computing-in-Memory Architectures
abstract
The demand for efficient machine learning (ML) accelerators is growing rapidly, driving the development of novel computing concepts such as resistive random access memory (RRAM)-based tiled computing-in-memory (CIM) architectures. CIM allows to compute within the memory unit, resulting in faster data processing and reduced power consumption. Efficient compiler algorithms are essential to exploit the potential of tiled CIM architectures. While conventional ML compilers focus on code generation for CPUs, GPUs, and other von Neumann architectures, adaptations are needed to cover CIM architectures. Cross-layer scheduling is a promising approach, as it enhances the utilization of CIM cores, thereby accelerating computations. Although similar concepts are implicitly used in previous work, there is a lack of clear and quantifiable algorithmic definitions for cross-layer scheduling for tiled CIM architectures. To close this gap, we present CLSA-CIM, a cross-layer scheduling algorithm for tiled CIM architectures. We integrate CLSA-CIM with existing weight-mapping strategies and compare performance against state-of-the-art (SOTA) scheduling algorithms. CLSA-CIM improves the utilization by up to 17.9 ×, resulting in an overall speedup increase of up to 29.2 × compared to SOTA.
Rebecca Pelke, José Cubero-Cascante, Nils Bosbach, Felix Staudigl, Rainer Leupers, Jan Moritz Joseph
DATE4
2023 Work-in-Progress: A Universal Instrumentation Platform for Non-Volatile Memories
abstract
Emerging non-volatile memories (NVMs) represent a disruptive technology that allows a paradigm shift from the conventional von Neumann architecture towards more efficient computing-in-memory (CIM) architectures. Several instrumentation platforms have been proposed to interface NVMs allowing the characterization of single cells and crossbar structures. However, these platforms suffer from low flexibility and are not capable of performing CIM operations on NVMs. Therefore, we recently designed and built the NeuroBreakoutBoard, a highly versatile instrumentation platform capable of executing CIM on NVMs. We present our preliminary results demonstrating a relative error < 5% in the range of 1 kΩ to 1 MΩ and showcase the switching behavior of a HfO2/Ti-based memristive cell.
Felix Staudigl, Mohammed Hossein, Tobias Ziegler 0005, Hazem Al Indari, Rebecca Pelke, Sebastian Siegel, Dirk J. Wouters, Dominik Germek, Jan Moritz Joseph, Rainer Leupers
CODES+ISSS1
2023 Fault Injection in Native Logic-in-Memory Computation on Neuromorphic Hardware
abstract
Logic-in-memory (LIM) describes the execution of logic gates within memristive crossbar structures, promising to improve performance and energy efficiency. Utilizing only binary values, LIM particularly excels in accelerating binary neural networks, shifting it in the focus of edge applications. Considering its potential, the impact of faults on BNNs accelerated with LIM still lacks investigation. In this paper, we propose faulty logic-in-memory (FLIM), a fault injection platform capable of executing full-fledged BNNs on LIM while injecting in-field faults. The results show that FLIM runs a single MNIST picture 66754× faster than the state of the art by offering a fine-grained fault injection methodology.
Felix Staudigl, Thorben Fetz, Rebecca Pelke, Dominik Germek, Jan Moritz Joseph, Letícia Maria Veiras Bolzani, Rainer Leupers
DAC1
2023 Mapping of CNNs on multi-core RRAM-based CIM architectures
abstract
Resistive random access memory (RRAM)-based multi-core systems improve the energy efficiency and performance of convolutional neural networks (CNNs). Thereby, the distributed parallel execution of convolutional layers causes critical data dependencies that limit the potential speedup. This paper presents synchronization techniques for parallel inference of convolutional layers on RRAM-based computing-in-memory (CIM) architectures. We propose an architecture optimization that enables efficient data exchange and discuss the impact of different architecture setups on the performance. The corresponding compiler algorithms are optimized for high speedup and low memory consumption during CNN inference. We achieve more than 99 % of the theoretical acceleration limit with a marginal data transmission overhead of less than 4 % for state-of-the-art CNN benchmarks.
Rebecca Pelke, Nils Bosbach, José Cubero-Cascante, Felix Staudigl, Rainer Leupers, Jan Moritz Joseph
VLSI-SoC4
2022 NeuroHammer: Inducing Bit-Flips in Memristive Crossbar Memories
abstract
Emerging non-volatile memory (NVM) technologies offer unique advantages in energy efficiency, latency, and features such as computing-in-memory. Consequently, emerging NVM technologies are considered an ideal substrate for computation and storage in future-generation neuromorphic platforms. These technologies need to be evaluated for fundamental reliability and security issues. In this paper, we present NeuroHammer, a security threat in ReRAM crossbars caused by thermal crosstalk between memory cells. We demonstrate that bit-flips can be deliberately induced in ReRAM devices in a crossbar by systematically writing adjacent memory cells. A simulation flow is developed to evaluate NeuroHammer and the impact of physical parameters on the effectiveness of the attack. Finally, we discuss the security implications in the context of possible attack scenarios.
Felix Staudigl, Hazem Al Indari, Daniel Schön, Dominik Germek, Farhad Merchant, Jan Moritz Joseph, Vikas Rana, Stephan Menzel, Rainer Leupers
DATE1
2022 EmuNoC: Hybrid Emulation for Fast and Flexible Network-on-Chip Prototyping on FPGAs
abstract
Networks-on-Chips (NoCs) recently became widely used, from multi-core CPUs to edge-AI accelerators. Emulation on FPGAs promises to accelerate their RTL modeling compared to slow simulations. However, realistic test stimuli are challenging to generate in hardware for diverse applications. In other words, both a fast and flexible design framework is required. The most promising solution is hybrid emulation, in which parts of the design are simulated in software, and the other parts are emulated in hardware. This paper proposes a novel hybrid emulation framework called EmuNoC. We introduce a clock-synchronization method and software-only packet generation that improves the emulation speed by 36.3 × to 79.3 × over state-of-the-art frameworks while retaining the flexibility of a pure-software interface for stimuli simulation. We also increased the area efficiency to model up to an NoC with 169 routers on a single FPGA, while previous frameworks only achieved 64 routers.
Yee Yang Tan, Felix Staudigl, Lukas Jünger 0001, Anna Drewes, Rainer Leupers, Jan Moritz Joseph
FPL2