VLDB 2026 Research / reviewers in the wild / expert
Rebecca Pelke
dblp:340/3898
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-5156-7072ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed-Precision Training and Compilation for RRAM-based Computing-in-Memory AcceleratorsabstractComputing-in-Memory (CIM) accelerators are a promising solution for accelerating Machine Learning (ML) workloads, as they perform Matrix-Vector Multiplications (MVMs) on crossbar arrays directly in memory. Although the bit widths of the crossbar inputs and cells are very limited, most CIM compilers do not support quantization below 8 bit. As a result, a single MVM requires many compute cycles, and weights cannot be efficiently stored in a single crossbar cell.To address this problem, we propose a mixed-precision training and compilation framework for CIM architectures. The biggest challenge is the massive search space, that makes it difficult to find good quantization parameters. This is why we introduce a reinforcement learning-based strategy to find suitable quantization configurations that balance latency and accuracy. In the best case, our approach achieves up to a 2.48× speedup over existing state-of-the-art solutions, with an accuracy loss of only 0.086 %. Rebecca Pelke, Joel Klein, José Cubero-Cascante, Nils Bosbach, Jan Moritz Joseph, Rainer Leupers |
DATE | 1 |
| 2025 | Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning WorkloadsabstractAccelerating Machine Learning (ML) workloads requires efficient methods due to their large optimization space. Autotuning has emerged as an effective approach for systematically evaluating variations of implementations. Traditionally, autotuning requires the workloads to be executed on the target hardware (HW). We present an interface that allows executing autotuning workloads on simulators. This approach offers high scalability when the availability of the target HW is limited, as many simulations can be run in parallel on any accessible HW.Additionally, we evaluate the feasibility of using fast instruction-accurate simulators for autotuning. We train various predictors to forecast the performance of ML workload implementations on the target HW based on simulation statistics.Our results demonstrate that the tuned predictors are highly effective. The best workload implementation in terms of actual run time on the target HW is always within the top 3% of predictions for the tested x86, ARM, and RISC-V-based architectures. In the best case, this approach outperforms native execution on the target HW for embedded architectures when running as few as three samples on three simulators in parallel. Rebecca Pelke, Nils Bosbach, Lennart M. Reimann, Rainer Leupers |
DAC | 1 |
| 2025 | High-Performance ARM-on-ARM Virtualization for Multicore SystemC-TLM-Based Virtual PlatformsabstractThe increasing complexity of hardware and software requires advanced development and test methodologies for modern systems on chips. This paper presents a novel approach to ARM-on-ARM virtualization within SystemC-based simulators using Linux's KVM to achieve high-performance simulation. By running target software natively on ARM-based hosts with hardware-based virtualization extensions, our method eliminates the need for instruction-set simulators, which significantly improves performance. We present a multicore SystemC-TLM-based CPU model that can be used as a drop-in replacement for an instruction-set simulator. It places no special requirements on the host system, making it compatible with various environments. Benchmark results show that our ARM-on-ARM-based virtual platform achieves up to 10 x speedup over traditional instruction-set-simulator-based models on compute-intensive workloads. Depending on the benchmark, speedups increase to more than 100 x. Nils Bosbach, Rebecca Pelke, Niko Zurstraßen, Jan Weinstock, Lukas Jünger 0001, Rainer Leupers |
DATE | 2 |
| 2025 | CIMFlow: Modelling Dataflow in Cross-Layer Compute-in-Memory Deep Learning AcceleratorsabstractTraditional Deep Learning Accelerators (DLAs) rely on off-chip memory to store large weight tensors, leading to high bandwidth demands and energy consumption. Compute-in-Memory (CIM) accelerators mitigate this by integrating high-density, non-volatile memory arrays, enabling a fully weight-stationary (FWS) dataflow. Multi-core CIM systems further enhance efficiency with cross-layer inference, where intermediate tensors stay on-chip, and cores operate in a pipeline. Despite diverse architecture proposals, no existing tool models the dataflow, memory access patterns and timing behaviour of multi-core CIM accelerators. We introduce CIMFlow, a modelling framework for cross-layer CIM architectures. Our flexible Hardware Architecture Model includes CIM and digital cores and leverages the buffets storage idiom for distributed token-based flow control. An Array-OL-based Workload Model captures CNNs’ multidimensional dependencies and applies hardware-aware transformations. These models are transformed into a timed cyclo-static dataflow graph for simulation. CIMFlow delivers latency, energy and traces for core and buffer utilisation. Our case studies on state-of-the-art CNNs show that cross-layer inference reduces latency by up to 52×. We also reveal that neglecting memory access delays results in throughput overestimations of up to 308%. To our knowledge, CIMFlow is the first tool focused on FWS cross-layer execution in CIM architectures that explicitly models data movement costs. It serves as a powerful cost model for design space exploration in next-generation CIM accelerators. José Cubero-Cascante, Lucas Tonini Rosenberg Schneider, Rebecca Pelke, Arunkumar Vaidyanathan, Rainer Leupers, Jan Moritz Joseph |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Towards High-Performance Virtual Platforms: A Parallelization Strategy for SystemC TLM-2.0 CPU ModelsabstractSystemC TLM-2.0 is currently the industry standard for simulating full Systems-on-a-Chip (SoCs). Although SystemC is designed to simulate the behavior of complex, parallel systems, the simulation itself is by default single-threaded. We present a technique to overcome this performance limitation by parallelizing the CPU model of a SystemC-TLM-2.0-based system-level simulator, a so-called Virtual Platform (VP). Our solution is fully compliant with the SystemC standard. To further increase the performance, we developed algorithms for asynchronous DMI pointer caching and we introduced a new tunable parameter called async_rate. This parameter controls the frequency used to annotate timing information to SystemC. Nils Bosbach, Niko Zurstraßen, Rebecca Pelke, Lukas Jünger 0001, Jan Weinstock, Rainer Leupers |
DAC | 3 |
| 2024 | CLSA-CIM: A Cross-Layer Scheduling Approach for Computing-in-Memory ArchitecturesabstractThe demand for efficient machine learning (ML) accelerators is growing rapidly, driving the development of novel computing concepts such as resistive random access memory (RRAM)-based tiled computing-in-memory (CIM) architectures. CIM allows to compute within the memory unit, resulting in faster data processing and reduced power consumption. Efficient compiler algorithms are essential to exploit the potential of tiled CIM architectures. While conventional ML compilers focus on code generation for CPUs, GPUs, and other von Neumann architectures, adaptations are needed to cover CIM architectures. Cross-layer scheduling is a promising approach, as it enhances the utilization of CIM cores, thereby accelerating computations. Although similar concepts are implicitly used in previous work, there is a lack of clear and quantifiable algorithmic definitions for cross-layer scheduling for tiled CIM architectures. To close this gap, we present CLSA-CIM, a cross-layer scheduling algorithm for tiled CIM architectures. We integrate CLSA-CIM with existing weight-mapping strategies and compare performance against state-of-the-art (SOTA) scheduling algorithms. CLSA-CIM improves the utilization by up to 17.9 ×, resulting in an overall speedup increase of up to 29.2 × compared to SOTA. Rebecca Pelke, José Cubero-Cascante, Nils Bosbach, Felix Staudigl, Rainer Leupers, Jan Moritz Joseph |
DATE | 1 |
| 2023 | Work-in-Progress: A Generic Non-Intrusive Parallelization Approach for SystemC TlM-2.0-Based Virtual Platforms
Nils Bosbach, Rebecca Pelke, Niko Zurstraßen, Lukas Jünger 0001, Jan Weinstock, Rainer Leupers |
CODES+ISSS | 2 |
| 2023 | Work-in-Progress: A Universal Instrumentation Platform for Non-Volatile MemoriesabstractEmerging non-volatile memories (NVMs) represent a disruptive technology that allows a paradigm shift from the conventional von Neumann architecture towards more efficient computing-in-memory (CIM) architectures. Several instrumentation platforms have been proposed to interface NVMs allowing the characterization of single cells and crossbar structures. However, these platforms suffer from low flexibility and are not capable of performing CIM operations on NVMs. Therefore, we recently designed and built the NeuroBreakoutBoard, a highly versatile instrumentation platform capable of executing CIM on NVMs. We present our preliminary results demonstrating a relative error < 5% in the range of 1 kΩ to 1 MΩ and showcase the switching behavior of a HfO2/Ti-based memristive cell. Felix Staudigl, Mohammed Hossein, Tobias Ziegler 0005, Hazem Al Indari, Rebecca Pelke, Sebastian Siegel, Dirk J. Wouters, Dominik Germek, Jan Moritz Joseph, Rainer Leupers |
CODES+ISSS | 5 |
| 2023 | Fault Injection in Native Logic-in-Memory Computation on Neuromorphic HardwareabstractLogic-in-memory (LIM) describes the execution of logic gates within memristive crossbar structures, promising to improve performance and energy efficiency. Utilizing only binary values, LIM particularly excels in accelerating binary neural networks, shifting it in the focus of edge applications. Considering its potential, the impact of faults on BNNs accelerated with LIM still lacks investigation. In this paper, we propose faulty logic-in-memory (FLIM), a fault injection platform capable of executing full-fledged BNNs on LIM while injecting in-field faults. The results show that FLIM runs a single MNIST picture 66754× faster than the state of the art by offering a fine-grained fault injection methodology. Felix Staudigl, Thorben Fetz, Rebecca Pelke, Dominik Germek, Jan Moritz Joseph, Letícia Maria Veiras Bolzani, Rainer Leupers |
DAC | 3 |
| 2023 | Mapping of CNNs on multi-core RRAM-based CIM architecturesabstractResistive random access memory (RRAM)-based multi-core systems improve the energy efficiency and performance of convolutional neural networks (CNNs). Thereby, the distributed parallel execution of convolutional layers causes critical data dependencies that limit the potential speedup. This paper presents synchronization techniques for parallel inference of convolutional layers on RRAM-based computing-in-memory (CIM) architectures. We propose an architecture optimization that enables efficient data exchange and discuss the impact of different architecture setups on the performance. The corresponding compiler algorithms are optimized for high speedup and low memory consumption during CNN inference. We achieve more than 99 % of the theoretical acceleration limit with a marginal data transmission overhead of less than 4 % for state-of-the-art CNN benchmarks. Rebecca Pelke, Nils Bosbach, José Cubero-Cascante, Felix Staudigl, Rainer Leupers, Jan Moritz Joseph |
VLSI-SoC | 1 |