VLDB 2026 Research / reviewers in the wild / expert
Bruno Endres Forlin
dblp:218/1093 · also Bruno E. Forlin
· DBLP profile ↗
13ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-4822-1841ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shooting Neutrons at Neurons: Radiation Testing of a Spiking Neural Network on Flash-Based FPGAs
Wim Nijsink, Bruno Endres Forlin, Amirreza Yousefzadeh, Marco Ottavi |
ETS | 2 |
| 2025 | Bloom Filters for Soft Error Detection: Neutron and Fault Injection ValidationabstractAs memory cells continue to shrink in modern semiconductor technologies, radiation-induced Single Event Effects, such as single- and multi-bit upsets, pose growing challenges to system reliability. While effective and efficient for single and double-bit errors, traditional error detection and correction approaches, such as Error Correcting Codes (ECC), incur substantial overhead and complexity when designed to detect and correct multiple-bit errors. This study investigates the use of probabilistic data structures (PDS) as lightweight detectors for multiple-bit soft errors in memories. Leveraging the space-efficient and low-latency properties of Bloom filters, we implement a lightweight error detector (checker) within a representative memory subsystem on a flash-based FPGA. The checker's performance is validated through extensive neutron beam irradiation and fault-injection campaigns, demonstrating effective detection of multiple-bit errors with a tunable false-positive rate. Elijah Cishugi, Tijmen T. Smit, Bruno Endres Forlin, Carlo Cazzaniga, Kuan-Hsun Chen, Marco Ottavi |
IOLTS | 3 |
| 2024 | Lightweight Instrumentation for Accurate Performance Monitoring in RTOSesabstractEvaluating performance metrics in embedded systems poses challenges, particularly due to the limited set of tools available for monitoring performance counters. In addition, performance evaluation frameworks for Real-Time Operating Systems (RTOSes) often lack the sophistication and capabilities available in general-purpose operating systems like Linux, which benefit from utilities such as perf_event. To bridge this gap, this paper presents an accurate and low-overhead instrumentation utility tailored for RTOSes. Our approach utilizes performance monitoring counters to observe individual user applications within the RTOS environment. Importantly, it enables comprehensive application monitoring by strategically placing probes at points of inherent system interference, thereby minimizing additional overhead. A pre-calibration of these probes allows for fine-grained measurements within user applications. This results in the elimination of 100 % of the overheads for most counters in our test configuration, impacting the context switch by only three additional instructions per monitored counter. Bruno Endres Forlin, Kuan-Hsun Chen, Nikolaos Alachiotis 0001, Luca Cassano, Marco Ottavi |
DATE | 1 |
| 2024 | Neutron Beam Evaluation of Probabilistic Data Structure-based Online CheckersabstractHigh-criticality applications are vulnerable to Single Event Effects (SEEs) and require highly reliable and customizable microprocessors. Online checkers have been used to detect security and reliability issues in such systems. Popular hardware redundancy techniques such as Triple Modular Redundancy (TMR) and Dual Modular Redundancy (DMR) provide a high error coverage at the cost of substantial redundancy; therefore, there is an interest in introducing lightweight checkers that could offer the same detection ability as DMR with a much lower overhead. A possible implementation of these online checkers can be based on Probabilistic Data Structure (PDS) such as the Bloom Filter (BF). They are a form of information redundancy and an excellent complement to Single Error Correction Double Error Detection (SECDED) codes because they allow for detecting higher-order upsets. In this work, we integrate an online checker into the open-source RISC-V core NEORV32 and deploy it on a flash-based FPGA. This paper presents the evaluation of the online checker’s performance conducted under a neutron beam. The neutron beam experiments demonstrate that the real-life error rates of such structures are comparably worse than the initial simulation would indicate and that other factors can impact their performance. Bruno Endres Forlin, Edian B. Annink, Elijah Cishugi, Carlo Cazzaniga, Paolo Rech, Gerard K. Rauwerda, Gianluca Furano, Marco Ottavi |
IOLTS | 1 |
| 2023 | An unprotected RISC-V Soft-core processor on an SRAM FPGA: Is it as bad as it sounds?abstractFast development, low cost, and reconfigurability are becoming critical factors for aerospace applications, making SRAM FPGAs attractive. However, SRAM FPGAs are prone to errors in the user and on the configuration bits. For their correct functioning, they must be capable of withstanding failures without sacrificing much performance. When adjusting a soft core for these applications, it is essential to know where redundancies are necessary, to avoid unnecessary overhead. We characterize the reliability of an unprotected RISC-V microcontroller using an accelerated neutron beam. Our investigation shows that, for our chosen benchmark and processor, the user data in the memory banks is the leading cause of the total number of errors in the application. By reversing the benchmark operations, we could root cause the origin of the observed errors and found that most of the data corruption detected during the runs stem from previously corrupt input data or from output data that were corrupted while transmitting. Bruno Endres Forlin, Wouter van Huffelen, Carlo Cazzaniga, Paolo Rech, Nikolaos Alachiotis 0001, Marco Ottavi |
ETS | 1 |
| 2023 | Plug N' PIM: An integration strategy for Processing-in-Memory accelerators
Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
Integr. | 2 |
| 2022 | Aggressive Performance Improvement on Processing-in-Memory Devices by Adopting HugepagesabstractProcessing-in-Memory (PIM) devices integrated into general-purpose systems demand virtual memory support. In this way, these devices can be seamlessly coupled to the software stack, while maintaining compatibility and security provided by address management via the Operating System (OS) without requiring disruptive programming efforts. Typically, PIM intends to access large volumes of data via vector operations, and thus can suffer severe penalties due to the high cost of page misses in the Translation Look-aside Buffer (TLB). Our study demonstrates the criticality of such penalties on the system's performance and that PIM must resort to large page sizes. The presented results exploit the native large pages available on the host, and they show substantial performance improvements$(84\times)$for wide-vector PIM operations with large pages. Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
ASAP | 2 |
| 2022 | Sim2PIM: A complete simulation framework for Processing-in-Memory
Bruno Endres Forlin, Paulo C. Santos 0001, Augusto E. Becker, Marco A. Z. Alves, Luigi Carro |
J. Syst. Archit. | 1 |
| 2021 | Providing Plug N' Play for Processing-in-Memory AcceleratorsabstractAlthough Processing-in-Memory (PIM) emerged as a solution to avoid unnecessary and expensive data movements to/from host and accelerators, their widespread usage is still difficult, given that to effectively use a PIM device, huge and costly modifications must be done at the host processor side to allow instructions offloading, cache coherence, virtual memory management, and communication between different PIM instances. The present work addresses these challenges by presenting non-invasive solutions for those requirements. We demonstrate that, at compile-time, and without any host modifications or programmer intervention, it is possible to exploit already available resources to allow efficient host and PIM communication and task partitioning, without disturbing neither host nor memory hierarchy. We present Plug&PIM, a plug n' play strategy for PIM adoption with minimal performance penalties. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
ASP-DAC | 2 |
| 2021 | Sim2PIM: A Fast Method for Simulating Host Independent & PIM Agnostic DesignsabstractProcessing-in-Memory (PIM), with the help of modern memory integration technologies, has emerged as a practical approach to mitigate the memory wall and improve performance and energy efficiency in contemporary applications. However, there is a need for tools capable of quickly simulating different PIMs designs and their suitable integration with different hosts. This work presents Sim2PIM, a Simple Simulator for PIM devices that seamlessly integrates any PIM architecture with the host processor and memory hierarchy. Sim2PIM's simulation environment allows the user to describe a PIM architecture in different user-defined abstraction levels. The application code runs natively on the Host, with minimal overhead from the simulator integration, allowing Sim2PIM to collect precise metrics from the Hardware Performance Counters (HPCs). Our simulator is available to download at https://pim.computer/. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
DATE | 2 |
| 2020 | G-PUF: An Intrinsic PUF Based on GPU Error SignaturesabstractPhysically Unclonable Functions (PUFs) are security primitives that provide trustworthy hardware for key-generation and device authentication. Among them, in contrast to dedicated PUFs, intrinsic PUFs are created from existing hardware components that exploit their variability through software. In this work we focus on GPUs and present G-PUF, a PUF implemented entirely in software on CUDA and hence does not require hardware modifications. Our results show that G-PUF has comparable characteristics to SRAM and DRAM PUFs in terms of uniformity 55.61% and reliability 90.09%. Bruno Endres Forlin, Ronaldo Husemann, Luigi Carro, Cezar Reinbrecht, Said Hamdioui, Mottaqiallah Taouil |
ETS | 1 |
| 2019 | Attacking Real-time MPSoCs: Preemptive NoCs are VulnerableabstractMulti-Processor System-on-Chip is one of the todays standard platforms which has being used in several applications, including time critical. In order to meet safety, thus attending real-time constraints, security may be put aside during the design stage. This is the case of the Priority-Preemptive NoCs, a widely used real-time interconnection structure. Their explicit behavior while dealing with communication flows constrained by tight deadlines creates security flaws. To this end, this work presents three contributions. First, we demonstrate for the first time an attack that exploits preemptive NoC-based MPSoCs. Second, we integrate security countermeasures that avoid these attacks while meeting hard deadlines. Third, we evaluate the impact of the attacks and the protected system. Results show that preemptive NoCs must be protected and that it is possible to effectively and efficiently mitigate the vulnerabilities while keeping the deterministic behavior required for time-critical applications. Bruno Endres Forlin, Cezar Reinbrecht, Martha Johanna Sepúlveda |
VLSI-SoC | 1 |
| 2018 | Earthquake - A NoC-based optimized differential cache-collision attack for MPSoCsabstractMulti-Processor Systems-on-Chips (MPSoCs) are a platform for a wide variety of applications and use-cases. The high on-chip connectivity, the programming flexibility, and the reuse of IPs, however, also introduce security concerns. Problems arise when applications with different trust and protection levels share resources of the MPSoC, such as processing units, cache memories and the Network-on-Chip (NoC) communication structure. If a program gets compromised, an adversary can observe the use of these resources and infer (potentially secret) information from other applications. In this work, we explore the cache-based attack by Bogdanov et al., which infers the cache activity of a target program through timing measurements and exploits collisions that occur when the same cache location is accessed for different program inputs. We implement this differential cache-collision attack on the MPSoC Glass and introduce an optimized variant of it, the Earthquake Attack, which leverages the NoC-based communication to increase attack efficiency. Our results show that Earthquake performs well under different cache line and MPSoC configurations, illustrating that cache-collision attacks are considerable threats on MPSoCs. Cezar Reinbrecht, Bruno Endres Forlin, Andreas Zankl, Martha Johanna Sepúlveda |
DATE | 2 |