EDBT 2026 Demo / reviewers in the wild / expert
Paul Palomero Bernardo
dblp:278/9119
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-6642-3976ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Smart Video Capsule Endoscopy: Raw Image-Based Localization for Enhanced GI Tract Investigation
Oliver Bause, Julia Werner, Paul Palomero Bernardo, Oliver Bringmann 0001 |
ICONIP (2) | 3 |
| 2025 | Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network AcceleratorsabstractImplementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the intended AI workload. To facilitate this, we present an automated generation approach for fast performance models to accurately estimate the latency of a DNN mapped onto systematically modeled and concisely described accelerator architectures. Using our accelerator architecture description method, we modeled representative DNN accelerators such as Gemmini, UltraTrail, Plasticine-derived, and a parameterizable systolic array. Together with DNN mappings for those modeled architectures, we perform a combined DNN/hardware dependency graph analysis, which enables us, in the best case, to evaluate only 154 loop kernel iterations to estimate the performance for 4.19 billion instructions achieving a significant speedup. We outperform regression and analytical models in terms of mean absolute percentage error (MAPE) compared with simulation results, while being several magnitudes faster than an RTL simulation. Konstantin Lübeck, Alexander Louis-Ferdinand Jung, Felix Wedlich, Mika Markus Müller, Federico Nicolás Peccia, Felix Thömmes, Jannik Steinmetz, Valentin Biermaier, Adrian Frischknecht, Paul Palomero Bernardo, Oliver Bringmann 0001 |
ACM Trans. Embed. Comput. Syst. | 10 |
| 2024 | A Scalable RISC-V Hardware Platform for Intelligent Sensor ProcessingabstractThis paper presents a demonstrator chip for an industrial audio event detection application developed as part of the Scale4Edge project. The project aims at enabling a comprehensive RISC-V based ecosystem to efficiently assemble well-tailored edge devices. The chip is manufactured in Globalfoundries' 22FDX technology and contains a RISC-V CPU with custom Instruction-Set-Architecture Extensions (ISAX) for fast AI and DSP processing, a low power neural network accelerator, and a scalable PLL to fulfill real-time processing requirements. By automated integration of these specialized hardware components, we achieve a speedup of ×2.15 while reducing the power by 27% compared to the unp[ntimized solution. Paul Palomero Bernardo, Patrick Schmid, Oliver Bringmann 0001, Mohammed Iftekhar, Babak Sadiye, Wolfgang Müller 0003, Andreas Koch 0001, Eyck Jentzsch, Axel Sauer, Ingo Feldner, Wolfgang Ecker |
DATE | 1 |
| 2024 | Energy-Efficient Seizure Detection Suitable for Low-Power ApplicationsabstractEpilepsy is the most common, chronic, neurological disease worldwide and is typically accompanied by reoccurring seizures. Neuro implants can be used for effective treatment by suppressing an upcoming seizure upon detection. Due to the restricted size and limited battery lifetime of those medical devices, the employed approach also needs to be limited in size and have low energy requirements. We present an energy-efficient seizure detection approach involving a TC-ResNet and time-series analysis which is suitable for low-power edge devices. The presented approach allows for accurate seizure detection without preceding feature extraction while considering the stringent hardware requirements of neural implants. The approach is validated using the CHB-MIT Scalp EEG Database with a 32-bit floating point model and a hardware suitable 4-bit fixed point model. The presented method achieves an accuracy of 95.28%, a sensitivity of 92.34% and an AUC score of 0.9384 on this dataset with 4-bit fixed point representation. Furthermore, the power consumption of the model is measured with the low-power AI accelerator UltraTrail, which only requires 495nW on average. Due to this low-power consumption this classification approach is suitable for real-time seizure detection on low-power wearable devices such as neural implants. Julia Werner, Bhavya Kohli, Paul Palomero Bernardo, Christoph Gerum, Oliver Bringmann 0001 |
IJCNN | 3 |
| 2024 | GOURD: Tensorizing Streaming Applications to Generate Multi-Instance Compute PlatformsabstractIn this article, we rethink the dataflow processing paradigm to a higher level of abstraction to automate the generation of multi-instance compute and memory platforms with interfaces to I/O devices (sensors and actuators). Since the different compute instances (NPUs, CPUs, DSPs, etc.) and I/O devices do not necessarily have compatible interfaces on a dataflow level, an automated translation is required. However, in multidimensional dataflow scenarios, it becomes inherently difficult to reason about buffer sizes and iteration order without knowing the shape of the data access pattern (DAP) that the dataflow follows. To capture this shape and the platform composition, we define a domain-specific representation (DSR) and devise a toolchain to generate a synthesizable platform, including appropriate streaming buffers for platform-specific tensorization of the data between incompatible interfaces. This allows platforms, such as sensor edge AI devices, to be easily specified by simply focusing on the shape of the data provided by the sensors and transmitted among compute units, giving the ability to evaluate and generate different dataflow design alternatives with significantly reduced design time. Patrick Schmid, Paul Palomero Bernardo, Christoph Gerum, Oliver Bringmann 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | The Scale4Edge RISC-V EcosystemabstractThis paper introduces the project Scale4Edge. The project is focused on enabling an effective RISC-V ecosystem for optimization of edge applications. We describe the basic components of this ecosystem and introduce the envisioned demonstrators, which will be used in their evaluation. Wolfgang Ecker, Peer Adelt, Wolfgang Müller 0003, Reinhold Heckmann, Milos Krstic, Vladimir Herdt, Rolf Drechsler, Gerhard Angst, Ralf Wimmer 0001, Andreas Mauderer, Rafael Stahl, Karsten Emrich, Daniel Mueller-Gritschneder, Bernd Becker 0001, Philipp M. Scholl, Eyck Jentzsch, Jan Schlamelcher, Kim Grüttner, Paul Palomero Bernardo, Oliver Bringmann 0001, Brindusa Mihaela Damian-Kosterhon, Julian Oppermann, Andreas Koch 0001, Jörg Bormann, Johannes Partzsch, Christian Mayr 0001, Wolfgang Kunz |
DATE | 19 |
| 2022 | Hardware Accelerator and Neural Network Co-Optimization for Ultra-Low-Power Audio Processing DevicesabstractThe increasing spread of artificial neural networks does not stop at ultralow-power edge devices. However, these very often have high computational demand and require specialized hardware accelerators to ensure the design meets power and performance constraints. The manual optimization of neural networks along with the corresponding hardware accelerators can be very challenging. This paper presents HANNAH (Hardware Accelerator and Neural Network seArcH), a framework for automated and combined hardware/software co-design of deep neural networks and hardware accelerators for resource and power-constrained edge devices. The optimization approach uses an evolution-based search algorithm, a neural network template technique and analytical KPI models for the configurable UltraTrail hardware accelerator template in order to find an optimized neural network and accelerator configuration. We demonstrate that HANNAH can find suitable neural networks with minimized power consumption and high accuracy for different audio classification tasks such as single-class wake word detection, multi-class keyword detection and voice activity detection, which are superior to the related work. Christoph Gerum, Adrian Frischknecht, Tobias Hald, Paul Palomero Bernardo, Konstantin Lübeck, Oliver Bringmann 0001 |
DSD | 4 |
| 2020 | UltraTrail: A Configurable Ultralow-Power TC-ResNet AI Accelerator for Efficient Keyword SpottingabstractRecent advances in machine learning show the superior behavior of temporal convolutional networks (TCNs) and especially their combination with residual networks (TC-ResNet) for intelligent sensor signal processing in comparison to classical CNNs and LSTMs. In this article, we propose UltraTrail, a configurable, ultralow-power TC-ResNet AI accelerator for sensor signal processing and its application to efficient keyword spotting (KWS). Following a strict hardware/model co-design approach, we have derived an optimized low-power hardware architecture for generalized TC-ResNet topologies consisting of a configurable array of processing elements and a distributed memory with dynamic content reallocation. We additionally extend the network with conditional computing to reduce the number of operations during inference and to provide the possibility for power-gating. The final accelerator implementation in Globalfoundries' 22FDX technology achieves a power consumption of 8.2 μW for the task of always-on KWS meeting the real-time requirement of 100 ms per inference with an accuracy of 93% on the Google Speech Command Dataset. Paul Palomero Bernardo, Christoph Gerum, Adrian Frischknecht, Konstantin Lübeck, Oliver Bringmann 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |