VLDB 2026 Research / reviewers in the wild / expert
Sergio A. Pertuz 0001
dblp:230/1150-1 · also Sergio Andres Pertuz Mendez
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-6311-3251ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoC-SLAM: FPGA-Based Hardware/Software Co-Design for Real-Time Visual SLAM Front-End and Back-End AccelerationabstractORB-SLAM3 is a state-of-the-art visual SLAM system, but its computational complexity poses major challenges for real-time deployment on embedded platforms. While prior work has largely focused on accelerating front-end tasks like feature extraction, back-end stages such as bundle adjustment remain less explored due to their algorithmic complexity and memory-intensive nature. Furthermore, most existing solutions accelerate specific modules without integrating them into a complete SLAM framework. In this work, we present SoC-SLAM, a novel FPGA-based hardware/software co-design that accelerates both the front-end and back-end stages of ORB-SLAM3 within a unified framework. Profiling identifies ORB feature extraction and local bundle adjustment as the primary performance bottlenecks, both critical for maintaining real-time system responsiveness. To address these, we develop modular FPGA-based accelerators for ORB extraction and key bundle adjustment solver steps, including Schur elimination, Cholesky decomposition, and back substitution, while retaining the remaining pipeline in software. Since the workload predominantly consists of operations on sparse block matrices, we develop and combine several optimization techniques, including matrix partitioning and pipelined block processing, to efficiently handle sparsity and maximize parallelism. Evaluation on the EuRoC MAV dataset shows $8 \times$ and $7.4 \times$ speedups for ORB extraction and local bundle adjustment, resulting in $2.4 \times$ and $3.3 \times$ improvements in the Tracking and Local Mapping threads, respectively. The system operates at 222 MHz, consumes 4.276 W, and achieves an RMSE of 0.02832 m. Our fully integrated pipeline demonstrates competitive performance and power efficiency compared to prior FPGA, ASIC, and GPU-based solutions. The proposed architecture is scalable and generalizable to other bundle adjustment modules, such as global bundle adjustment, welding bundle adjustment, and essential graph optimization, offering an extensible hardware acceleration design for real-time visual SLAM on resource-constrained platforms. Bhavay Arora, Paul Gottschaldt, Ariel Podlubne, Sergio A. Pertuz 0001, Diana Göhringer |
DSD | 4 |
| 2024 | Towards an Embedded System for Failure Diagnosis in Drones Using AI and SAC-DM on FPGAabstractWe present a way of failure detection in real-time unmanned aerial vehicles (UAVs) by integrating Chaos Theory and AI techniques on an FPGA board. The Signal Analysis based on Chaos using the Density of Maxima (SAC-DM) validates the input of the Machine Learning (ML) model due to the relation between the density of maxima and autocorrelation length. While the accuracies achieved solely by SAC-DM are not remarkably high, the ML model demonstrates an accuracy of 92.46% when utilizing sac-dm results as inputs. The unprecedented integration of SAC-DM on FPGA board serves as a solution for high-speed onboard processing, parallel integrated data synchronization and fusion, and an enhanced low-power architecture. Rafael Batista, Matthias Nickel, Alexander Lehnert, Sergio A. Pertuz 0001, Marc Reichenbach, Diana Göhringer, Alisson Brito |
DATE | 4 |
| 2024 | Hardware-level Access Control and Scheduling of Shared Hardware AcceleratorsabstractWith the trend to consolidate hardware on a single platform, FPGA virtualization plays an increasingly important role in the embedded domain. FPGA virtualization allows multiple software tasks or even guest operating systems to share recon- figurable resources. However, state-of-the-art approaches assign each hardware accelerator to a single software task for a fixed duration. This becomes a problem when the number of hardware accelerators required by software tasks concurrently exceeds the FPGA area. If several software tasks request to accelerate the same functionality, accelerators can be shared. Embedded reconfigurable systems face the challenge of a uniform address space. When several tasks use a memory-mapped communication interface that allows to directly access the accelerator's address space, access control and the protection from unauthorized access must be ensured. Existing software-based approaches lead to high latencies. Thus, we propose a hardware-level scheduler that schedules hardware tasks in spatial and temporal respect. The allocation to a hardware accelerator is combined with the assignment of access rights. Any unauthorized access leads to a page fault. When hardware tasks share an accelerator, they are scheduled according to the Earliest Deadline First (EDF) policy. Buffers ensure data isolation. Compared to hardware task scheduling in software, a performance increase of 7.02 times is reached. Cornelia Wulf, Sergio A. Pertuz 0001, Diana Göhringer |
DSD | 2 |
| 2023 | An Efficient Accelerator for Nonlinear Model Predictive ControlabstractThe computational complexity of Nonlinear Model Predictive Control (NMPC) often hinders their application to cyber-physical systems with fast dynamics, such as mobile robots or Unmanned Aerial Vehicles. This complexity overhead comes from the control algorithm's backbone, an iterative solver that must ensure convergence and often takes the form of a highly structured convex Quadratic Program (QP). Such overhead could be overcome using specialized computer architectures. Field Programmable Gate Arrays are good candidates for making hardware accelerators that comply with the realtime constraints of fast-dynamic cyber-physical systems. Nevertheless, QP-solvers have been demonstrated to be complex to implement as a hardware accelerator. With this in mind, the present paper proposes a novel accelerator architecture that uses Knowledge-based Particle Swarm Optimization (PSO) as a solver while exploring its parallel nature. PSO is a stochastic global optimization algorithm that creates a fast and precise solution for NMPC. The proposed strategy in this papergrants system control stability for short sampling frequencies and long prediction horizons. It can also meet realtime constraints while achieving low hardware consumption. Additionally, it is generalized, so it can potentially be adapted to any application and is compatible with the Robot Operating System (ROS). The architecture is tested with two applications: an inverted pendulum swing-up procedure and a quadrotor drone with control and state constraints. Following, we analyze the accelerator performance and highlight our solution's advantages to other works in the literature. Namely, our architecture solves more complex problems with a greater dimension and longer horizon while using similar resources. The proposed solution also has good computational performance (29ms and 11ms) for both the quadrotor and inverted pendulum, respectively, while achieving the realtime requirements (50ms and 100ms, respectively). Parallelly, ad-hoc embedded architectures are important for a low-end, low-cost, and low-power MPSoC+FPGA device. Our solution uses less than 50% of a low-end, low-power MPSoC device (ZU3EG), while others rely on large, more power-hungry devices (e.g., Kintex7 and XC7Z045). Sergio A. Pertuz 0001, Ariel Podlubne, Diana Göhringer |
ASAP | 1 |
| 2023 | EuFRATE: European FPGA Radiation-hardened Architecture for TelecommunicationsabstractThe EuFRATE project aims to research, develop and test radiation-hardening methods for telecommunication payloads deployed for Geostationary-Earth Orbit (GEO) using Commercial-Off- The-Shelf Field Programmable Gate Arrays (FPGAs). This project is conducted by Argotec Group (Italy) with the collaboration of two partners: Politecnico di Torino (Italy) and Technische Universität Dresden (Germany). The idea of the project focuses on high-performance telecommunication algorithms and the design and implementation strategies for connecting an FPGA device into a robust and efficient cluster of multi-FPGA systems. The radiation-hardening techniques currently under development are addressing both device and cluster levels, with redundant datapaths on multiple devices, comparing the results and isolating fatal errors. This paper introduces the current state of the project's hardware design description, the composition of the FPGA cluster node, the proposed cluster topology, and the radiation hardening techniques. Intermediate stage experimental results of the FPGA communication layer performance and fault detection techniques are presented. Finally, a wide summary of the project's impact on the scientific community is provided.1 Ludovica Bozzoli, Antonino Catanese, Emilio Fazzoletto, Eugenio Scarpa, Diana Göhringer, Sergio A. Pertuz 0001, Lester Kalms, Cornelia Wulf, Najdet Charaf, Luca Sterpone, Sarah Azimi, Daniele Rizzieri, Salvatore Gabriele La Greca, David Merodio Codinachs |
DATE | 6 |
| 2022 | Model-based Generation of Hardware/Software Architectures for Robotics SystemsabstractRobotic systems compute data from multiple sensors to perform several actions (e.g., path planning, object detection). FPGA - based architectures for such systems may consist of several accelerators to process compute-intensive algorithms. Designing and implementing such complex systems tends to be an arduous task. This work proposes a modeling approach to generate architectures for such applications, compliant with existing robotics middlewares (e.g., ROS, ROS2). The challenge is to have a compact, yet expressive description of the system with just enough information to generate all required components and to integrate existing algorithms. This system model must be generalizable, so it is not application-dependent, and it must exploit the benefits of FPGAs over software solutions. Previous work mainly focused on individual accelerators rather than all components involved in a system and their interactions. The proposed approach exploits the advantages of model-driven engineering and model-based code generation to produce all components, i.e., message converters acting as middleware interfaces and wrappers to integrate algorithms. Data type and data flow analysis are performed to derive the necessary information to generate the components and their connections. Solutions to several identified challenges for generating entire systems from such models are evaluated using four different use cases. Ariel Podlubne, Johannes Mey, Sergio A. Pertuz 0001, Uwe Aßmann, Diana Göhringer |
FPL | 3 |
| 2022 | Reflections on "Rock, Paper, Scissors": Communicating Science to the Public through a DemonstratorabstractCommunicating science to the public is increasingly important. Demonstrators are a valuable and established tool for communication in technology research and development. However, their role in communicating current science and technology to the public has not received much attention neither in research nor practice. This paper reflects on the design and usage of the demonstrator “Rock, Paper, Scissors”, which we developed to communicate current advances in Human-Robot Interaction to public audiences. We discuss two years of “Rock, Paper, Scissors” in action and its evolution within this period. We conclude with an outlook to future work regarding technology development and evaluation of science communication. Tina Bobbe, Hans Winger, Ariel Podlubne, Florian Wieczorek, Lisa-Marie Lüneburg, Ievgen Kharabet, Jens Wagner, Sergio A. Pertuz 0001 |
HRI | 8 |
| 2021 | Optimized Deep Learning Object Recognition for Drones using Embedded GPUabstractNowadays, drones can be seen in various applications in industry like surveillance and transportation. Industrial drones leverage fully-fledged computer vision techniques, such as object detection based on Deep Learning Neural Networks (DNN), to efficiently perform these objectives. Those techniques come with a high computational effort and are implemented on distributed schemes using ground devices with high performance and power consumption. This limits a drone's operational range since it has to communicate with the ground devices constantly. To alleviate such constraints, an optimized, low-power perception system on the drone is desirable. This work improves a trained DNN architecture to navigate a UAV introduced by the University of Zurich called DroNet. DroNet is computationally expensive and has a high power consumption, making it unsuitable for embedded platforms because of low memory and computational power. In this paper, a ROS-based architecture is first designed to port DroNet on a low-power Jetson Nano board, which conducts the drone's perception and control tasks. Secondly, tuning parameters and various schemes have been carried out to run the inference of the DNN efficiently. To implement the different layers in DNNs, Nvidia's TensorRT SDK is used to compile a high-performance inference engine for the Jetson Nano. Results showed that the Jetson Nano can achieve real-time performance, with 47 frames per second using a Winograd convolution and well-tuned parallelization parameters. The implementation can also achieve a speedup of 2× as compared with the Jetson Nanos ARM CPU while increasing the power consumption by 54%. Finally, the Jetson Nano's usability for drone inference algorithm is shown, achieving real-time response using the DroNet DNN without losing detection accuracy. Pedram Amini Rad, Danny Hofmann, Sergio A. Pertuz 0001, Diana Göhringer |
ETFA | 3 |