Federico Nicolás Peccia

dblp:335/2522 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-3587-0415ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LLM-aided Test Generation for Custom Neural Network Hardware Accelerators
Federico Nicolás Peccia, Tobias Hald, Oliver Bringmann 0001
ETS1
2025 Tensor Program Optimization for the RISC-V Vector Extension Using Probabilistic Programs
abstract
RISC-V provides a flexible and scalable platform for applications ranging from embedded devices to high-performance computing clusters. Particularly, its RISC-V Vector Extension (RVV) becomes of interest for the acceleration of AI workloads. But writing software that efficiently utilizes the vector units of RISC-V CPUs without expert knowledge requires the programmer to rely on the autovectorization features of compilers or hand-crafted libraries like muRISCV-NN. Smarter approaches, like autotuning frameworks, have been missing the integration with the RISC-V RVV extension, thus heavily limiting the efficient deployment of complex AI workloads. In this paper, we present a workflow based on the TVM compiler to efficiently map AI workloads onto RISC-V vector units. Instead of relying on hand-crafted libraries, we integrated the RVV extension into TVM’s MetaSchedule framework, a probabilistic program framework for tensor operation tuning. We implemented different RISC-V SoCs on an FPGA and tuned a wide range of AI workloads on them. We found that our proposal shows a mean improvement of 46% in execution latency when compared against the autovectorization feature of GCC, and 29% against muRISCV-NN. Moreover, the binary resulting from our proposal has a smaller code memory footprint, making it more suitable for embedded devices. Finally, we also evaluated our solution on a commercially available RISC-V SoC implementing the RVV 1.0 Vector Extension and found our solution is able to find mappings that are 35% faster on average than the ones proposed by LLVM. We open-sourced our proposal for the community to expand it to target other RISC-V extensions.
Federico Nicolás Peccia, Frederik Haxel, Oliver Bringmann 0001
ICCAD1
2025 Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
abstract
Implementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the intended AI workload. To facilitate this, we present an automated generation approach for fast performance models to accurately estimate the latency of a DNN mapped onto systematically modeled and concisely described accelerator architectures. Using our accelerator architecture description method, we modeled representative DNN accelerators such as Gemmini, UltraTrail, Plasticine-derived, and a parameterizable systolic array. Together with DNN mappings for those modeled architectures, we perform a combined DNN/hardware dependency graph analysis, which enables us, in the best case, to evaluate only 154 loop kernel iterations to estimate the performance for 4.19 billion instructions achieving a significant speedup. We outperform regression and analytical models in terms of mean absolute percentage error (MAPE) compared with simulation results, while being several magnitudes faster than an RTL simulation.
Konstantin Lübeck, Alexander Louis-Ferdinand Jung, Felix Wedlich, Mika Markus Müller, Federico Nicolás Peccia, Felix Thömmes, Jannik Steinmetz, Valentin Biermaier, Adrian Frischknecht, Paul Palomero Bernardo, Oliver Bringmann 0001
ACM Trans. Embed. Comput. Syst.5
2024 DIAPASON: Differentiable Allocation, Partitioning and Fusion of Neural Networks for Distributed Inference
abstract
Concerns in areas such as privacy, energy consumption, climate gas emissions, and costs, push the trend of migrating neural network inference from being executed on the cloud to embedded edge devices. We present our novel approach DIA-PASON to overcome restrictions brought on by the computing and application requirements that impede their execution on resource-constrained embedded devices. Our approach addresses these challenges by distributing the inference across multiple computing instances, which could be anything ranging from a multi-CPU configuration on the same SoC to geographically distributed devices. In contrast to recent efforts which tend to apply heuristics to reduce the search space of their problem definition to solve it in a timely fashion, our novel problem definition applies the concept of continuous relaxation to the categorical selection of partitioning, layer fusion, and allocation opportunities. This approach overcomes the problem of the poor exploration of the actual search space that arises when removing potential distribution opportunities during problem simplifications. We conduct numerical simulations and ablation experiments on each one of the configuration parameters of our algorithm by distributing several widely used neural networks. Finally, we compare DIAPASON against the commonly used MoDNN baseline and a state-of-the-art approach, CoopAI, achieving a 44 % and 12 % mean speed-up respectively.
Federico Nicolás Peccia, Alexander Viehl, Oliver Bringmann 0001
DATE1
2024 Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
abstract
The growing concerns regarding energy consumption and privacy have prompted the development of AI solutions deployable on the edge, circumventing the substantial CO2 emissions associated with cloud servers and mitigating risks related to sharing sensitive data. But deploying Convolutional Neural Networks (CNNs) on non-off-the-shelf edge devices remains a complex and labor-intensive task. In this paper, we present an end-to-end workflow for the deployment of CNNs on Field Programmable Gate Arrays (FPGAs) using the Gemmini accelerator, which we modified for efficient implementation on FPGAs. We describe how we leverage the use of open-source software on each optimization step of the deployment process, the customizations we added to them and their impact on the final system's performance. We were able to achieve real-time performance by deploying a YOLOv7 model on a Xilinx ZCU102 FPGA with an energy efficiency of 36.5 GOP/s/W. Our FPGA-based solution demonstrates superior power efficiency compared with other embedded hardware devices and even outperforms other FPGA reference implementations. Finally, we present how this kind of solution can be integrated into a wider system, by testing our proposed platform in a traffic monitoring scenario.
Federico Nicolás Peccia, Svetlana Pavlitska, Tobias Fleck, Oliver Bringmann 0001
DSD1
2024 Iterative Filter Pruning for Concatenation-based CNN Architectures
abstract
Model compression and hardware acceleration are essential for the resource-efficient deployment of deep neural networks. Modern object detectors have highly interconnected convolutional layers with concatenations. In this work, we study how pruning can be applied to such architectures, exemplary for YOLOv7. We propose a method to handle concatenation layers, based on the connectivity graph of convolutional layers. By automating iterative sensitivity analysis, pruning, and subsequent model fine-tuning, we can significantly reduce model size both in terms of the number of parameters and FLOPs, while keeping comparable model accuracy. Finally, we deploy pruned models to FPGA and NVIDIA Jetson Xavier AGX. Pruned models demonstrate a 2x speedup for the convolutional layers in comparison to the unpruned counterparts and reach real-time capability with 14 FPS on FPGA. Our code is available at https://github.com/fzi-forschungszentrum-informatik/iterative-yolo-pruning.
Svetlana Pavlitska, Oliver Bagge, Federico Nicolás Peccia, Toghrul Mammadov, Johann Marius Zöllner
IJCNN3