EDBT 2026 Demo / reviewers in the wild / expert
Vincenzo Maisto
dblp:321/0185
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-1631-1597ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A hardware/software architecture for multi-threaded offloading of erasure codes in distributed file systemsabstractBig Data analytics and cloud computing impose an ever-growing demand for data-center providers in terms of computational requirements, latency, and storage. Distributed file systems offer the strategic advantage of scaling-out computing and storage resources, hence allowing for notable speed-ups with massively parallel and distributed computing paradigms. On the other hand, such distributed clusters are constantly challenged with storage failures. Data replication is often deployed to ensure fault tolerance and business continuity, typically in a 3x configuration. This results in expensive 200 % overheads in storage space, write propagation, and energy costs. Erasure codes offer an alternative approach for fault tolerance by allowing reconstruction of erased data chunks, while reducing storage overhead down to 30 %. However, a considerable share of CPU cycles and energy is spent computing such codes, effectively reducing the cluster’s efficiency and starving other user and system tasks. Offloading on a custom accelerator is a non-trivial issue, due to the highly multi-threaded nature of such tasks and the lack of robust multi-threading support in conventional accelerator runtimes. In this work, we present a heterogeneous hardware/software architectural design for large-scale and multi-threaded acceleration of distributed erasure codes on PCIe accelerators, and a new abstraction and integration model for distributed accelerators in fault-tolerant storage systems. We enable safe and seamless deployment of multi-threaded SYCL-based IP cores through a hardware thread proxying layer providing software thread-isolation, and integration with cluster-level middlewares. In addition, our design allows for heterogeneous cluster configurations, with full compatibility and transparent integration of heterogeneously-accelerated and CPU-only nodes. We systematically evaluate the individual layers of our architecture and validate design’s integration in a container-based HDFS cluster, comparing performance against the state-of-the-art AVX-512-accelerated ISA-L library and other SYCL substrates, such as GPUs and single-threaded FPGAs. Vincenzo Maisto, Alessandro Cilardo, Emilio Billi, Chuck Fader |
Future Gener. Comput. Syst. | 1 |
| 2026 | Distilling knowledge for low-energy AIoTabstractThe Artificial Intelligence of Things (AIoT) empowers IoT devices to leverage the advantages of AI near data-sources, reducing data movement, latency, and mitigating privacy issues. However, AI workloads are notoriously energy-intensive, posing significant challenges for energy-constrained IoT devices. Since such devices are often deployed in thousands of instances, even minor inefficiencies can significantly increase carbon emissions and energy consumptions. Model compression techniques have been employed to enable AI inference in resource-constrained environments. For example, Knowledge Distillation (KD) is an elaborate approach targeting low-footprint and high-accuracy models, although introducing further complexity during training due to inefficient grid searches of additional hyperparameters. The emerging wave of AIoT, however, calls for prioritizing energy-awareness both in inference and training. To address this shortcoming, this work proposes a three-stage design workflow for low-energy AIoT applications, driven primarily by an input energy budget characterizing the target IoT scenario. Given a specific CNN architecture and IoT platform, our workflow identifies the most effective student under the imposed energy constrained and derives an efficient configuration of the KD hyperparameters that maximizes student accuracy, while avoiding inefficient and expensive grid-search. Hence, this approach enable energy-efficient CNN inference while substantially reducing overall training costs. We validate our workflow with a systematic experimental campaign using ResNets and DenseNets on CIFAR-10, CIFAR-100,and Tiny Imagenet datasets, on an AMD Xilinx Zynq Ultrascale+ ZCU102 MPSoC. Our proposal maintains high accuracy while lowering energy consumption by up to 80%, highlighting the potential of our flow for real-world AIoT applications. Franca Rocco di Torrepadula, Vincenzo Maisto, Alessandro Cilardo, Nicola Mazzocca |
J. Syst. Archit. | 2 |
| 2026 | The Simply-V Framework: An Extensible RISC-V Reconfigurable Soft-SoC for Open Research and Fast PrototypingabstractThe recent rise of open hardware, mainly driven by the momentum of the RISC-V ecosystem, has sparked significant innovation in the development of open-source CPUs and SoCs. This movement has enabled broad exploration across academia and industry, fostering collaboration and reuse. However, the diversity and openness that empower this space also introduce challenges: academic projects often fall short of industry-grade robustness, and meaningful comparison across hardware platforms remains difficult due to ad hoc infrastructures, lack of standardization, and simulation limitations. To ease the work of researchers some key challenges must be faced in open hardware development: platforms’ reconfigurability, ease of integration of third-party IPs, and support for technological heterogeneity. A core problem lies in validating and comparing CPUs and SoC components across varying protocols, toolchains, and design languages, especially in real hardware settings. To address these issues, we present Simply-V, a flexible FPGA-based soft-SoC platform designed for rapid prototyping and open hardware research. Simply-V enables plug-and-play support for multiple CPUs, IPs and accelerators, offers structured configurability across embedded and high-performance profiles, and supports the integration of both RTL and HLS-based components. Capabilities such as a high-level configuration flow, frequency scaling, and cross-device portability make our platform a powerful tool to simplify open hardware research. We demonstrate the SoC generator’s capabilities through multi-task FreeRTOS examples, platform-fair CPU benchmarking and the iterative development of HLS-designed convolutional accelerators. Moreover, we validate multi-accelerator and multi-CPU scalability and compare with the state-of-the-art SoC generators. Our platform showcases simplified fast prototyping, configurability, scalability and heterogeneous IP support on real hardware. Simply-V is openly available at https://github.com/HiSA-Team/Simply-V . Vincenzo Maisto, Stefano Mercogliano, Manuel Maddaluno, Alessandro Cilardo |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | An Approach to the Systematic Characterization of Multitask Accelerated CNN Inference in Edge MPSoCsabstractDeep Learning is ubiquitous today and is increasingly moving from the cloud down to the edge of networked infrastructures, where it enables embedded applications to perform complex inference tasks close to the data sources, reducing long-distance data movement and alleviating the need for a powerful cloud infrastructure. Edge-class multi-processor system on chip (MPSoC) devices featuring an on-chip FPGA fabric offer key advantages for Deep Learning inference tasks, especially for complex applications where multiple models may be run concurrently in the same platform. In this work, we propose an approach and a practical framework for the systematic characterization of multithreaded Deep Learning inference on edge FPGA MPSoCs. We instantiate the framework into a real-world MPSoC platform, targeting Xilinx Vitis-AI as a representative example of a commercial Deep Learning acceleration toolkit for edge environments. We design a comprehensive experimental campaign and apply it to the platform for several convolutional neural networks, each trained on three different datasets. We show that our approach can be used for both hardware- and software-level analysis of a target system. Among other findings, the analysis revealed a suboptimal behavior of the underlying toolkit runtime, involving the utilization of the accelerator cores and the uneven software latency of the support library, influenced by the shapes of the input tensors. Alessandro Cilardo, Vincenzo Maisto, Nicola Mazzocca, Franca Rocco di Torrepadula |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | A Pluggable Vector Unit for RISC-V Vector ExtensionabstractVector extensions have become increasingly important for accelerating data-parallel applications in areas like multimedia, data-streaming, and Machine Learning. This interactive presentation in-troduces a microarchitectural design of a vector unit compliant with the RISC- V vector extension v1.0. While we targeted a specific core for demonstration, CVA6, our architecture is designed so as to ensure extensibility, maintainability, and re-usability in other cores. Furthermore, as a distinctive feature, we support speculative execution and precise vector traps. The paper provides an overview of the main motivation, design choices, and implementation details, followed by a qualitative and quantitative discussion of the results collected from the synthesis of the extended CVA6 RISC-V core. Vincenzo Maisto, Alessandro Cilardo |
DATE | 1 |