EDBT 2026 Demo / reviewers in the wild / expert
Marcello Barbirotta
dblp:275/7192
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1902-7188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Impact and Characterization of Single Event Failures in Hyperdimensional Computing Hardware Accelerator
Marcello Barbirotta, Marco Angioli, Rocco Martino, Federico Frontera, Antonio Mastrandrea, Mauro Olivieri |
ETS | 1 |
| 2026 | RAS Enhancement of ECC-Protected Vector Register File for the RISC-V Architecture via RERI-Compliant Interface
Marcello Barbirotta, Nicasio Canino, Giovanni Mazzini, Mauro Olivieri, Daniele Rossi 0001, Sergio Saponara |
IOLTS | 1 |
| 2026 | Configurable Hardware Acceleration for Hyperdimensional Computing Extension on RISC-VabstractHyperdimensional Computing (HDC) is a neuro-inspired computational model that represents and manipulates information using high-dimensional distributed representations that are combined and compared using simple and highly parallel vector operations. In this work, we present a highly flexible hardware acceleration unit for HDC learning tasks based on the binary spatter-code model. Integrated into the execution stage of the Klessydra-T03 RISC-V core, the unit accelerates the core arithmetic operations of HDC and can be configured at synthesis time in terms of hardware parallelism, supported operations, and size of the local memories, trading off execution time with hardware resources to match application needs. A custom RISC-V Instruction Set Extension efficiently controls the accelerator, with instructions fully integrated into the GNU Compiler Collection toolchain and exposed to the programmer as intrinsics. Dedicated Control and Status Registers allow specifying the characteristics of the high-dimensional space and the target learning tasks at runtime, controlling the hardware loops of the accelerator and enabling the same hardware architecture to be used for various tasks. The dual flexibility coming from hardware configuration and software programmability sets this work apart from application-specific solutions in the literature, offering a unique, versatile accelerator adaptable to a wide range of applications and learning tasks. Rocco Martino, Marco Angioli, Antonello Rosato, Marcello Barbirotta, Abdallah Cheikh, Mauro Olivieri |
IEEE Trans. Computers | 4 |
| 2025 | HD-CBBIN: A Lightweight Approach for Contextual Bandit Learning in Real-Time ApplicationsabstractAs the Internet of Things expands, the need to embed artificial intelligence algorithms into resource-constrained devices for real-time applications is growing. These systems require efficient and scalable algorithms to perform rapid and autonomous decision-making without relying on centralized cloud servers. Hyperdimensional Computing (HDC) has recently emerged as a compelling paradigm for learning tasks in such environments, offering computational efficiency, exceptional parallelism, and scalability.In this work, we present HD-CBBIN, a lightweight and efficient implementation of the HD-CB framework for modeling and automating sequential decision-making Contextual Bandits (CB) problems on embedded systems. By introducing modifications to the original algorithm, HD-CBBINexclusively uses binary hypervectors, significantly reducing computational demands. We benchmark the performance of HD-CBBINon synthetic datasets, comparing it to the real-valued counterpart and the traditional state-of-the-art LinUCB algorithm. We also evaluate its execution time, computational complexity, and memory requirements on various embedded platforms and demonstrate additional gains through hardware acceleration. The results show that our approach achieves linear execution time with respect to the context vector size and up to a 141× speedup over LinUCB while maintaining competitive performance, establishing a new milestone in contextual bandit algorithms for time-critical, resource-constrained applications. Marco Angioli, Antonello Rosato, Marcello Barbirotta, Rocco Martino, Andrea Marcelli, Antonio Mastrandrea, Mauro Olivieri |
IJCNN | 3 |
| 2025 | Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning SystemsabstractAs the Internet of Things expands, embedding Artificial Intelligence algorithms in resource-constrained devices has become increasingly important to enable real-time, autonomous decision-making without relying on centralized cloud servers. However, implementing and executing complex algorithms in embedded devices poses significant challenges due to limited computational power, memory, and energy resources. 2This article presents algorithmic and hardware techniques to efficiently implement two LinearUCB Contextual Bandits algorithms on resource-constrained embedded devices. Algorithmic modifications based on the Sherman–Morrison–Woodbury formula streamline model complexity, while vector acceleration is harnessed to speed up matrix operations. We analyze the impact of each optimization individually and then combine them in a two-pronged strategy. The results show notable improvements in execution time and energy consumption, demonstrating the effectiveness of combining algorithmic and hardware optimizations to enhance learning models for edge computing environments with low-power and real-time requirements. Marco Angioli, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Mauro Olivieri |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | AeneasHDC: An Automatic Framework for Deploying Hyperdimensional Computing Models on FPGAsabstractHyperdimensional Computing (HDC) is a bio-inspired learning paradigm, that models neural pattern activities using high-dimensional distributed representations. HDC leverages parallel and simple vector arithmetic operations to combine and compare different concepts, enabling cognitive and reasoning tasks. The computational efficiency and parallelism of this approach make it particularly suited for hardware implementations, especially as a lightweight, energy-efficient solution for performing learning tasks on resource-constrained edge devices. The HDC pipeline, including encoding, training, and comparison stages, has been extensively explored with various approaches in the literature. However, while these techniques are mainly oriented to improve the model accuracy, their influence on hardware parameters remains largely unexplored. This work presents AeneasHDC, an automatic and open-source platform for the streamlined deployment of HDC models in both software and hardware for classification, regression and clustering tasks. AeneasHDC supports an extensive range of techniques commonly adopted in literature, automates the design of flexible hardware accelerators for HDC, and empowers users to easily assess the impact of different design choices on model accuracy, memory usage, execution time, power consumption, and area requirements. Marco Angioli, Saeid Jamili, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Antonello Rosato, Mauro Olivieri |
IJCNN | 3 |
| 2024 | Design, Implementation and Evaluation of a New Variable Latency Integer Division SchemeabstractInteger division is key for various applications and often represents the performance bottleneck due to its inherent mathematical properties that limit its parallelization. This paper presents a new datadependent variable latency division algorithm, derived from the classic non-performing restoring method. The proposed technique exploits the relationship between the number of leading zeros in the divisor and in the partial remainder to dynamically detect and skip those iterations that result in a simple left shift. While a similar principle has been exploited in previous works, the proposed approach outperforms existing variable latency divider schemes in average latency and power consumption. We detail the algorithm and its implementation in four variants, offering versatility for the specific application requirements. For each variant, we report the average latency evaluated with different benchmarks, and we analyze the synthesis results for both FPGA and ASIC deployment, reporting clock speed, average execution time, hardware resources, and energy consumption, compared with existing fixed and variable latency dividers. Marco Angioli, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Saeid Jamili, Mauro Olivieri |
IEEE Trans. Computers | 2 |
| 2020 | ADMM Consensus for Deep LSTM NetworksabstractIn modern real-world applications, the need of using a decentralized data processing approach has progressively increased, facing complexity and handling issues. Pervasive data and ubiquitous computational capacity have enabled the proficient use of distributed implementation of machine learning algorithms, especially for forecasting problems. We provide in this paper a new, fully distributed prediction approach based on the Long Short-Term Memory deep neural network. When placed in a network of interconnected agents, the single predictors are able to improve the prediction accuracy by means of the Alternating Direction Method of Multipliers consensus procedure on some network parameters. Experimental tests on real-world time series prove the efficacy of the proposed approach, which regulates the information exchange in the network through high-level structures in the considered models. Antonello Rosato, Federico Succetti, Marcello Barbirotta, Massimo Panella |
IJCNN | 3 |