Enrico Russo 0002

dblp:35/11048-2 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-7598-146XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 POSTER: Reinforcement Learning-based QoS-aware Online Scheduling for Multi-Tenant DNN Inference on Heterogeneous Accelerators
abstract
Deep Neural Networks (DNNs) are increasingly deployed through cloud services such as Inference-as-a-Service, where multiple tenants submit deadline-constrained inference requests to shared hardware infrastructures.
Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti
CF2
2026 Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI Workloads
abstract
Energy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts.
Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci
DATE12
2026 Timechain-level modeling and analysis of the bitcoin lightning network
Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi
Comput. Networks3
2026 Assessing the Role of Communication in Modular Multi-Core Quantum Systems
abstract
The scalability of quantum computing is constrained by the physical and architectural limitations of monolithic quantum processors. Modular multi-core quantum architectures, which interconnect multiple quantum cores (QCs) via classical and quantum-coherent links, offer a promising alternative to address these challenges. However, transitioning to a modular architecture introduces communication overhead, where classical communication plays a crucial role in executing quantum algorithms by transmitting measurement outcomes and synchronizing operations across QCs. Understanding the impact of classical communication on execution time is therefore essential for optimizing system performance. In this work, we introduce qcomm , an open-source simulator designed to evaluate the role of classical communication in modular quantum computing architectures. qcomm provides a high-level execution and timing model that captures the interplay between quantum gate execution, entanglement distribution, teleportation protocols, and classical communication latency. We conduct an extensive experimental analysis to quantify the impact of classical communication bandwidth, interconnect types, and quantum circuit mapping strategies on overall execution time. Furthermore, we assess classical communication overhead when executing real quantum benchmarks mapped onto a cryogenically-controlled multi-core quantum system. Our results show that, while classical communication is generally not the dominant contributor to execution time, its impact becomes increasingly relevant in optimized scenarios—such as improved quantum technology, large-scale interconnects, or communication-aware circuit mappings. These findings provide useful insights for the design of scalable modular quantum architectures and highlight the importance of evaluating classical communication as a performance-limiting factor in future systems.
Maurizio Palesi, Enrico Russo 0002, Giuseppe Ascia, Hamaad Rafique, Davide Patti, Vincenzo Catania, Sergi Abadal, Abhijit Das 0002, Pau Escofet, Eduard Alarcón, Carmen G. Almudéver
ACM Trans. Design Autom. Electr. Syst.2
2025 A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
abstract
Graph Neural Networks (GNNs) have shown significant promise in various domains, such as recommendation systems, bioinformatics, and network analysis. However, the irregularity of graph data poses unique challenges for efficient computation, leading to the development of specialized GNN accelerator architectures that surpass traditional CPU and GPU performance. Despite this, the structural diversity of input graphs results in varying performance across different GNN accelerators, depending on their dataflows. This variability in performance due to differing dataflows and graph properties remains largely unexplored, limiting the adaptability of GNN accelerators. To address this, we propose a data-driven framework for dataflow-aware latency prediction in GNN inference. Our approach involves training regressors to predict the latency of executing specific graphs on particular dataflows, using simulations on synthetic graphs. Experimental results indicate that our regressors can predict the optimal dataflow for a given graph with up to 91.28% accuracy and a Mean Absolute Percentage Error (MAPE) of 3.78%. Additionally, we introduce an online scheduling algorithm that uses these regressors to enhance scheduling decisions. Our experiments demonstrate that this algorithm achieves up to 3.17× speedup in mean completion time and 6.26× speedup in mean execution time compared to the best feasible baseline across all datasets.
Pol Puigdemont, Enrico Russo 0002, Axel Wassington, Abhijit Das 0002, Sergi Abadal, Maurizio Palesi
ASP-DAC2
2025 Optimizing Qubit Assignment in Modular Quantum Systems via Attention-Based Deep Reinforcement Learning
abstract
Modular, distributed, and multi-core architectures are considered a promising solution for scaling quantum computing systems. Optimising communication is crucial to preserve quantum coherence. The compilation and mapping of quantum circuits should minimise state transfers while adhering to architec-tural constraints. To address this problem efficiently, we propose a novel approach using Reinforcement Learning (RL) to learn heuristics for a specific multi-core architecture. Our RL agent uses a Transformer encoder and Graph Neural Networks, encoding quantum circuits with self-attention and producing outputs via an attention-based pointer mechanism to match logical qubits with physical cores efficiently. Experimental results show our method outperform the baseline reducing by 28% inter-core communications for random circuits while minimising time-to-solution.
Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DATE1
2025 An Anomaly Detection Model for RISC-V in Automotive Applications: A Domain-Specific Accelerator Perspective
abstract
Early anomaly detection in automotive systems is crucial for enhancing user safety and enabling timely corrective actions, thereby minimizing the risks associated with system malfunctions. This paper presents an approach for implementing Artificial Intelligence (AI)-based algorithms for anomaly detection in the automotive domain, leveraging the RISC-V architecture in conjunction with Domain-Specific Accelerators (DSAs). By exploiting the efficiency of DSAs, the proposed system aims to achieve faster anomaly detection compared to traditional processing methods. A detailed comparison is conducted between the performance of executing the AI-based anomaly detection algorithm on the RISC-V core versus offloading it to an optimized hardware accelerator tailored to the specific AI model. The goal of this work is to provide valuable insights into the potential of RISCV and DSAs to enhance AI-driven safety mechanisms, contributing to the development of more reliable automotive systems.
Elio Vinciguerra, Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia
PDP2
2024 A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
abstract
Currently, there is a growing trend of outsourcing the execution of DNNs to cloud services. For service providers, managing multitenancy and ensuring high-quality service delivery, particularly in meeting stringent execution time constraints, assumes paramount importance, all while endeavoring to maintain cost-effectiveness. In this context, the utilization of heterogeneous multi-accelerator systems becomes increasingly relevant. This paper presents RELMAS, a low-overhead deep reinforcement learning algorithm designed for the online scheduling of DNNs in multi-tenant environments, taking into account the dataflow heterogeneity of accelerators and memory bandwidths contentions. By doing so, service providers can employ the most efficient scheduling policy for user requests, optimizing Service-Level-Agreement (SLA) satisfaction rates and enhancing hardware utilization. The application of RELMAS to a heterogeneous multi-accelerator system composed of various instances of Simba and Eyeriss sub-accelerators resulted in up to a 173% improvement in SLA satisfaction rate compared to state-of-the-art scheduling techniques across different workload scenarios, with less than a 1.5% energy overhead.
Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DAC2
2024 Abstracting Bitcoin Lightning Network Complexity with Ultraviolet
abstract
With this work, we introduce the concept of Timechain-level model, along with an open-source implementation (Ultraviolet), to abstract the complexity of the Lightning Network (LN) while still providing a vision of base layer events and protocol internals. After depicting how each element of the model is mapped into the LN architectural stack, we show a case study to demonstrate its usage in investigating large-scale scenarios for research, development and educational purposes. Finally, we present a comparison to properly contextualize our contribution to the current state of the art of LN modelling, highlighting the advancements introduced by a Timechain-level and future directions of research it opens.
Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi, Vincenzo Catania
ICBC3
2024 Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
abstract
This paper addresses the critical challenge of managing Quality of Service (QoS) in cloud services, focusing on the nuances of individual tenant expectations and varying Service Level Indicators (SLIs). It introduces a novel approach utilizing Deep Reinforcement Learning for tenant-specific QoS management in multi-tenant, multi-accelerator cloud environments. The chosen SLI, deadline hit rate, allows clients to tailor QoS for each service request. A novel online scheduling algorithm for Deep Neural Networks in multi-accelerator systems is proposed, with a focus on guaranteeing tenant-wise, model-specific QoS levels while considering real-time constraints.
Enrico Russo 0002, Francesco Giulio Blanco, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Vincenzo Catania
ISCAS1
2024 Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-Based Accelerators
abstract
The need to efficiently execute different Deep Neural Networks (DNNs) on the same computing platform, coupled with the requirement for easy scalability, makes Multi-Chip Module (MCM)-based accelerators a preferred design choice. Such an accelerator brings together heterogeneous sub-accelerators in the form of chiplets, interconnected by a Network-on-Package (NoP). This paper addresses the challenge of selecting the most suitable sub-accelerators, configuring them, determining their optimal placement in the NoP, and mapping the layers of a predetermined set of DNNs spatially and temporally. The objective is to minimise execution time and energy consumption during parallel execution while also minimising the overall cost, specifically the silicon area, of the accelerator.This paper presents MOHaM, a framework for multi-objective hardware-mapping co-optimisation for multi-DNN workloads on chiplet-based accelerators. MOHaM exploits a multi-objective evolutionary algorithm that has been specialised for the given problem by incorporating several customised genetic operators. MOHaM is evaluated against state-of-the-art Design Space Exploration (DSE) frameworks on different multi-DNN workload scenarios. The solutions discovered by MOHaM are Pareto optimal compared to those by the state-of-the-art. Specifically, MOHaM-generated accelerator designs can reduce latency by up to 96% and energy by up to 96.12%.
Abhijit Das 0002, Enrico Russo 0002, Maurizio Palesi
IEEE Trans. Computers2
2024 Correction to: The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi
J. Supercomput.2
2023 Memory-Aware DNN Algorithm-Hardware Mapping via Integer Linear Programming
abstract
Mapping a deep neural network (DNN) layer onto domain-specific accelerators can require an intractable number of choices regarding loop factorization, ordering, and spatial unrolling. Determining the optimal mapping that achieves the best figures in terms of latency and energy efficiency can be difficult due to the vast number of possible candidates that need to be exhaustively evaluated. Many techniques have been recently proposed for fast and efficient mapping space exploration; some of them adopt a black-box optimization approach, others make assumptions on the underlying accelerator memory hierarchy or require time-consuming model retraining. We propose an integer linear programming (ILP) approach and formulate a mathematical model, namely LEMON, that takes into account number of accesses to each buffer, energy costs and buffer bandwidths in the accelerator and is flexible enough to work with different memory hierarchies. Compared with state-of-the-art techniques, LEMON achieves up to 83% energy-delay product reduction when compared to another ILP-based approach (CoSA) and 27% when compared to a genetic algorithm approach (GAMMA).
Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Salvatore Monteleone, Vincenzo Catania
CF1
2023 Multiobjective End-to-End Design Space Exploration of Parameterized DNN Accelerators
abstract
Deep neural network (DNN) hardware accelerators enable the execution of complex DNN inferences on resource-constrained IoT devices. Inference performance and energy figures depend on how the DNN layers are mapped into the accelerator and how the architecture of the accelerator fits the variety of layers’ shapes of the actual DNN. The mapping determines the execution order of the operations, both temporally and spatially. Thus, selecting the best mapping that allows fitting the DNN model to the specific accelerator is of paramount importance to meet the strong constraints imposed by resource-scarce IoT platforms. Although several mapping space exploration techniques have been proposed in the literature, they are focused on determining the best mapping for a given layer, for a given architecture, and for optimizing a single objective. This article largely extends the scope of the exploration by considering the huge design space spanned by mapping related and architectural parameters, considering all the layers of the DNN, and optimizing multiple objectives simultaneously. We present EPOCA, end-to-end Pareto optimization of DNN accelerators, whose goal is to determine the accelerator’s architecture and the mapping for each layer that optimizes end-to-end and in a multiobjective fashion a set of conflicting design criteria. We assess EPOCA on different DNN models on a parameterized hardware accelerator designed for IoT applications and compare them with a state-of-the-art mapping space explorer, considering the area, inference latency, and inference energy as optimization metrics. We show that the set of Pareto solutions found by EPOCA provides the designer with a range of choices from which to select the best tradeoff with respect to the specific application.
Enrico Russo 0002, Maurizio Palesi, Davide Patti, Salvatore Monteleone, Giuseppe Ascia, Vincenzo Catania
IEEE Internet Things J.1
2023 The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi
J. Supercomput.2
2022 MEDEA: A Multi-objective Evolutionary Approach to DNN Hardware Mapping
abstract
Deep Neural Networks (DNNs) embedded domain-specific accelerators enable inference on resource-constrained devices. Making optimal design choices and efficiently scheduling neural network algorithms on these specialized architectures is challenging. Many choices can be made to schedule computation spatially and temporally on the accelerator. Each choice influences the access pattern to the buffers of the architectural hierarchy, affecting the energy and latency of the inference. Each mapping also requires specific buffer capacities and a number of spatial components instances that translate in different chip area occupation. The space of possible combinations, the mapping space, is so large that automatic tools are needed for its rapid ex-ploration and simulation. This work presents MEDEA, an open-source multi-objective evolutionary algorithm based approach to DNNs accelerator mapping space exploration. MEDEA leverages the Timeloop analytical cost model. Differently from the other schedulers that optimize towards a single objective, MEDEA allows deriving the Pareto set of mappings to optimize towards multiple, sometimes conflicting, objectives simultaneously. We found that solutions found by MEDEA dominates in most cases those found by state-of-the-art mappers.
Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DATE1
2022 DNN Model Compression for IoT Domain-Specific Hardware Accelerators
abstract
Machine learning techniques, particularly those based on neural networks, are always more often used at the edge of the network by Internet of Things (IoT) nodes. Unfortunately, the computation capabilities demanded by those applications, together with their energy efficiency-related constraints, exceed those exposed by embedded general-purpose processors. For this reason, the use of domain-specific hardware accelerators (DSAs) is considered the most viable solution to the unsustainable “Turing tariff” of general-purpose hardware. Starting from the observation that memory and communication traffic account for a large fraction of the overall latency and energy in deep neural network (DNN) inferences, this article proposes a new compression technique aimed at: 1) reducing the memory footprint for storing the model parameters of a DNN and 2) improving DNN inference latency and energy on resource-constrained IoT devices. The proposed compression technique, namely, LineCompress, is applied on a set of representative convolutional neural networks (CNNs) for object recognition mapped on a state-of-the-art DSA targeted for resource-constrained IoT devices. We show that on average,$7.4\times $memory footprint reduction can be obtained, thus reducing the memory and communication traffic that result to 77% and 87% inference latency and energy reduction, respectively, trading-off efficiency versus accuracy.
Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Andrea Mineo, Giuseppe Ascia, Vincenzo Catania
IEEE Internet Things J.1