Emilio Paolini

dblp:325/4078 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0001-7486-9268ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Federated Intrusion Detection with Key-Value Cached Transformers in 5G RANs
Andrea Di Matteo, Emilio Paolini, Luca Valcarenghi, Nicola Andriolli
HPSR2
2026 Incremental Vision Transformers for Zero-Day Attack Detection in NextG Wireless Networks
Andrea Di Matteo, Emilio Paolini, Luca Valcarenghi, Nicola Andriolli
WCNC2
2026 From packets to predictions on GPU: Accelerated graph-based intrusion detection system
abstract
• From Packets to Predictions On GPU: accelerated Graph-based Intrusion Detection System Ahmed Salah Tawfik Ibrahim, Emilio Paolini, Filippo Cugini, Francesco Paolucci This manuscript presents an In-GPU GNN-based intrusion detection system. Below, we summarize the novel contributions introduced in this work: • End-to-End In-GPU Pipeline - Problem: CPU-side graph construction and host-device memory transfers dominate prediction latency in GNN-based IDS pipelines, hindering real-time deployment. - New contribution: We redesign node aggregation and adjacency generation as CUDA kernels and execute both graph construction and GNN inference entirely on the GPU. This eliminates copy overheads and exploits thread-level parallelism to reduce end-to-end latency while preserving detection accuracy. • Reduced Memory Requirements - Problem: Previous GNN-based intrusion detection systems are memory-hungry. - New contribution: In this version, this problem is mitigated by exploiting the sparse properties of the input graph to reduce the memory size required for the system to run. • Optimized Memory Access Pattern - Problem: GPU memory accesses patterns are typically not considered. - New contribution: Accesses to GPU global memory are optimized achieving better performance and lower execution time especially after reducing the memory requirements. Graph Neural Networks (GNNs) are effective in detecting cyberattacks thanks to their ability to model network traffic, capturing structural relationships within the network. However, the high latency of the graph construction phase directly translates into higher overall prediction time, posing a challenge to real-time deployment. Therefore, an optimized GPU-accelerated framework that leverages the structural properties of the traffic graph is proposed in this work. It builds upon the notion of a precomputed adjacency matrix that gets modified by each graph instance. GPU threads are further used to construct the nodes, achieving an end-to-end graph generation and inference fully within the GPU. The sparsity of the graph and the coalesced memory access pattern within the GPU threads are further exploited to optimize the process. This allows the GPU to build large graphs much faster than the CPU without affecting the classification accuracy. This speedup is prominent for large graphs of 700 packets with the GPU being almost 4.3 times faster, highlighting the effectiveness of the proposed approach in real-time deployments.
Ahmed Salah Tawfik Ibrahim, Emilio Paolini, Filippo Cugini, Francesco Paolucci
Comput. Networks2
2026 End-to-end latency assurance for distributed augmented reality over programmable 6G networks: A DESIRE6G demonstration
abstract
6G networks are expected to deliver ultra-low latency, high reliability, and real-time intelligence for emerging services such as interactive Augmented Reality (AR), autonomous robotics, and digital twins. Achieving these requirements in practice demands tight coordination between networking, computing, and control domains, spanning RAN, transport, edge, and cloud. However, current 5G deployments lack pervasive telemetry, fine-grained observability, and automated control mechanisms capable of reacting at the time scales required by latency-sensitive applications. This paper presents a full integrated demonstration of DESIRE6G, a cloud-native 6G-ready architecture that leverages programmable data planes with P4 for flexible routing and telemetry using an implementation of novel data plane protocols, achieves distributed optimization of service deployment and runtime monitoring and reconfiguration via secure multi-agent systems (MAS), combined with intent-based orchestration layer for end-to-end service assurance. The system is validated on the federated ARNO testbed using a real distributed AR application involving a remotely-controlled drone as a User-Equipment that is equipped with a camera streaming a live video through the DESIRE6G network to a Kubernetes edge cluster that executes serverless inference functions for object detection and recognition, the video is then augmented with object information and shown on a Quest 3 AR headset. The MAS monitors the end-to-end latency in real time through P4 Telemetry and responds to changes in network conditions by reconfiguring the affected segments, while the Kubernetes monitoring provides real-time visibility and scalability across different segments. Overall, three hierarchical service assurance loops are demonstrated: (i) In-Network Control (INC) executing microsecond-scale congestion recovery in the P4 data plane, (ii) Infrastructure Management Layer (IML) performing millisecond-scale function migration and scaling, and (iii) MAS-driven cross-domain optimization operating at sub-second time scales to resolve RAN latency anomalies. Evaluation results show stable end-to-end latency in the 15–25 ms range in steady-state conditions, with fast recovery during induced congestion ≤ 1 . 5 ms data plane reroute via P4 INC.
Francesco Paolucci, Emilio Paolini, Faris Alhamed, Massimo Satler, Domenico Uomo, Michelangelo Guaitolini, Pol González, Marc Ruiz 0001, Luis Velasco 0001, Sándor Laki, Dávid Kis, Gergely Pongrácz, Attila Mihály, Anestis Dalgkitsis, Chrysa Papagianni, Anastassios Nanos, Stephen Parker, Vincent Lefebvre, M. Angoustures, Juan Jose Vegas Olmos, Andrea Sgambelluri
Comput. Networks2
2026 Noise-resilient photonic neural networks through adaptive quantization
abstract
Abstract Photonic neural networks have emerged as a promising solution to overcome limitations of traditional hardware for neuromorphic computations, offering advantages in bandwidth, latency, and power efficiency. However, their performance is constrained by the limited precision of analog photonic computing, which is affected by inherent noise sources such as thermal and shot noise, and distortions. These effects degrade the photonic neural network performance, reducing the bit resolution achievable in photonic hardware typically to 2–4 bits. Traditional quantization strategies fail to account for these noise contributions, resulting in a substantial accuracy loss during inference. This paper introduces an adaptive quantization method called Adaptive-Quantization Photonic-Aware Neural Network (AQ-PANN) to address the challenges posed by different noise sources in analog photonic hardware. The proposed method uses a learnable step size quantization scheme to achieve high accuracy and stability under varying noise levels, introducing a scheme that unifies quantization step adaptation with noise injection exactly where photonic distortions arise. This design incurs only a minor training-time overhead, as it involves learning a small number of per-layer quantization step sizes and does not affect inference. Experimental evaluations on three commonly used test datasets (MNIST, SVHN, and CIFAR-10) with different bit resolutions show the robustness of AQ-PANN. On MNIST, an accuracy drop of only 2% was observed from low to high noise levels in a 4-bit configuration, while traditional DoReFa quantization suffered a 29% drop. For the SVHN dataset, AQ-PANN obtained a mean accuracy of 92% under high noise with 4-bit quantization, outperforming DoReFa by over 45%. On CIFAR-10, AQ-PANN maintained close to 60% accuracy under high noise in the 4-bit configuration, whereas DoReFa and PACT both collapsed below 40%. These results highlight the effectiveness of AQ-PANN in sustaining model performance across different noise intensities, enabling practical photonic neural network deployment.
Emilio Paolini, Lorenzo De Marinis, Peter Seigo Kincaid, Luca Valcarenghi, Giampiero Contestabile, Ioannis Roumpos, Miltiadis Moralis-Pegios, Nikos Pleros, Nicola Andriolli
Neural Comput. Appl.1
2026 Programmable In-Network Aggregation for Communication-Aware Federated Learning in 5G RANs
abstract
Federated Learning (FL) enables collaborative model training without sharing raw data, making it attractive for privacy-preserving applications at the wireless edge. However, when executed over real 5G networks, FL performance degrades due to uplink congestion, heterogeneous client capabilities, and intermittent connectivity. Most existing approaches attempt to mitigate these issues indirectly by optimizing clients (through adaptive participation, local training, or selection strategies) or by optimizing models (via pruning, quantization, or compression), but they ignore potential network bottlenecks. This paper introduces FLAG, an FL architecture that embeds innetwork aggregation directly into 5G gNodeBs, transforming the network into an active participant in the learning process. In particular, FLAG performs parameter aggregation at line rate within the 5G Service Data Adaptation Protocol layer and incorporates three mechanisms: Partial-Contribution Correction for loss-tolerant averaging, a timer-driven pipeline for real-time scheduling, and a deadline-based grouping strategy to mitigate stragglers. Experiments with realistic wireless emulation show that FLAG achieves up to 5.1× faster time-to-accuracy and maintains accuracy within 0.8% of a loss-free baseline, while reducing gNB-to-server bandwidth by aggregating pergNB rather than per-client. FLAG requires no modifications to clients or the parameter server, demonstrating how 5G-aware system design can make federated learning scalable, efficient, and resilient under real-world wireless conditions.
Emilio Paolini, Andrea Pinto, Luca Valcarenghi, Flavio Esposito
IEEE Trans. Netw. Serv. Manag.1
2025 Demo: Design and Implementation of Hierarchical Cross-Domain Orchestration Using TeraFlowSDN
abstract
This demonstration showcases the autonomous creation of optical lightpaths across two geographically optical network testbeds, using TeraFlowSDN (TFS) as a high-level intent-based orchestrator for optical service provisioning. Unlike existing approaches that create lightpaths independently within a single domain, our system highlights a hierarchical control model in which a centralized TFS instance coordinates two heterogeneous domain controllers: a vendor-specific controller at Politecnico di Milano (Italy) and a local TFS instance at Sant’Anna School of Advanced Studies (Italy). Southbound adapters enable the translation of high-level service intents into device-specific configurations, making it possible to integrate different controllers and vendors without modifying the underlying infrastructure. The live demo demonstrates automated provisioning of optical lightpaths triggered via a user-friendly graphical interface. The process includes endpoint discovery, transceiver selection, and lightpath establishment, all performed autonomously across multiple domains to support a video streaming service. This work demonstrates the novelty of hierarchical cross-domain orchestration, showing how TFS can unify multivendor environments under a single platform with minimal configuration overhead. This lays the groundwork for future developments in automated service provisioning, closed-loop control, and scalable cross-domain networking.
Anouar El Hachimi, Aryanaz Attarpour, Gabriele Nanni, Memedhe Ibrahimi, Sebastian Troia, Andrea Sgambelluri, Emilio Paolini, Massimo Tornatore, Francesco Musumeci 0001
CNSM7
2025 Flecto: Cross-Layer Adaptive Congestion Control with Reinforcement Learning
abstract
Effective congestion control is critical for wireless networks, where rapidly varying channel conditions and diverse traffic demands can severely degrade performance. Traditional congestion control algorithms rely on static heuristics that are often ill-suited for dynamic wireless environments. In this paper, we introduce Flecto, a Reinforcement Learning (RL)-based congestion control solution integrated into the QUIC protocol that, leveraging cross-layer metrics, including Signal-to-Noise Ratio, Block Error Rates, and Round-Trip Time measurements, can take decisions using a comprehensive view of network conditions. We implemented Flecto on a 5G testbed using OpenAirInterface and ETTUS USRP B210 radios, showing how it adapts transmission rates in real-time to maximize throughput and minimize latency while maintaining stability. Experimental results show that Flecto achieves an average throughput of 4539.5 KB/s approximately 6% higher both than Cubic (4267.2 KB/s) and New Reno (2674.1 KB/s) while reducing the average Round-Trip Time to 21.8 ms, significantly lower than Cubic’s 27.6 ms and New Reno’s 174.9 ms. These performance gains underscore the promise of integrating RL with cross-layer feedback for adaptive, efficient congestion control in next-generation wireless networks. Moreover, the modular design of Flecto facilitates its extension to other transport protocols and multi-user scheduling frameworks, paving the way for broader adoption in future wireless systems.
Cristiano Serra, Emilio Paolini, Roger Immich, Alessio Sacco, Guido Marchetto, Flavio Esposito
HPSR2
2024 Efficient Distributed Learning Over Lossy Wireless Networks
abstract
In the context of NextG Wireless Networks, addressing the challenges of wireless communication link reliability is paramount to ensure efficient Distributed Learning systems. However, many recent solutions have overlooked key challenges, such as packet-level losses and the impact of TCP retransmissions, which are crucial for the robustness of these systems. In this paper, we propose the integration of fountain codes into the distributed learning process to offer a robust mechanism to counteract packet loss. Specifically, we propose a cumulative strategy logic based on fountain codes specifically tailored for packet exchanges in Distributed Learning applications. Our evaluation shows that fountain codes significantly enhance the efficiency and reliability of distributed learning model updates under severe packet loss conditions, e.g., a packet reduction of ≈ 84% (≈ 60%) at the UE (gNB) side compared to traditional TCP methods when packet loss probability reaches 0.9 in Federated Learning context. However, under low packet loss scenarios, fountain codes computational overhead becomes non-negligible. These results highlight the potential of fountain codes to serve as a robust alternative to conventional communication protocols in distributed learning systems, particularly in environments characterized by unstable network conditions.
Emilio Paolini, Andrea Pinto, Luca Valcarenghi, Nicola Andriolli, Luca Maggiani, Flavio Esposito
CNSM1
2024 A Programmable 5G DU-RU SmartNIC based on MPSoC FPGA
abstract
The adoption of disaggregated, virtualized, and open gNodeB in the next generation Radio Access Network offers benefits such as cost reduction and improved network performance. However, meeting specific 5G and beyond requirements, e.g., Ultra Reliable Low Latency Communications, requires offloading selected gNodeB functions onto accelerated hardware. This study proposes to implement a 5G Distributed Unit (DU)Radio Unit (RU) in a System-on-a-Programmable-Chip (SoPC) where the FFT of Orthogonal Frequency Division Multiplexing in uplink transmission is offloaded onto an FPGA. The proposed solution is programmable and pluggable, allowing flexibility in function implementation and integration into various devices. The experimental evaluation shows that the proposed SmartNIC-based DU-RU achieves $15 \times$ speedup in processing time when compared to a server’s CPU with FPGA-accelerated Low-PHY processing. In addition, it shows that employing highperformance off-chip memory leads to about $14 \%$ reduction in processing time compared to the use of an FPGA-internal block memory.
Abdelghani Bourenane, Emilio Paolini, Nicola Andriolli, Luca Valcarenghi
HPSR2
2024 Enabling Lightweight Federated Learning in NextG Wireless Networks
abstract
NextG wireless will heavily rely on Federated Learning (FL) applications to learn context-aware AI solutions from the massive amount of generated data. Ensuring the reliability of wireless links for such applications is paramount, especially for FL where packet loss can severely hamper performance and efficiency. Traditional approaches fall short under the high packet loss characteristics of wireless networks. This demo shows how the integration of Fountain Codes (FC) into the FL process can bring notable improvements in packet transmission efficiency, especially under high packet loss conditions.
Emilio Paolini, Luca Valcarenghi, Nicola Andriolli, Luca Maggiani, Flavio Esposito
NetSoft1
2024 Hierarchical Software-Defined Control for coordinated RAN and PON-based Transport Scaling
abstract
This demonstration shows the effectiveness of a hierarchical Software-Defined approach in coordinating Radio Access Network (RAN) and Passive Optical Network (PON)-based RAN transport to proactively scale virtual Distributed Units (vDUs) / Radio Units (RUs) and to jointly reconfigure the mid-haul transport.
Alessandro Pacini, Andrea Sgambelluri, Carlo Centofanti, Andrea Marotta, Emilio Paolini, Alessio Giorgetti, Luca Valcarenghi
NOMS5
2023 Cascaded Look Up Table Distillation of P4 Deep Neural Network Switches
abstract
In-network function offloading represents a key enabler of the SDN-based data plane programmability to enhance network operation and awareness while speeding up applications and reducing the energy footprint. The offload of network functions exploiting machine learning and artificial intelligence has been recently considered with intermediate solutions such as feature extraction acceleration and mixed architectures including AI-specific platforms (e.g., GPU, FPGA). Indeed, the P4 language enables the programmability of deep neural networks inside the pipelines of both software and hardware switches and NICs. However, programmable hardware pipeline chipsets suffer from significant computing capability limitations (e.g., missing arithmetic logic units, limited and slow stateful registers) preventing the plain programmability of a deep neural network (DNN) operating at wirespeed. This paper proposes an innovative knowledge distillation technique that maps a DNN into a cascade of lookup tables (i.e., flow tables) with limited entry size. The proposed mapping avoids stateful elements and maths operators, whose requirement prevented the deployment of DNNs within hardware switches up to now. The evaluation is carried out considering a cyber security use case targeting a DDoS mitigator network function, showing negligible impact due to the lossless mapping reduction and feature quantization.
Lorenzo De Marinis, Emilio Paolini, Rana Abubakar, Filippo Cugini, Francesco Paolucci
GLOBECOM2
2023 CHARLES: A C++ fixed-point library for Photonic-Aware Neural Networks
Emilio Paolini, Lorenzo De Marinis, Luca Maggiani, Marco Cococcioni, Nicola Andriolli
Neural Networks1
2022 Photonic-aware Neural Networks for Packet Classification in URLLC scenarios
abstract
Ultra Reliable Low Latency Communications (URLLC) scenarios require very low latency and high reliability, imposing an optimization of every aspect of 5G data processing, transmission, and networking. Artificial Intelligence (AI)-based tools can be helpful resources in this context, enhancing multiple functionalities, from network resource allocation to network security. In this paper we propose a solution placed at the next generation eNB (gNB)-Central Unit (CU) level, relying on Neural Networks (NNs), capable of classifying incoming packets. The developed system increases the security of 5G and B5G architectures, protecting the 5G Core (5GC) from potential attacks. To comply with URLLC requirements on latency, we present an architecture leveraging photonic hardware to speed-up NN computations. The proposed solution, namely Photonic-Aware Neural Network (PANN), complies with physical layer constraints raised by photonic analog computing and can achieve high throughput and time-of-flight latency. The classification performance of the devised PANN model has been assessed through simulation on the distilled Kitsune dataset, suited for 5G scenarios. The experiments proved that PANN significantly lowers the chance of transmitting malicious packets to the 5GC with a classification performance increasing with the bit resolution supported by the analog photonic physical layer.
Emilio Paolini, Federico Civerchia, Lorenzo De Marinis, Luca Valcarenghi, Luca Maggiani, Nicola Andriolli
HPSR1
2022 Photonic-aware neural networks
abstract
Abstract Photonics-based neural networks promise to outperform electronic counterparts, accelerating neural network computations while reducing power consumption and footprint. However, these solutions suffer from physical layer constraints arising from the underlying analog photonic hardware, impacting the resolution of computations (in terms of effective number of bits), requiring the use of positive-valued inputs, and imposing limitations in the fan-in and in the size of convolutional kernels. To abstract these constraints, in this paper we introduce the concept of Photonic-Aware Neural Network (PANN) architectures, i.e., deep neural network models aware of the photonic hardware constraints. Then, we devise PANN training schemes resorting to quantization strategies aimed to obtain the required neural network parameters in the fixed-point domain, compliant with the limited resolution of the underlying hardware. We finally carry out extensive simulations exploiting PANNs in image classification tasks on well-known datasets (MNIST, Fashion-MNIST, and Cifar-10) with varying bitwidths (i.e., 2, 4, and 6 bits). We consider two kernel sizes and two pooling schemes for each PANN model, exploiting $$2\times 2$$ 2 × 2 and $$3\times 3$$ 3 × 3 convolutional kernels, and max and average pooling, the latter more amenable to an optical implementation. $$3\times 3$$ 3 × 3 kernels perform better than $$2\times 2$$ 2 × 2 counterparts, while max and average pooling provide comparable results, with the latter performing better on MNIST and Cifar-10. The accuracy degradation due to the photonic hardware constraints is quite limited, especially on MNIST and Fashion-MNIST, demonstrating the feasibility of PANN approaches on computer vision tasks.
Emilio Paolini, Lorenzo De Marinis, Marco Cococcioni, Luca Valcarenghi, Luca Maggiani, Nicola Andriolli
Neural Comput. Appl.1