Federica Filippini

dblp:286/7745 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6549-924XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Distributed replica allocation and load balancing for Edge-Cloud FaaS
abstract
The Function-as-a-Service (FaaS) paradigm supports many Cloud-native applications, but rising demand for low-latency services exceeds what the Cloud alone can deliver. Edge computing addresses this limitation; however, its heterogeneity and fragmented administrative domains, together with the workload dynamicity, greatly complicate resource coordination and function replica allocation across the Edge-Cloud continuum. This paper addresses these challenges through two complementary contributions. First, we formalize the Function Replica Allocation and Load Balancing (FRALB) problem in the Edge-Cloud continuum as a Distributed orchestration problem referred to as DiFRALB. The formulation jointly optimizes function placement, horizontal offloading among Edge nodes, and vertical offloading toward Cloud resources. Computation offloading plays a central role in this setting, as it enables workloads to be distributed across heterogeneous Edge and Cloud infrastructures while meeting performance requirements. Second, recognizing that centralized orchestration approaches suffer from scalability limitations and raise privacy concerns in multi-stakeholder environments, we propose FaaS-MACrO, a distributed multi-agent orchestration architecture designed to solve the DiFRALB problem. In FaaS-MACrO, each Edge node operates as an independent agent making local decisions, while a lightweight coordinator ensures global consistency by iteratively updating offloading prices to resolve conflicts among neighboring nodes. Crucially, coordination requires only minimal information exchange, thereby preserving operational privacy. Our solution approach jointly optimizes processing and offloading decisions by capturing the trade-offs among local execution efficiency, horizontal offloading latency, and vertical offloading costs. Extensive experiments across heterogeneous node configurations, diverse network topologies, and varying function characteristics demonstrate that FaaS-MACrO achieves solutions within 0.03-12.14% of the centralized optimum on average while significantly improving scalability, reducing solution times by up to three orders of magnitude in large-scale deployments with 200 nodes.
Federica Filippini, Marin Lujak, Michele Ciavotta
J. Syst. Archit.1
2026 Tabular Reinforcement Learning Methods for Artificial Intelligence Tasks Offloading in Smart Eye-Wears
abstract
Virtual and Extended Reality technologies are increasingly adopted in fields such as healthcare, entertainment, and education. These applications heavily rely on Smart Eye-Wears (SEWs) and AI to provide users with new ways to perceive their environment. However, SEWs face limitations in computational power, memory, and battery life. Offloading computations to external servers is a prominent example of edge computation. However, this also presents considerable challenges due to delays caused by varying network conditions and server workloads. This article proposes self-adaptive techniques based on tabular reinforcement learning (RL) to optimize the offloading of Deep Neural Network tasks between the SEW, the user’s smartphone, and cloud servers. The goal is to maintain a high-quality user experience while minimizing energy consumption and 5G connection costs. We evaluated our framework under varying 5G and WiFi bandwidths and cloud latency. The results show that Q-learning, SARSA, and Expected SARSA achieve near-optimal policies, with Q-learning demonstrating superior performance in reducing execution time violations (approximately at 10%) and improving agent stability. Additionally, our approach offers a more favorable tradeoff between energy efficiency and execution time violations compared to two baseline methods. Real-system experiments reveal that the proposed solution can double SEW battery life with respect to local computation while maintaining a good quality of service, with only 11% execution time violations. These findings highlight the effectiveness of our approach in managing resources and enhancing the overall user experience in SEW AI applications.
Abednego Wamuhindo Kambale, Hamta Sedghani, Federica Filippini, Giacomo Verticale, Danilo Ardagna
ACM Trans. Auton. Adapt. Syst.3
2025 Federated Reinforcement Learning for Runtime Optimization of AI Applications in Smart Eyewears
Hamta Sedghani, Abednego Wamuhindo Kambale, Federica Filippini, Francesca Palermo, Diana Trojaniello, Danilo Ardagna
MASCOTS3
2025 OSCAR-P and aMLLibrary: Profiling and predicting the performance of FaaS-based applications in computing continua
abstract
This paper proposes an automated framework for efficient application profiling and training of Machine Learning (ML) performance models, composed of two parts: OSCAR-P and aMLLibrary. OSCAR-P is an auto-profiling tool designed to automatically test serverless application workflows running on multiple hardware and node combinations in cloud and edge environments. OSCAR-P obtains relevant profiling information on the execution time of the individual application components. These data are later used by aMLLibrary to train ML-based performance models. This makes it possible to predict the performance of applications on unseen configurations. We test our framework on clusters with different architectures (x86 and arm64) and workloads, considering multi-component use-case applications. This extensive experimental campaign proves the efficiency of OSCAR-P and aMLLibrary, significantly reducing the time needed for the application profiling, data collection, and data processing. The preliminary results obtained on the ML performance models accuracy show a Mean Absolute Percentage Error lower than 30% in all the considered scenarios.
Roberto Sala, Bruno Guindani, Enrico Galimberti, Federica Filippini, Hamta Sedghani, Danilo Ardagna, Sebastián Risco, Germán Moltó, Miguel Caballer
J. Syst. Softw.4
2025 Decentralized Edge Workload Forecasting With Gossip Learning
abstract
Edge computing has emerged as a crucial paradigm for addressing the growing demands of interconnected devices and large-scale mobile applications by relocating computation and storage services closer to end-users. Edge workloads are inherently volatile and challenging to forecast due to their dependence on factors such as human mobility patterns and geographically-distributed infrastructure, combined with the dynamic nature of edge nodes. Traditional centralized approaches to workload forecasting are inadequate in the context of decentralized and failure-prone edge environments. To address this challenge, this paper investigates workload forecasting using Gossip Learning (GL), an asynchronous peer-to-peer learning protocol. GL allows for the training of forecasting models in a fully-decentralized manner, thereby mitigating single point of failure risks and enhancing overall system robustness. We extended the original protocol across multiple dimensions to improve convergence, reduce communication overhead, and enhance resilience to failures. We evaluated the proposed approach through extensive simulations; the obtained results demonstrate its effectiveness with respect to classical methods, rendering it a promising solution to enhance load balancing and task offloading strategies at the edge, thereby ensuring Quality-of-Service (QoS) and reducing Service Level Agreement (SLA) violations.
Alessandro Tundo, Federica Filippini, Francesco Regonesi, Michele Ciavotta, Marco Savi
IEEE Trans. Netw. Serv. Manag.2
2024 Greening AI: A Framework for Energy-Aware Resource Allocation of ML Training Jobs with Performance Guarantees
Roberto Sala, Federica Filippini, Danilo Ardagna, Daniele Lezzi, Francesc Lordan, Patrick Thiem
AINA (5)2
2024 A Stochastic Approach for Scheduling AI Training Jobs in GPU-Based Systems
abstract
In this work, we optimize the scheduling of Deep Learning (DL) training jobs from the perspective of a Cloud Service Provider running a data center, which efficiently selects resources for the execution of each job to minimize the average energy consumption while satisfying time constraints. To model the problem, we first develop a Mixed-Integer Non-Linear Programming formulation. Unfortunately, the computation of an optimal solution is prohibitively expensive, and to overcome this difficulty, we design a heuristic STochastic Scheduler (STS). Exploiting the probability distribution of early termination, STS determines how to adapt the resource assignment during the execution of the jobs to minimize the expected energy cost while meeting the job due dates. The results of an extensive experimental evaluation show that STS guarantees significantly better results than other methods in the literature, effectively avoiding due date violations and yielding a percentage total cost reduction between 32% and 80% on average. We also prove the applicability of our method in real-world scenarios, as obtaining optimal schedules for systems of up to 100 nodes and 400 concurrent jobs requires less than 5 seconds. Finally, we evaluated the effectiveness of GPU sharing, i.e., running multiple jobs in a single GPU. The obtained results demonstrate that depending on the workload and GPU memory, this further reduces the energy cost by 17–29% on average.
Federica Filippini, Jonatha Anselmi, Danilo Ardagna, Bruno Gaujal
IEEE Trans. Cloud Comput.1
2024 SPACE4AI-D: A Design-Time Tool for AI Applications Resource Selection in Computing Continua
abstract
Nowadays, Artificial Intelligence (AI) applications are becoming increasingly popular in a wide range of industries, mainly thanks to Deep Neural Networks (DNNs) that needs powerful resources. Cloud computing is a promising approach to serve AI applications thanks to its high processing power, but this sometimes results in an unacceptable latency because of long-distance communication. Vice versa, edge computing is close to where data are generated and therefore it is becoming crucial for their timely, flexible, and secure management. Given the more distributed nature of the edge and the heterogeneity of its resources, efficient component placement and resource allocation approaches become critical in orchestrating the application execution. In this paper, we formulate the resource selection and AI applications component placement problem in a computing continuum as a Mixed Integer Non-Linear Problem (MINLP), and we propose a design-time tool for its efficient solution. We first propose a Random Greedy algorithm to minimize the cost of the placement while guaranteeing some response time performance constraints. Then, we develop some heuristic methods such as Local Search, Tabu Search, Simulated Annealing and Genetic Algorithms, to improve the initial solutions provided by the Random Greedy. To evaluate our proposed approach, we designed an extensive experimental campaign, comparing the heuristics methods with one another and then the best heuristic against Best Cost Performance Constraint (BCPC) algorithm, a state-of-the-art approach. The results demonstrate that our proposed approach finds lower-cost solution than BCPC (27.6% on average) under the same time limit in large-scale systems. Finally, during the validation in a real edge system including FaaS resources our approach finds the globally optimal solution, suffering a deviation of around 12% between actual and predicted costs.
Hamta Sedghani, Federica Filippini, Danilo Ardagna
IEEE Trans. Serv. Comput.2
2023 Performance Models for Distributed Deep Learning Training Jobs on Ray
abstract
Deep Learning applications are pervasive today, and efficient strategies are designed to reduce the computational time and resource demand of the training process. The Distributed Deep Learning (DDL) paradigm yields a significant speed-up by partitioning the training into multiple, parallel tasks. The Ray framework supports DDL applications exploiting data parallelism by enhancing the scalability with minimal user effort. This work aims at evaluating the performance of DDL training applications, by profiling their execution on a Ray cluster and developing Machine Learning-based models to predict the training time when changing the dataset size, the number of parallel workers and the amount of computational resources. Such performance-prediction models are crucial to forecast computational resources usage and costs in Cloud environments. Experimental results prove that our models achieve average prediction errors between 3 and 15% for both interpolation and extrapolation, thus demonstrating their applicability to unforeseen scenarios.
Federica Filippini, Boris Lublinsky, Maximilien de Bayser, Danilo Ardagna
SEAA1
2023 A Path Relinking Method for the Joint Online Scheduling and Capacity Allocation of DL Training Workloads in GPU as a Service Systems
abstract
The Deep Learning (DL) paradigm gained remarkable popularity in recent years. DL models are used to tackle increasingly complex problems, making the training process require considerable computational power. The parallel computing capabilities offered by modern GPUs partially fulfill this need, but the high costs related to GPU as a Service solutions in the cloud call for efficient capacity planning and job scheduling algorithms to reduce operational costs via resource sharing. In this work, we jointly address the online capacity planning and job scheduling problems from the perspective of cloud end-users. We present a Mixed Integer Linear Programming (MILP) formulation, and a path relinking-based method aiming at optimizing operational costs by (i) rightsizing Virtual Machine (VM) capacity at each node, (ii) partitioning the set of GPUs among multiple concurrent jobs on the same VM, and (iii) determining a due-date-aware job schedule. An extensive experimental campaign attests the effectiveness of the proposed approach in practical scenarios: costs savings up to 97% are attained compared with first-principle methods based on, e.g., Earliest Deadline First, cost reductions up to 20% are obtained with respect to a previously proposed Hierarchical Method and up to 95% against a dynamic programming-based method from the literature. Scalability analyses show that systems with up to 100 nodes and 450 concurrent jobs can be managed in less than 7 seconds. The validation in a prototype cloud environment shows a deviation below 5% between real and predicted costs.
Federica Filippini, Marco Lattuada 0001, Michele Ciavotta, Arezoo Jahani, Danilo Ardagna, Edoardo Amaldi
IEEE Trans. Serv. Comput.1
2021 A Randomized Greedy Method for AI Applications Component Placement and Resource Selection in Computing Continua
abstract
Artificial Intelligence (AI) and Deep Learning (DL) are pervasive today, with applications spanning from personal assistants to healthcare. Nowadays, the accelerated migration towards mobile computing and Internet of Things, where a huge amount of data is generated by widespread end devices, is determining the rise of the edge computing paradigm, where computing resources are distributed among devices with highly heterogeneous capacities. In this fragmented scenario, efficient component placement and resource allocation algorithms are crucial to orchestrate at best the computing continuum resources. In this paper, we propose a tool to effectively address the component placement problem for AI applications at design time. Through a randomized greedy algorithm, it identifies the placement of minimum cost providing performance guarantees across heterogeneous resources including edge devices, cloud GPU-based Virtual Machines and Function as a Service solutions.
Hamta Sedghani, Federica Filippini, Danilo Ardagna
JCC2