Valeria Mele

dblp:19/11405 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-2643-3483ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 AgentAI: A comprehensive survey on autonomous agents in distributed AI for industry 4.0
abstract
AgentAI represents a transformative approach within distributed Artificial Intelligence (AI) in which autonomous agents work either individually or collaboratively in decentralized environments to address challenging problems. AgentAI enhances scalability, robustness, and flexibility by utilizing advanced communication, learning, and decision-making capabilities, making it integral to diverse applications in Industry 4.0. The ability of AI systems to interpret sensory data in open-world environments has seen significant advancements in recent years. This progress emphasizes the need to move beyond reductionist approaches and embrace more embodied and cohesive systems, which integrate foundational models into agent-driven actions. Existing surveys often focus on isolated domains or specific autonomy levels, lacking a cohesive analysis that spans the full spectrum of AgentAI development in Industry 4.0. This survey explicitly fills this gap by introducing a multi-domain taxonomy and by systematically analyzing both non-autonomous and fully autonomous AgentAI systems, offering a comprehensive synthesis not previously available in the literature. Additionally, the paper extends the discussion to Industry 5.0 and 6.0, exploring the evolution of AgentAI from automation to collaboration and, ultimately, to fully autonomous systems. This comprehensive analysis highlights the potential of AgentAI in driving industries toward a more efficient, sustainable, and adaptable future.
Francesco Piccialli, Diletta Chiaro, Sundas Sarwar, Donato Cerciello, Pian Qi, Valeria Mele
Expert Syst. Appl.6
2025 FLAIR: Federated Learning for Augmented Industrial Retrieval
abstract
Deep learning (DL) has significantly advanced Industry 4.0 by leveraging data from the Industrial Internet of Things (IIoT) to enable smart manufacturing, predictive maintenance, and data-driven product marketing. However, multimodal industrial data presents challenges for traditional frameworks, including scalability, data privacy, and integration efficiency. This paper introduces an efficient product retrieval framework for e-commerce systems, addressing privacy and performance challenges through federated learning (FL). Specifically, we propose FLAIR (Federated Learning for Augmented Industrial Retrieval), a novel part retrieval system where distributed warehouses collaboratively train a multimodal foundation model, CLIP (Contrastive Language-Image Pre-Training), by fine-tuning only the Adapter module via FL, ensuring data privacy and efficiency. To address the limited availability of multimodal industrial data, our framework incorporates effective data augmentation strategies to enhance the diversity and quality of the training dataset. Comprehensive experiments on the Industrial Language-Image Dataset (ILID) highlight that FLAIR holds effective privacy safeguards and strong retrieval capabilities. Additionally, an advanced e-commerce recommendation system built on FLAIR showcases its practical effectiveness. FLAIR represents the first application of FL for industrial product retrieval, optimizing part searches, inventory management, and customer experience while maintaining data security. The complete code is available at https://github.com/MODAL-UNINA/FLAIR.
Diletta Chiaro, Pian Qi, Valeria Mele, Francesco Piccialli
IEEE Internet Things J.3
2024 Generalized Ware-Amdhal Law
abstract
We herein briefly describe a multilevel approach to analyze parallel algorithms performances. The main outcome of such an approach is that the algorithm is described using a set of operators related to each other according to the problem decomposition. A set of block matrices (called decomposition and execution matrices) highlights fundamental characteristics of the algorithm, such as inherent parallelism and sources of overheads, and all the involved factors and their relationships, in a general but modular, flexible and adaptive fashion. The work aims to show how we can rewrite the well-known Ware-Amdhal's Law: in a previous work we already gave an expression for the Amdhal's law in our framework, but, even with the same meaning, it wasn't immediately comparable with the original one. Here we focus on that law and show how it comes exactly from our parameters. Moreover, we extend the law with a more general expression, that we call Generalized Amdhal's Law: the classical law will come as a particular case of the generalized one.
Valeria Mele, Diego Romano
PDP1
2024 Eco-FL: Enhancing Federated Learning sustainability in edge computing through energy-efficient client selection
abstract
In the realm of edge cloud computing (ECC), Federated Learning (FL) revolutionizes the decentralization of machine learning (ML) models by enabling their training across multiple devices. In this way, FL preserves privacy and minimizes the need for centralized data by processing data near the source. From a communication standpoint, only the model weights are exchanged between devices. By avoiding the need to send data to a centralized location for processing, FL reduces the energy required for data transfer and supports more efficient use of computing resources at the edge. FL is particularly advantageous for resource-constrained devices, such as smartphones and IoT devices. However, this limited computational power and battery capacity and the challenge of energy consumption are critical aspects of FL systems. This paper introduces Eco-FL, an innovative methodology designed to optimize energy consumption in FL systems, in the field of Green Edge Cloud Computing (GECC). Our approach employs a device selection process that considers the entropy of the data held by the devices and their available energy reserves. This ensures that devices with lower energy availability are less likely to participate in the training rounds, prioritizing those with higher energy capacities. To evaluate the efficacy of our methodology, we utilize FedEntropy, an entropy-based aggregation method, alongside established aggregation methods such as FedAvg and FedProx for performance comparison. The effectiveness of Eco-FL in reducing energy consumption without compromising the accuracy of the FL process is demonstrated through analyses conducted on three distinct datasets. These analyses vary the β parameter of the Dirichlet distribution and account for scenarios with both homogeneous and heterogeneous initial device charges. Our findings validate Eco-FL’s potential to enhance the sustainability of FL systems by judiciously managing client participation based on energy criteria, presenting a significant step forward in the development of energy-efficient FL.
Martina Savoia, Edoardo Prezioso, Valeria Mele, Francesco Piccialli
Comput. Commun.3
2024 Toward a new linpack-like benchmark for heterogeneous computing resources
abstract
Summary This work describes some first efforts to design a new Linpack‐like benchmark useful to evaluate the performance of Heterogeneous Computing Resources. The benchmark is based on the Schur Complement reformulation of the solution of a linear equation system. Details about its implementation and evaluation, mainly in terms of performance scalability, are presented for a computing environment based on multi NVIDIA GP‐GPUs nodes connected by an Infiniband network.
Luisa Carracciuolo, Valeria Mele, Gianluca Sabella
Concurr. Comput. Pract. Exp.2
2024 Special Issue on the pervasive nature of HPC (PN-HPC)
abstract
Summary This special issue on the Pervasive Nature of HPC (PN‐HPC) collects an extension of the most valuable works presented at the sixth Workshop on Models, Algorithms and Methodologies for Hybrid Parallelism in New HPC Systems (MAMHYP‐22), held in Gdansk (Poland) in September 2022, jointly with the 14th conference on Parallel Processing and Applied Mathematics (PPAM‐22). New original papers related to the workshop themes are also included. The final aim is to provide a glimpse of the current state of knowledge related to the development of efficient methodologies and algorithms for HPC systems with multiple forms of parallelism.
Marco Lapegna, Valeria Mele, Raffaele Montella, Lukasz Szustak
Concurr. Comput. Pract. Exp.2
2023 Machine Learning Insights for Behavioral Data Analysis Supporting the Autonomous Vehicles Scenario
abstract
The advent of the digital innovation era is changing service, use, and resources management paradigms, offering a wide range of new and essential opportunities. In particular, the advent of the Internet of Things (IoT), i.e., the ability to connect individual objects to the Internet, also capable of communicating autonomously, has its particular declination on the connected vehicle. It is combined with the potential of advanced sensors placed pervasively on vehicles, which offer multifunctional monitoring capabilities of the entire system: from individual components up to the whole vehicle, including driver behavior and conditions and many exogenous parameters to the vehicle (road and weather conditions, congestion, risk situations, changes to mobility plans, etc.). In this perspective, machine learning (ML) models can transform raw data into new knowledge; they can contribute in an innovative way to define and suggest decisions, strategies, and criteria for resource use. Nowadays, most intelligent mobility projects also integrate artificial intelligence (AI) and ML solutions. In this article, we present and discuss the application of unsupervised learning techniques on a vehicular IoT data set. The main goal is to generate new knowledge about a geographical zone by analyzing historical drivers behavioral data. The autonomous vehicle’s framework can exploit the generated valuable insights to optimize the routes and prevent critical issues.
Edoardo Prezioso, Fabio Giampaolo, Carlo Mazzocca, Armir Bujari, Valeria Mele, Flora Amato
IEEE Internet Things J.5
2022 Classification of urban functional zones through deep learning
Stefano Izzo, Edoardo Prezioso, Fabio Giampaolo, Valeria Mele, Vittorio Di Somma, Gang Mei
Neural Comput. Appl.4
2021 About the granularity portability of block-based Krylov methods in heterogeneous computing environments
abstract
Summary Large‐scale problems in engineering and science often require the solution of sparse linear algebra problems and the Krylov subspace iteration methods (KM) have led to a major change in how users deal with them. But, for these solvers to use extreme‐scale hardware efficiently a lot of work was spent to redesign both the KM algorithms and their implementations to address challenges like extreme concurrency, complex memory hierarchies, costly data movement, and heterogeneous node architectures. All the redesign approaches bases the KM algorithm on block‐based strategies which lead to the Block‐KM (BKM) algorithm which has high granularity (i.e., the ratio of computation time to communication time). The work proposes novel parallel revisitation of the modules used in BKM which are based on the overlapping of communication and computation. Such revisitation is evaluated by a model of their granularity and verified on the basis of a case study related to a classical problem from numerical linear algebra.
Luisa Carracciuolo, Valeria Mele, Lukasz Szustak
Concurr. Comput. Pract. Exp.2
2021 A scalable Kalman filter algorithm: Trustworthy analysis on constrained least square model
abstract
Summary Kalman filter (KF) is one of the most important and common estimation algorithms. We introduce an innovative designing of Kalman filter algorithm based on domain decomposition (we call it DD‐KF). DD‐KF involves decomposition of the whole computational problem, partitioning of the solution and a slight modification of KF algorithm allowing a correction at run‐time of local solutions. The resulted parallel algorithm consists of concurrent copies of KF algorithm, each one requiring the same amount of computations on each subdomain and an exchange of boundary conditions between adjacent subdomains. Main advantage of this approach is that it can be potentially applied in a moderately nonintrusive manner to existing codes for tracking and controlling systems in location, navigation, in computer graphics and in much more state estimation problems. To highlight the capability of DD‐KF of exploiting the computing power provided by future designs of microprocessors based on multi/many‐cores CPU/GPU technologies, we consider DD both at physical core level and at microprocessor level and we discuss scalability of DD‐KF algorithm at coarse and fine grained level. Throughout the present work, we derive and discuss DD‐KF algorithm for solving constrained least square model, which underlies any data sampling and estimation problem.
Luisa D'Amore, Rosalba Cacciapuoti, Valeria Mele
Concurr. Comput. Pract. Exp.3
2021 Special Issue on High-end Heterogeneous Architectures, Methodologies, and Algorithms (HHAMA20)
abstract
TEST 02 - Elsevier's Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields.
Sokol Kosta, Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella
Concurr. Comput. Pract. Exp.4
2020 Designing a GPU-parallel algorithm for raw SAR data compression: A focus on parallel performance estimation
Diego Romano, Marco Lapegna, Valeria Mele, Giuliano Laccetti
Future Gener. Comput. Syst.3
2020 Performance enhancement of a dynamic K-means algorithm through a parallel adaptive strategy on multicore CPUs
Giuliano Laccetti, Marco Lapegna, Valeria Mele, Diego Romano, Lukasz Szustak
J. Parallel Distributed Comput.3
2020 Correlation of Performance Optimizations and Energy Consumption for Stencil-Based Application on Intel Xeon Scalable Processors
abstract
This article provides a comprehensive study of the impact of performance optimizations on the energy efficiency of a real-world CFD application called MPDATA, as well as an insightful analysis of performance-energy interaction of these optimizations with the underlying hardware that represents the first generation of Intel Xeon Scalable processors. Considering the MPDATA iterative application as a use case, we explore the fundamentals of energy and performance analysis for a memory-bound application when exposed to a set of optimization steps that increase the application performance, by improving the operational intensity of code and utilizing resources more efficiently. It is shown that for memory-bound applications, optimizing toward high performance could be a powerful strategy for improving the energy efficiency as well. In fact, for the considered performance optimizations, the energy gain is correlated with the performance gain but with varying degrees. As a result, these optimizations allow improving both performance and energy consumption radically, up to about 10.9 and 8.8 times, respectively. The impact of the Intel AVX-512 SIMD extension on the energy consumption and performance is demonstrated. Also, we discover limitations on the usability of CPU frequency scaling as a tool for balancing energy savings with admissible performance losses.
Lukasz Szustak, Roman Wyrzykowski, Tomasz Olas, Valeria Mele
IEEE Trans. Parallel Distributed Syst.4
2019 An adaptive algorithm for high-dimensional integrals on heterogeneous CPU-GPU systems
abstract
Summary In this paper, we introduce an adaptive procedure for the numerical computation of a high‐dimensional integrals on HPC systems with heterogeneous nodes composed of multi‐core CPU and GPU devices. To this aim, we have integrated together two different approaches: a first one is in charge of a fair workload among the threads running on the multi‐core CPU, while a second one is in charge of an efficient execution of the computational kernels on the GPU. We tested the resulting algorithm on several test functions on a system where the nodes are provided with two Intel ten‐core CPU and one NVIDIA GPU device.
Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella
Concurr. Comput. Pract. Exp.3
2018 A PETSc parallel-in-time solver based on MGRIT algorithm
abstract
Summary We address the development of a modular implementation of the MGRIT (MultiGrid‐In‐Time) algorithm to solve linear and nonlinear systems that arise from the discretization of evolutionary models with a parallel‐in‐time approach in the context of the PETSc (the Portable, Extensible Toolkit for Scientific computing) library. Our aim is to give the opportunity of predicting the performance gain achievable when using the MGRIT approach instead of the Time Stepping integrator (TS). To this end, we analyze the performance parameters of the algorithm that provide a‐priori the best number of processing elements and grid levels to use to address the scaling of MGRIT, regarded as a parallel iterative algorithm proceeding along the time dimension.
Valeria Mele, Emil M. Constantinescu, Luisa Carracciuolo, Luisa D'Amore
Concurr. Comput. Pract. Exp.1
2014 Algorithm 946: ReLIADiff - A C++ Software Package for Real Laplace Transform Inversion based on Algorithmic Differentiation
abstract
Algorithm 662 of the ACM TOMS library is a software package, based on the Weeks method, which is used for calculating function values of the inverse Laplace transform. The software requires transform values at arbitrary points in the complex plane. We developed a software package, called ReLIADiff, which is a modification of Algorithm 662 using transform values at arbitrary points on real axis. ReLIADiff, implemented in C++, relies on TADIFF software package designed for Algorithmic Differentiation. In this article, we present ReLIADiff focusing on its design principles, performance, and use.
Luisa D'Amore, Rosanna Campagna, Valeria Mele, Almerico Murli
ACM Trans. Math. Softw.3