Eugenio Gianniti

dblp:180/5886 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
5since 2021 · last 2023
0000-0002-5647-4024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2023 Erratum to: A Hybrid Machine Learning Approach for Performance Modeling of Cloud-Based Big Data Applications
Ehsan Ataie, Athanasia Evangelinou, Eugenio Gianniti, Danilo Ardagna
Comput. J.3
2022 A Hybrid Machine Learning Approach for Performance Modeling of Cloud-Based Big Data Applications
abstract
Abstract Nowadays, Apache Hadoop and Apache Spark are two of the most prominent distributed solutions for processing big data applications on the market. Since in many cases these frameworks are adopted to support business critical activities, it is often important to predict with fair confidence the execution time of submitted applications, for instance when service-level agreements are established with end-users. In this work, we propose and validate a hybrid approach for the performance prediction of big data applications running on clouds, which exploits both analytical modeling and machine learning (ML) techniques and it is able to achieve a good accuracy without too many time consuming and costly experiments on a real setup. The experimental results show how the proposed approach attains improvement in accuracy, number of experiments to be run on the operational system and cost over applying ML techniques without any support from analytical models. Moreover, we compare our approach with Ernest, an ML-based technique proposed in the literature by the Spark inventors. Experiments show that Ernest can accurately estimate the performance in interpolating scenarios while it fails to predict the performance when configurations with increasing number of cores are considered. Finally, a comparison with a similar hybrid approach proposed in the literature demonstrates how our approach significantly reduce prediction errors especially when few experiments on the real system are performed.
Ehsan Ataie, Athanasia Evangelinou, Eugenio Gianniti, Danilo Ardagna
Comput. J.3
2022 Optimal Resource Allocation of Cloud-Based Spark Applications
abstract
Nowadays, the big data paradigm is consolidating its central position in the industry, as well as in society at large. Lots of applications, across disparate domains, operate on huge amounts of data and offer great advantages both for business and research. According to analysts, cloud computing adoption is steadily increasing to support big data analyses and Spark is expected to take a prominent market position for the next decade. As big data applications gain more and more importance over time and given the dynamic nature of cloud resources, it is fundamental to develop an intelligent resource management system to provide Quality of Service guarantees to end-users. This article presents a set of run-time optimization-based resource management policies for advanced big data analytics. Users submit Spark applications characterized by a priority and by a hard or soft deadline. Optimization policies address two scenarios: i) identification of the minimum capacity to run a Spark application within the deadline; ii) re-balance of the cloud resources in case of heavy load, minimising the weighted soft deadline application tardiness. The solution relies on an initial non-linear programming model formulation and a search space exploration based on simulation-optimization procedures. Spark application execution times are estimated by relying on a gamut of techniques, including machine learning, approximated analyses, and simulation. The benefits of the approach are evaluated on Microsoft Azure HDInsight and on a private cloud cluster based on POWER8 by considering the TPC-DS industry benchmark and SparkBench. The results obtained in the first scenario demonstrate that the percentage error of the prediction of the optimal resource usage with respect to system measurement and exhaustive search is in the range 4–29 percent while literature-based techniques present an average error in the range 6–63 percent. Moreover, in the second scenario, the proposed algorithms can address complex problems like computing the optimal redistribution of resources among tens of applications in less than a minute with an error of 8 percent on average. On the same considered tests, literature-based approaches obtain an average error of about 57 percent.
Marco Lattuada 0001, Enrico Barbierato, Eugenio Gianniti, Danilo Ardagna
IEEE Trans. Cloud Comput.3
2021 Optimizing Quality-Aware Big Data Applications in the Cloud
abstract
The last years witnessed a steep rise in data generation worldwide and, consequently, the widespread adoption of software solutions able to support data-intensive application. Competitiveness and innovation have strongly benefited from these new platforms and methodologies, and there is a great deal of interest around the new possibilities that Big Data analytics promise to make reality. Many companies currently engage in data-intensive processes as part of their core businesses; however, fully embracing the data-driven paradigm is still cumbersome, and establishing a production-ready, fine-tuned deployment is time-consuming, expensive, and resource-intensive. This situation calls for innovative models and techniques to streamline the process of deployment configuration for Big Data applications. In particular, the focus in this paper is on the rightsizing of Cloud deployed clusters, which represent a cost-effective alternative to installation on premises. This paper proposes a novel tool, integrated in a wider DevOps-inspired approach, implementing a parallel and distributed simulation-optimization technique that efficiently and effectively explores the space of alternative Cloud configurations, seeking the minimum cost deployment that satisfies quality of service constraints. The soundness of the proposed solution has been thoroughly validated in a vast experimental campaign encompassing different applications and Big Data platforms.
Eugenio Gianniti, Michele Ciavotta, Danilo Ardagna
IEEE Trans. Cloud Comput.1
2021 Predicting the performance of big data applications on the cloud
Danilo Ardagna, Enrico Barbierato, Eugenio Gianniti, Marco Gribaudo, Túlio B. M. Pinto, Ana Paula Couto da Silva, Jussara M. Almeida
J. Supercomput.3
2019 Machine Learning for Performance Prediction of Spark Cloud Applications
abstract
Big data applications and analytics are employed in many sectors for a variety of goals: improving customers satisfaction, predicting market behavior or improving processes in public health. These applications consist of complex software stacks that are often run on cloud systems. Predicting execution times is important for estimating the cost of cloud services and for effectively managing the underlying resources at runtime. Machine Learning (ML), providing black box solutions to model the relationship between application performance and system configuration without requiring in-detail knowledge of the system, has become a popular way of predicting the performance of big data applications. We investigate the cost-benefits of using supervised ML models for predicting the performance of applications on Spark, one of today's most widely used frameworks for big data analysis. We compare our approach with Ernest (an ML-based technique proposed in the literature by the Spark inventors) on a range of scenarios, application workloads, and cloud system configurations. Our experiments show that Ernest can accurately estimate the performance of very regular applications, but it fails when applications exhibit more irregular patterns and/or when extrapolating on bigger data set sizes. Results show that our models match or exceed Ernest's performance, sometimes enabling us to reduce the prediction error from 126-187% to only 5-19%.
Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida, Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna
CLOUD6
2019 Performance Prediction of GPU-based Deep Learning Applications
abstract
Recent years saw an increasing success in the application of deep learning methods across various domains and for tackling different problems, ranging from image recognition and classification to text processing and speech recognition. In this paper we propose and validate an approach to model the execution time for training convolutional neural networks (CNNs) deployed on GPGPUs. We demonstrate that our approach is generally applicable to a variety of CNN models and different types of G PG PU s with high accuracy, aiming at the preliminary design phases for system sizing.
Eugenio Gianniti, Li Zhang 0002, Danilo Ardagna
CLOSER1
2019 Gray-Box Models for Performance Assessment of Spark Applications
abstract
Big data applications are among the most suitable applications to be executed on cluster resources because of their high requirements of computational power and data storage.Correctly sizing the resources devoted to their execution does not guarantee they will be executed as expected.Nevertheless, their execution can be affected by perturbations which can change the expected execution time.Identifying when these types of issue occurred by comparing their actual execution time with the expected one is mandatory to identify potentially critical situations and to take the appropriate steps to prevent them.To fulfill this objective, accurate estimates are necessary.In this paper, machine learning techniques coupled with a posteriori knowledge are exploited to build performance estimation models.Experimental results show how the models built with the proposed approach are able to outperform a reference state-of-the-art method (i.e., Ernest method), reducing in some scenarios the error from the 221.09-167.07%to 13.15-30.58%.
Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna, Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida
CLOSER2
2019 Analytical composite performance models for Big Data applications
Soroush Karimian Aliabadi, Danilo Ardagna, Reza Entezari-Maleki, Eugenio Gianniti, Ali Movaghar-Rahimabadi
J. Netw. Comput. Appl.4
2018 Performance Prediction of GPU-Based Deep Learning Applications
abstract
Recent years saw an increasing success in the application of deep learning methods across various domains and for tackling different problems, ranging from image recognition and classification to text processing and speech recognition. In this paper we propose and validate an approach to model the execution time for training convolutional neural networks (CNNs) deployed on GPGPUs. We demonstrate that our approach is generally applicable to a variety of CNN models and different types of G PG PU s with high accuracy, aiming at the preliminary design phases for system sizing.
Eugenio Gianniti, Li Zhang 0002, Danilo Ardagna
SBAC-PAD1
2018 Performance Prediction of Cloud-Based Big Data Applications
abstract
Data heterogeneity and irregularity are key characteristics of big data applications that often overwhelm the existing software and hardware infrastructures. In such context, the exibility and elasticity provided by the cloud computing paradigm over a natural approach to cost-effectively adapting the allocated resources to the application's current needs. Yet, the same characteristics impose extra challenges to predicting the performance of cloud-based big data applications, a central step in proper management and planning. This paper explores two modeling approaches for performance prediction of cloud-based big data applications. We evaluate a queuing-based analytical model and a novel fast ad-hoc simulator in various scenarios based on different applications and infrastructure setups. Our results show that our approaches can predict average application execution times with 26% relative error in the very worst case and about 12% on average. Moreover, our simulator provides performance estimates 70 times faster than state of the art simulation tools.
Danilo Ardagna, Enrico Barbierato, Athanasia Evangelinou, Eugenio Gianniti, Marco Gribaudo, Túlio B. M. Pinto, Anna Guimarães, Ana Paula Couto da Silva, Jussara M. Almeida
ICPE4
2018 An optimization framework for the capacity allocation and admission control of MapReduce jobs in cloud systems
Marzieh Malekimajd, Danilo Ardagna, Michele Ciavotta, Eugenio Gianniti, Mauro Passacantando, Alessandro Maria Rizzi
J. Supercomput.4
2017 A Game-Theoretic Approach for Runtime Capacity Allocation in MapReduce
abstract
Nowadays many companies have available large amounts of raw, unstructured data. Among Big Data enabling technologies, a central place is held by the MapReduce framework and, in particular, by its open source implementation, Apache Hadoop. For cost effectiveness considerations, a common approach entails sharing server clusters among multiple users. The underlying infrastructure should provide every user with a fair share of computational resources, ensuring that service level agreements (SLAs) are met and avoiding wastes. In this paper we consider mathematical models for the optimal allocation of computational resources in a Hadoop 2.x cluster with the aim to develop new capacity allocation techniques that guarantee better performance in shared data centers. Our goal is to get a substantial reduction of power consumption while respecting the deadlines stated in the SLAs and avoiding penalties associated with job rejections. The core of this approach is a distributed algorithm for runtime capacity allocation, based on Game Theory models and techniques, that mimics the MapReduce dynamics by means of interacting players, namely the central Resource Manager and Class Managers.
Eugenio Gianniti, Danilo Ardagna, Michele Ciavotta, Mauro Passacantando
CCGrid1
2016 Modeling Performance of Hadoop Applications: A Journey from Queueing Networks to Stochastic Well Formed Nets
Danilo Ardagna, Simona Bernardi 0001, Eugenio Gianniti, Soroush Karimian Aliabadi, Diego Perez-Palacin, José Ignacio Requeno
ICA3PP3
2016 D-SPACE4Cloud: A Design Tool for Big Data Applications
Michele Ciavotta, Eugenio Gianniti, Danilo Ardagna
ICA3PP2