EDBT 2026 Demo / reviewers in the wild / expert
Marco Lattuada 0001
dblp:67/2653
· DBLP profile ↗
21ranked-venue papers
6as first author
7since 2021 · last 2023
0000-0003-0062-6049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | aMLLibrary: An AutoML Approach For Performance PredictionabstractaMLLibrary is an open-source, high-level Python package that allows the parallel building of multiple Machine Learning (ML) regression models. It is focused on performance modeling and supports several methods for feature engineering/selection and hyperparameter tuning. The library implements fault tolerance mechanisms to recover from system crashes, and only a simple declarative text file is required to launch a full experimental campaign for all required models. Its modular structure allows users to implement their own plugins and model-building wrappers and easily add them to the library. We test aMLLibrary on building the performance models of neural networks and image processing applications, with the best model produced often having less than 20% prediction error. Bruno Guindani, Marco Lattuada 0001, Danilo Ardagna |
ECMS | 2 |
| 2023 | µ-FF: On-Device Forward-Forward Training Algorithm for MicrocontrollersabstractDeliver intelligence into low-cost hardware e.g., Microcontroller Units (MCUs) for the realization of low-power tailored applications nowadays is an emerging research area. However, the training of deep learning models on embedded systems is still challenging mainly due to their low amount of memory, available energy, and computing power which significantly limit the complexity of the tasks that can be executed, thus making impossible use of traditional training algorithms such as backpropagation (BP). During these years techniques such as weights compression and quantization have emerged as solutions, but they only address the inference phase. Forward-Forward (FF) is a novel training algorithm that has been recently proposed as a possible alternative to BP when the available resources are limited. This is achieved by training the layers of a neural network separately, thus reducing the required energy and memory. In this paper, we propose µ-FF, a variation of the original FF which tackles the training process with a multivariate Ridge regression approach and allows to find closed-form solution by using the Mean Squared Error (MSE) as loss function. Such an approach does not use BP and does not need to compute gradients, thus saving memory and computing resources to enable the on-device training directly on MCUs of the STM32 family. Experimental results conducted on the Fashion-MNIST dataset demonstrate the effectiveness of the proposed approach in terms of memory and accuracy. Fabrizio De Vita, Rawan M. A. Nawaiseh, Dario Bruneo, Valeria Tomaselli, Marco Lattuada 0001, Mirko Falchetto |
SMARTCOMP | 5 |
| 2023 | A Path Relinking Method for the Joint Online Scheduling and Capacity Allocation of DL Training Workloads in GPU as a Service SystemsabstractThe Deep Learning (DL) paradigm gained remarkable popularity in recent years. DL models are used to tackle increasingly complex problems, making the training process require considerable computational power. The parallel computing capabilities offered by modern GPUs partially fulfill this need, but the high costs related to GPU as a Service solutions in the cloud call for efficient capacity planning and job scheduling algorithms to reduce operational costs via resource sharing. In this work, we jointly address the online capacity planning and job scheduling problems from the perspective of cloud end-users. We present a Mixed Integer Linear Programming (MILP) formulation, and a path relinking-based method aiming at optimizing operational costs by (i) rightsizing Virtual Machine (VM) capacity at each node, (ii) partitioning the set of GPUs among multiple concurrent jobs on the same VM, and (iii) determining a due-date-aware job schedule. An extensive experimental campaign attests the effectiveness of the proposed approach in practical scenarios: costs savings up to 97% are attained compared with first-principle methods based on, e.g., Earliest Deadline First, cost reductions up to 20% are obtained with respect to a previously proposed Hierarchical Method and up to 95% against a dynamic programming-based method from the literature. Scalability analyses show that systems with up to 100 nodes and 450 concurrent jobs can be managed in less than 7 seconds. The validation in a prototype cloud environment shows a deviation below 5% between real and predicted costs. Federica Filippini, Marco Lattuada 0001, Michele Ciavotta, Arezoo Jahani, Danilo Ardagna, Edoardo Amaldi |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Svelto: High-Level Synthesis of Multi-Threaded Accelerators for Graph AnalyticsabstractGraph analytics are an emerging class of irregular applications. Operating on very large datasets, they present unique behaviors, such as fine-grained, unpredictable memory accesses, and highly unbalanced task level parallelism, that make existing high-performance general-purpose processors or accelerators (e.g., GPUs) suboptimal. To address these issues, research and industry are developing a variety of custom accelerator designs for this application area, including solutions based on reconfigurable devices (Field Programmable Gate Arrays). These new approaches often employ High-Level Synthesis (HLS) to Speed up the development of the accelerators. In this paper, we propose a novel architecture template for the automatic generation of accelerators for graph analytics and irregular applications. The architecture template includes a dynamic task scheduling mechanism, a parallel array of accelerators that enables supporting task-level parallelism with context switching, and a related multi-channel memory interface that decouples communication from computation and provides support for fine-grained atomic memory operations. We discuss the integration of the architectural template in an HLS flow, presenting the necessary modifications to enable automatic generation of the custom architectures starting from OpenMP annotated code. We evaluate our approach first by synthesizing and exploring triangle counting, a common graph algorithm, and then by synthesizing custom designs for a set of graph database benchmark queries, representing series of graph pattern matching routines. We compare the synthesized accelerators with previous state-of-the-art methodologies for the synthesis of parallel architectures, showing that the proposed approach allows reducing resource usage by optimizing the number of accelerators replicas without any performance penalty. Marco Minutoli, Vito Giovanni Castellana, Nicola Saporetti, Stefano Devecchi, Marco Lattuada 0001, Pietro Fezzardi, Antonino Tumeo, Fabrizio Ferrandi |
IEEE Trans. Computers | 5 |
| 2022 | Architectural Design of Cloud Applications: A Performance-Aware Cost Minimization ApproachabstractCloud Computing has assumed a relevant role in the ICT, profoundly influencing the life-cycle of modern applications in the manner they are designed, developed, and deployed and operated. In this article, we tackle the problem of supporting the design-time analysis of Cloud applications to identify a cost-optimized strategy for allocating components onto Cloud Virtual Machine infrastructural services, taking performance requirements into account. We present an approach and a tool, SPACE4Cloud, that supports users in modeling the architecture of an application, in defining performance requirements as well as deployment constraints, and then in mapping each architecture component into a corresponding VM service, minimizing total costs. An optimization algorithm supports the mapping and determines the Cloud configuration that minimizes the execution costs of the application over a daily time horizon. The benefits of this approach are demonstrated in the context of an industrial case study. Furthermore, we show that SPACE4Cloud leads to a cost reduction up to 60 percent, when compared to a first-principle technique based on utilization thresholds, like the ones typically used in practice, and that our solution is able to solve large problem instances within a time frame compatible with a fast-paced design process (less than half an hour in the worst case). Finally, we show that SPACE4Cloud is suitable to model even microservice-based applications and to compute the corresponding optimized deployment configuration which is compared with a state-of-the art meta-heuristic alternative method, achieving savings between 21 and 85 percent. Michele Ciavotta, Giovanni Paolo Gibilisco, Danilo Ardagna, Elisabetta Di Nitto, Marco Lattuada 0001, Marcos Aurélio Almeida da Silva |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | Optimal Resource Allocation of Cloud-Based Spark ApplicationsabstractNowadays, the big data paradigm is consolidating its central position in the industry, as well as in society at large. Lots of applications, across disparate domains, operate on huge amounts of data and offer great advantages both for business and research. According to analysts, cloud computing adoption is steadily increasing to support big data analyses and Spark is expected to take a prominent market position for the next decade. As big data applications gain more and more importance over time and given the dynamic nature of cloud resources, it is fundamental to develop an intelligent resource management system to provide Quality of Service guarantees to end-users. This article presents a set of run-time optimization-based resource management policies for advanced big data analytics. Users submit Spark applications characterized by a priority and by a hard or soft deadline. Optimization policies address two scenarios: i) identification of the minimum capacity to run a Spark application within the deadline; ii) re-balance of the cloud resources in case of heavy load, minimising the weighted soft deadline application tardiness. The solution relies on an initial non-linear programming model formulation and a search space exploration based on simulation-optimization procedures. Spark application execution times are estimated by relying on a gamut of techniques, including machine learning, approximated analyses, and simulation. The benefits of the approach are evaluated on Microsoft Azure HDInsight and on a private cloud cluster based on POWER8 by considering the TPC-DS industry benchmark and SparkBench. The results obtained in the first scenario demonstrate that the percentage error of the prediction of the optimal resource usage with respect to system measurement and exhaustive search is in the range 4–29 percent while literature-based techniques present an average error in the range 6–63 percent. Moreover, in the second scenario, the proposed algorithms can address complex problems like computing the optimal redistribution of resources among tens of applications in less than a minute with an error of 8 percent on average. On the same considered tests, literature-based approaches obtain an average error of about 57 percent. Marco Lattuada 0001, Enrico Barbierato, Eugenio Gianniti, Danilo Ardagna |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | Invited: Bambu: an Open-Source Research Framework for the High-Level Synthesis of Complex ApplicationsabstractThis paper presents the open-source high-level synthesis (HLS) research framework Bambu. Bambu provides a research environment to experiment with new ideas across HLS, high-level verification and debugging, FPGA/ASIC design, design flow space exploration, and parallel hardware accelerator design. The tool accepts as input standard C/C++ specifications and compiler intermediate representations (IRs) coming from the well-known Clang/LLVM and GCC compilers. The broad spectrum and flexibility of input formats allow the electronic design automation (EDA) research community to explore and integrate new transformations and optimizations. The easily extendable modular framework already includes many optimizations and HLS benchmarks used to evaluate the QoR of the tool against existing approaches [1]. The integration with synthesis and verification backends (commercial and open-source) allows researchers to quickly test any new finding and easily obtain performance and resource usage metrics for a given application. Different FPGA devices are supported from several different vendors: AMD/Xilinx, Intel/Altera, Lattice Semiconductor, and NanoXplore. Finally, integration with the OpenRoad open-source end-to-end silicon compiler perfectly fits with the recent push towards open-source EDA. Fabrizio Ferrandi, Vito Giovanni Castellana, Serena Curzel, Pietro Fezzardi, Michele Fiorito, Marco Lattuada 0001, Marco Minutoli, Christian Pilato, Antonino Tumeo |
DAC | 6 |
| 2019 | Machine Learning for Performance Prediction of Spark Cloud ApplicationsabstractBig data applications and analytics are employed in many sectors for a variety of goals: improving customers satisfaction, predicting market behavior or improving processes in public health. These applications consist of complex software stacks that are often run on cloud systems. Predicting execution times is important for estimating the cost of cloud services and for effectively managing the underlying resources at runtime. Machine Learning (ML), providing black box solutions to model the relationship between application performance and system configuration without requiring in-detail knowledge of the system, has become a popular way of predicting the performance of big data applications. We investigate the cost-benefits of using supervised ML models for predicting the performance of applications on Spark, one of today's most widely used frameworks for big data analysis. We compare our approach with Ernest (an ML-based technique proposed in the literature by the Spark inventors) on a range of scenarios, application workloads, and cloud system configurations. Our experiments show that Ernest can accurately estimate the performance of very regular applications, but it fails when applications exhibit more irregular patterns and/or when extrapolating on bigger data set sizes. Results show that our models match or exceed Ernest's performance, sometimes enabling us to reduce the prediction error from 126-187% to only 5-19%. Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida, Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna |
CLOUD | 5 |
| 2019 | Software defined architectures for data analyticsabstractData analytics applications increasingly are complex workflows composed of phases with very different program behaviors (e.g., graph algorithms and machine learning, algorithms operating on sparse and dense data structures, etc). To reach the levels of efficiency required to process these workflows in real time, upcoming architectures will need to leverage even more workload specialization. If, at one end, we may find even more heterogenous processors composed by a myriad of specialized processing elements, at the other end we may see novel reconfigurable architectures, composed of sets of functional units and memories interconnected with (re)configurable on-chip networks, able to adapt dynamically to adapt the workload characteristics. Field Programmable Gate Arrays are more and more used for accelerating various workloads and, in particular, inferencing in machine learning, providing higher efficiency than other solutions. However, their fine-grained nature still leads to issues for the design software and still makes dynamic reconfiguration impractical. Future, more coarse-grained architectures could offer the features to execute diverse workloads at high efficiency while providing better reconfiguration mechanisms for dynamic adaptability. Nevertheless, we argue that the challenges for reconfigurable computing remain in the software. In this position paper, we describe a possible toolchain for reconfigurable architectures targeted at data analytics. Vito Giovanni Castellana, Marco Minutoli, Antonino Tumeo, Marco Lattuada 0001, Pietro Fezzardi, Fabrizio Ferrandi |
ASP-DAC | 4 |
| 2019 | Gray-Box Models for Performance Assessment of Spark ApplicationsabstractBig data applications are among the most suitable applications to be executed on cluster resources because of their high requirements of computational power and data storage.Correctly sizing the resources devoted to their execution does not guarantee they will be executed as expected.Nevertheless, their execution can be affected by perturbations which can change the expected execution time.Identifying when these types of issue occurred by comparing their actual execution time with the expected one is mandatory to identify potentially critical situations and to take the appropriate steps to prevent them.To fulfill this objective, accurate estimates are necessary.In this paper, machine learning techniques coupled with a posteriori knowledge are exploited to build performance estimation models.Experimental results show how the models built with the proposed approach are able to outperform a reference state-of-the-art method (i.e., Ernest method), reducing in some scenarios the error from the 221.09-167.07%to 13.15-30.58%. Marco Lattuada 0001, Eugenio Gianniti, Marjan Hosseini, Danilo Ardagna, Alexandre Maros, Fabricio Murai, Ana Paula Couto da Silva, Jussara M. Almeida |
CLOSER | 1 |
| 2019 | BIGSEA: A Big Data analytics platform for public transportation information
Andy S. Alic, Jussara M. Almeida, Giovanni Aloisio, Nazareno Andrade, Nuno Antunes, Danilo Ardagna, Rosa M. Badia, Tânia Basso, Ignacio Blanquer, Tarciso Braz, Andrey Brito, Donatello Elia, Sandro Fiore, Dorgival O. Guedes, Marco Lattuada 0001, Daniele Lezzi, Matheus Maciel, Wagner Meira Jr., Demetrio Gomes Mestre, Regina Lúcia de Oliveira Moraes, Fábio Morais 0001, Carlos Eduardo S. Pires, Nádia P. Kozievitch, Walter Santos, Paulo Silva 0002, Marco Vieira |
Future Gener. Comput. Syst. | 15 |
| 2019 | A Design Flow Engine for the Support of Customized Dynamic High Level Synthesis FlowsabstractHigh Level Synthesis is a set of methodologies aimed at generating hardware descriptions starting from specifications written in high-level languages. While these methodologies share different elements with traditional compilation flows, there are characteristics of the addressed problem which require ad hoc management. In particular, differently from most of the traditional compilation flows, the complexity and the execution time of the High Level Synthesis techniques are much less relevant than the quality of the produced results. For this reason, fixed-point analyses, as well as successive refinement optimizations, can be accepted, provided that they can improve the quality of the generated designs. This article presents a design flow engine for the description and the execution of complex and customized synthesis flows. It supports dynamic addition of passes and dependencies, cyclic dependencies, and selective pass invalidation. Experimental results show the benefits of such type of design flows with respect to static linear design flows when applied to High Level Synthesis. Marco Lattuada 0001, Fabrizio Ferrandi |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2017 | Exploiting vectorization in high level synthesis of nested irregular loops
Marco Lattuada 0001, Fabrizio Ferrandi |
J. Syst. Archit. | 1 |
| 2017 | Using Efficient Path Profiling to Optimize Memory Consumption of On-Chip Debugging for High-Level SynthesisabstractHigh-Level Synthesis (HLS) for FPGAs is attracting popularity and is increasingly used to handle complex systems with multiple integrated components. To increase performance and efficiency, HLS flows now adopt several advanced optimization techniques. Aggressive optimizations and system level integration can cause the introduction of bugs that are only observable on-chip. Debugging support for circuits generated with HLS is receiving a considerable attention. Among the data that can be collected on chip for debugging, one of the most important is the state of the Finite State Machines (FSM) controlling the components of the circuit. However, this usually requires a large amount of memory to trace the behavior during the execution. This work proposes an approach that takes advantage of the HLS information and of the structure of the FSM to compress control flow traces and to integrate optimized components for on-chip debugging. The generated checkers analyze the FSM execution on-fly, automatically notifying when a bug is detected, localizing it and providing data about its cause. The traces are compressed using a software profiling technique, called Efficient Path Profiling (EPP), adapted for the debugging of hardware accelerators generated with HLS. With this technique, the size of the memory used to store control flow traces can be reduced up to 2 orders of magnitude, compared to state-of-the-art. Pietro Fezzardi, Marco Lattuada 0001, Fabrizio Ferrandi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | A Dynamically Scheduled Architecture for the Synthesis of Graph Database QueriesabstractData analytics applications, such as graph databases, exibit irregular behaviors that make their acceleration non-trivial. These applications expose a significant amount of Task Level Parallelism (TLP), but they present fine grained memory accesses. Marco Minutoli, Vito Giovanni Castellana, Antonino Tumeo, Fabrizio Ferrandi, Marco Lattuada 0001 |
FCCM | 5 |
| 2016 | A dynamically scheduled architecture for the synthesis of graph methodsabstractPresents a collection of slides covering the following: scheduling; graph methods; data analysis; parallel processing; memory interface controller; and load balancing. Marco Minutoli, Vito Giovanni Castellana, Antonino Tumeo, Marco Lattuada 0001, Fabrizio Ferrandi |
Hot Chips Symposium | 4 |
| 2016 | Efficient synthesis of graph methods: a dynamically scheduled architectureabstractRDF databases naturally map to a graph representation and employ languages, such as SPARQL, that implements queries as graph pattern matching routines. Graph methods exhibit an irregular behavior: they present unpredictable, fine-grained data accesses, and are synchronization intensive. Graph data structures expose large amounts of dynamic parallelism, but are difficult to partition without generating load unbalance. In this paper, we present a novel architecture to improve the synthesis of graph methods. Our design addresses the issues of these algorithms with two components: a Dynamic Task Scheduler (DTS), which reduces load unbalance and maximize resource utilization, and a Hierarchical Memory Interface controller (HMI), which provides support for concurrent memory operations on multi-ported/multi-banked shared memories. We evaluate our approach by generating the accelerators for a set of SPARQL queries from the Lehigh University Benchmark (LUBM). We first analyze the load unbalance of these queries, showing that execution time among tasks can differ even of order of magnitudes. We then synthesize the queries and compare the performance of the resulting accelerators against the current state of the art. Experimental results show that our solution provides a speedup over the serial implementation close to the theoretical maximum and a speedup up to 3.45 over a baseline parallel implementation. We conclude our study by exploring the design space to achieve maximum memory channels utilization. The best design used at least three of the four memory channels for more than 90% of the execution time. Marco Minutoli, Vito Giovanni Castellana, Antonino Tumeo, Marco Lattuada 0001, Fabrizio Ferrandi |
ICCAD | 4 |
| 2015 | High Level Synthesis of RDF Queries for Graph AnalyticsabstractIn this paper we present a set of techniques that enable the synthesis of efficient custom accelerators for memory intensive, irregular applications. To address the challenges of irregular applications (large memory footprint, unpredictable fine-grained data accesses, and high synchronization intensity), and exploit their opportunities (thread level parallelism, memory level parallelism), we propose a novel accelerator design that employs an adaptive and Distributed Controller (DC) architecture, and a Memory Interface Controller (MIC) that supports concurrent and atomic memory operations on a multi-ported/multi-banked shared memory. Among the multitude of algorithms that may benefit from our solution, we focus on the acceleration of graph analytics applications and, in particular, on the synthesis of SPARQL queries on Resource Description Framework (RDF) databases. We achieve this objective by incorporating the synthesis techniques into Bambu, an Open Source high-level synthesis tools, and interfacing it with GEMS, the Graph database Engine for Multithreaded Systems. The GEMS' front-end generates optimized C implementations of the input queries, modeled as graph pattern matching algorithms, which are then automatically synthesized by Bambu. We validate our approach by synthesizing several SPARQL queries from the Lehigh University Benchmark (LUBM). Vito Giovanni Castellana, Marco Minutoli, Alessandro Morari, Antonino Tumeo, Marco Lattuada 0001, Fabrizio Ferrandi |
ICCAD | 5 |
| 2015 | Code Transformations Based on Speculative SDC SchedulingabstractCode motion and speculations are usually exploited in the High Level Synthesis of control dominated applications to improve the performances of the synthesized designs. Selecting the transformations to be applied is not a trivial task: their effects can indeed indirectly spread across the whole design, potentially worsening the quality of the results. In this paper we propose a code transformation flow, based on a new extension of the System of Difference Constraints (SDC) scheduling algorithm, which introduces a large number of transformations, whose profitability is guaranteed by SDC formulation. Experimental results show that the proposed technique in average reduces the execution time of control dominated applications by 37% with respect to a commercial tool without increasing the area usage. Marco Lattuada 0001, Fabrizio Ferrandi |
ICCAD | 1 |
| 2012 | Performance estimation of embedded software with confidence levelsabstractSince time constraints are a very critical aspect of an embedded system, performance evaluation can not be postponed to the end of the design flow, but it has to be introduced since its early stages. Estimation techniques based on mathematical models are usually preferred during this phase since they provide quite accurate estimation of the application performance in a fast way. However, the estimation error has to be considered during design space exploration to evaluate if a solution can be accepted (e.g., by discarding solutions whose estimated time is too close to constraint). Evaluate if the possible error can be significant analyzing a punctual estimation is not a trivial task. In this paper we propose a methodology, based on statistical analysis, that provides a prediction interval on the estimation and a confidence level on meeting a time constraint. This information can drive design space exploration reducing the number of solutions to be validated. The results show how the produced intervals effectively capture the estimation error introduced by a linear model. Marco Lattuada 0001, Fabrizio Ferrandi |
ASP-DAC | 1 |
| 2009 | Performance estimation for task graphs combining sequential path profiling and control dependence regionsabstractThe speed-up estimation of parallelized code is crucial to efficiently compare different parallelization techniques or task graph transformations. Unfortunately, most of the time, during the parallelization of a specification, the information that can be extracted by profiling the corresponding sequential code (e.g. the most executed paths) are not properly taken into account. In particular, correlating sequential path profiling with the corresponding parallelized code can help in the identification of code hot spots, opening new possibilities for automatic parallelization. For this reason, starting from a well-known profiling technique, the Efficient Path Profiling, we propose a methodology that estimates the speed-up of a parallelized specification, just using the corresponding hierarchical task graph representation and the information coming from the dynamic profiling of the initial sequential specification. Experimental results show that the proposed solution outperforms existing approaches. Fabrizio Ferrandi, Marco Lattuada 0001, Christian Pilato, Antonino Tumeo |
MEMOCODE | 2 |