Lorenzo Ferretti

dblp:185/7240 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-8935-6796ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Approximate Logic Synthesis Via Iterative SMT-Based Subcircuit Rewriting
abstract
This paper presents a novel iterative approach to achieve effective and efficient approximate logic synthesis (ALS). The core idea is to perform circuit rewriting in a way that is both local, i.e., is applied piece-wise to selected subcircuits, and extensive, i.e., systematically explores the design space for good solutions. Concretely, we propose SubXPAT, a new Boolean rewriting framework which iteratively employs satisfiability modulo theories (SMT) solving to select and approximate key parts of a circuit. Selection aims at finding subcircuits that at the same time include a significant number of gates and can be efficiently approximated, which is done by searching for large convex subcircuits with a limited number of inputs and outputs. Approximation is guided by the use of a parametric template, structured as a sum of products, which allows for fine-grained control over the subcircuit characteristics. SubXPAT was implemented as an open-source tool and compared against other ALS tools implementing state-of-the-art techniques. Our experimental evaluation used a broad range of arithmetic circuits with different bit-widths and our results indicate that SubXPAT generates approximate circuits that are more area-efficient than those generated by state-of-the-art techniques in 72% of the cases.
Morteza Rezaalipour, Marco Biasion, Francesco Costa, Cristian Tirelli, Lorenzo Ferretti, Rodrigo Otoni, George A. Constantinides, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 SAT-Based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
abstract
Coarse-Grain Reconfigurable Arrays (CGRAs) represent emerging low-power architectures designed to accelerate Compute-Intensive Loops (CILs). The effectiveness of CGRAs in providing acceleration relies on the quality of mapping: how efficiently the CIL is compiled onto the platform. State-of-the-Art (SoA) compilation techniques utilize modulo scheduling to minimize the Iteration Interval (II) and use graph algorithms like Max-Clique Enumeration to address mapping challenges. Our work approaches the mapping problem through a satisfiability (SAT) formulation. We introduce the Kernel Mobility Schedule (KMS), an ad hoc schedule used with the Data Flow Graph and CGRA architectural information to generate Boolean statements that, when satisfied, yield a valid mapping. Experimental results demonstrate SAT-MapIt outperforming SoA alternatives in almost 50% of explored benchmarks. Additionally, we evaluated the mapping results in a synthesizable CGRA design and emphasized the runtime metrics trends, i.e., energy efficiency and latency, across different CILs and CGRA sizes. We show that a hardware-agnostic analysis performed on compiler-level metrics can optimally prune the architectural design space, while still retaining Pareto-optimal configurations. Moreover, by exploring how implementation details impact cost and performance on real hardware, we highlight the importance of holistic software-to-hardware mapping flows, as the one presented herein.
Cristian Tirelli, Juan Sapriza, Rubén Rodríguez Álvarez, Lorenzo Ferretti, Benoît W. Denkinger, Giovanni Ansaloni, José Miranda 0001, David Atienza 0001, Laura Pozzi 0001
ACM J. Emerg. Technol. Comput. Syst.4
2023 ErrorEval: an Open-Source Worst-Case-Error Evaluation Framework for Approximate Computing
abstract
Approximate Computing is a design paradigm that allows for a small loss in accuracy in an application in exchange for improved efficiency and/or reduced power consumption. Approximate Logic Synthesis (ALS) is a process through which an inexact (approximate) version of a circuit is generated, assuring that the error introduced by approximation does not exceed a certain threshold [1].
Morteza Rezaalipour, Lorenzo Ferretti, Ilaria Scarabottolo, George A. Constantinides, Laura Pozzi 0001
CF2
2023 SAT-MapIt: An Open Source Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
abstract
The need for low power and high performance architectures - that can efficiently handle compute-intensive tasks while working with tight power and resource constraints - has been steadily rising due to the constant growth of computational demands in everyday applications.
Cristian Tirelli, Lorenzo Ferretti, Laura Pozzi 0001
CF2
2023 SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
abstract
Coarse-Grain Reconfigurable Arrays (CGRAs) are emerging low-power architectures aimed at accelerating compute-intensive application loops. The acceleration that a CGRA can ultimately provide, however, heavily depends on the quality of the mapping, i.e. on how effectively the loop is compiled onto the given platform. State of the Art compilation techniques achieve mapping through modulo scheduling, a strategy which attempts to minimize the II (Iteration Interval) needed to execute a loop, and they do so usually through well known graph algorithms, such as Max-Clique Enumeration. We address the mapping problem through a SAT formulation, instead, and thus explore the solution space more effectively than current SoA tools. To formulate the SAT problem, we introduce an ad-hoc schedule called the kernel mobility schedule (KMS), which we use in conjunction with the data-flow graph and the architectural information of the CGRA in order to create a set of boolean statements that describe all constraints to be obeyed by the mapping for a given II. We then let the SAT solver efficiently navigate this complex space. As in other SoA techniques, the process is iterative: if a valid mapping does not exist for the given II, the II is increased and a new KMS and set of constraints is generated and solved. Our experimental results show that SAT-MapIt obtains better results compared to SoA alternatives in 47.72% of the benchmarks explored: sometimes finding a lower II, and others even finding a valid manning when none could previously be found.
Cristian Tirelli, Lorenzo Ferretti, Laura Pozzi 0001
DATE2
2023 Graph Neural Networks for High-Level Synthesis Design Space Exploration
abstract
High-level Synthesis (HLS) Design-Space Exploration (DSE) aims at identifying Pareto-optimal synthesis configurations whose exhaustive search is unfeasible due to the design-space dimensionality and the prohibitive computational cost of the synthesis process. Within this framework, we address the design automation problem by proposing graph neural networks that jointly predict acceleration performance and hardware costs of a synthesized behavioral specification given optimization directives. Learned models can be used to rapidly approach the Pareto curve by guiding the DSE, taking into account performance and cost estimates. The proposed method outperforms traditional HLS-driven DSE approaches, by accounting for the arbitrary length of computer programs and the invariant properties of the input. We propose a novel hybrid control and dataflow graph representation that enables training the graph neural network on specifications of different hardware accelerators. Our approach achieves prediction accuracy comparable with that of state-of-the-art simulators without having access to analytical models of the HLS compiler. Finally, the learned representation can be exploited for DSE in unexplored configuration spaces by fine-tuning on a small number of samples from the new target domain. The outcome of the empirical evaluation of this transfer learning shows strong results against state-of-the-art baselines in relevant benchmarks.
Lorenzo Ferretti, Andrea Cini, Georgios Zacharopoulos 0001, Cesare Alippi, Laura Pozzi 0001
ACM Trans. Design Autom. Electr. Syst.1
2022 INCLASS: Incremental Classification Strategy for Self-Aware Epileptic Seizure Detection
abstract
Wearable Health Companions allow the unobtrusive monitoring of patients affected by chronic conditions. In particular, by acquiring and interpreting bio-signals, they enable the detection of acute episodes in cardiac and neurological ailments. Nevertheless, the processing of bio-signals is computationally complex, especially when a large number of features are required to obtain reliable detection outcomes. Addressing this challenge, we present a novel methodology, named INCLASS, that iteratively extends employed feature sets at run-time, until a confidence condition is satisfied. INCLASS builds such sets based on code analysis and profiling information. When applied to the challenging scenario of detecting epileptic seizures based on ECG and SpO2 acquisitions, INCLASS obtains savings of up to 54%, while incurring in a negligible loss of detection performance (1.1% degradation of specificity and sensitivity) with respect to always computing and evaluating all features.
Lorenzo Ferretti, Giovanni Ansaloni, Renaud Marquis, Tomás Teijeiro, Philippe Ryvlin, David Atienza 0001, Laura Pozzi 0001
DATE1
2020 Leveraging Prior Knowledge for Effective Design-Space Exploration in High-Level Synthesis
abstract
High-Level Synthesis (HLS) tools allow the generation of a large variety of hardware implementations from the same specification by setting different optimization directives. Each combination of HLS directives returns an implementation of the target application that is based on a particular microarchitecture. Designers are interested only in the subset of implementations that correspond to Pareto-optimal points in the performance versus cost design space. Finding this subset is hard because the relationship between the HLS directives and the Pareto-optimal implementations cannot be foreseen. Hence, designers must default to an exploration of the design space through many time-consuming HLS runs. We present a methodology that infers knowledge from past design explorations to identify high-quality directives for new target applications. To this end, we formulate a novel abstract representation of applications and their associated configuration spaces, introduce a similarity metric to compare quantitatively the configuration spaces of different applications, and a method to infer actionable information from a source space to a target space. The experimental results with the MachSuite benchmarks show that our approach retrieves close approximations of the Pareto frontier of best-performing implementations for the target application, in exchange for a small number of HLS runs.
Lorenzo Ferretti, Jihye Kwon, Giovanni Ansaloni, Giuseppe Di Guglielmo, Luca P. Carloni, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Tailoring SVM Inference for Resource-Efficient ECG-Based Epilepsy Monitors
abstract
Event detection and classification algorithms are resilient towards aggressive resource-aware optimisations. In this paper, we leverage this characteristic in the context of smart health monitoring systems. In more detail, we study the attainable benefits resulting from tailoring Support Vector Machine (SVM) inference engines devoted to the detection of epileptic seizures from ECG-derived features. We conceive and explore multiple optimisations, each effectively reducing resource budgets while minimally impacting classification performance. These strategies can be seamlessly combined, which results in 12.5X and 16X gains in energy and area, respectively, with a negligible loss, 3.2% in classification performance.
Lorenzo Ferretti, Giovanni Ansaloni, Laura Pozzi 0001, Amir Aminifar, David Atienza 0001, Leila Cammoun, Philippe Ryvlin
DATE1
2019 Compiler-Assisted Selection of Hardware Acceleration Candidates from Application Source Code
abstract
Hardware design is a difficult task. Beside ensuring functional correctness of an implementation, hardware developers are confronted with multiple and often conflicting constraints, such as performance and area cost targets, that require lengthy explorations. This issue is compounded when considering the acceleration of complex applications, of which some parts are implemented in software, and others are accelerated in hardware. Hardware/Software partitioning must be settled early in the development cycle, and is far from trivial, since at this stage detailed performance measurements are not available, while wrong choices can lead to vastly sub-optimal solutions or to wasted implementation efforts. To address this challenge, we present a framework to automatically identify, from un-modified software code, software segments that are promising candidates for hardware acceleration, to evaluate their potential speedup and resource requirements, and to select a subset of them under resource constraint. Our strategy is based on Intermediate Representation (IR) analysis passes, which we embed in the LLVM compiler toolchain, and does not require any time-consuming synthesis. We explore its effectiveness on the reference software implementation of a complex application, the H.264 Decoder from University of Illinois, and demonstrate that our methodology selects higher-performance sets of accelerators, when compared to strategies only based on profiling information.
Georgios Zacharopoulos 0001, Lorenzo Ferretti, Giovanni Ansaloni, Giuseppe Di Guglielmo, Luca P. Carloni, Laura Pozzi 0001
ICCD2
2019 RegionSeeker: Automatically Identifying and Selecting Accelerators From Application Source Code
abstract
Embedded systems present stringent and often conflicting requirements. On the one side, the need for high performance within a tight energy budget favors inflexible Application Specific Integrated Circuit (ASIC) implementations; on the other side, a short time-to-market demands programmability. Hybrid architectures such as special-purpose customized processors represent an attractive solution, as they are programmable by software, but use dedicated hardware to accelerate parts of the computation. In such a scenario, the capability of automatically identifying the computation parts to be realized in hardware is highly desirable, in order to reduce design time and effort. This paper aims at advancing the state-of-the-art in this field. We recognize that subgraphs of control flow graphs having a single input control point and a single output control point, that we call regions, are good targets for the synthesis of application specific hardware accelerators. We therefore provide a method to identify them and an LLVM-based toolchain (named RegionSeeker) that, analyzing a software application, automatically selects its most profitable regions given an area constraint. Experimental evidence shows that the accelerators identified by RegionSeeker provide a speedup of up to $4.6\boldsymbol {\times }$ and, on average, approximately 30% higher speedup is achieved compared to state-of-the-art identification techniques.
Georgios Zacharopoulos 0001, Lorenzo Ferretti, Emanuele Giaquinta, Giovanni Ansaloni, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Lattice-Traversing Design Space Exploration for High Level Synthesis
abstract
This paper describes a design space exploration methodology for High Level Synthesis (HLS) frameworks. Inputs of HLS tools are a description (usually in C/C++) of the functionality of an intended hardware, and a set of optimisation directives that specify its implementation, hence allowing the generation of many design variants with widely varying performance and required resources. The relationship between directives and performance/cost is nonetheless not straightforward, and highly influenced by application-specific characteristics. A major challenge facing designers is then to define effective values for the directives while avoiding time-consuming - and often infeasible - exhaustive explorations. We herein address it by proposing a novel HLS exploration approach which employs a lattice representation of the design space, and a methodology for its navigation. We base our strategy on the observation that Pareto-implementations share a low variance among their configurations. We therefore guide the selection of HLS directives minimising the variance of new candidate solutions, with respect to the best performing ones that have already been visited. By only requiring local searches in the lattice space, our methodology gracefully scales to complex designs. It results in close approximations of the real Pareto frontier, while requiring a lower workload and fewer synthesis runs with respect to existing approaches.
Lorenzo Ferretti, Giovanni Ansaloni, Laura Pozzi 0001
ICCD1
2016 An Indoor Localization System for Telehomecare Applications
abstract
In this paper, we present a novel probabilistic technique, based on the Bayes filter, able to estimate the user location, even with unreliable sensor data coming only from fixed sensors in the monitored environment. Our approach has been extensively tested in a home-like environment, as well as in a real home, and achieves very good results. We present results on two datasets, representative of real life conditions, collected during the testing phase. We detect the patient location with subroom accuracy, an improvement over the state of the art for localization using only environmental sensors. The main drawback is that it is only suitable for applications where a single person is present in the environment, like as with other approaches that do not use any mobile device. For this reason, we introduced the “telehomecare” term, therefore differentiating from generic telemedicine applications, where many people can be in the same environment at the same time.
Augusto Luis Ballardini, Lorenzo Ferretti, Simone Fontana, Axel Furlan, Domenico G. Sorrenti
IEEE Trans. Syst. Man Cybern. Syst.2