EDBT 2026 Demo / reviewers in the wild / expert
Emanuele Del Sozzo
dblp:179/3087
· DBLP profile ↗
24ranked-venue papers
7as first author
13since 2021 · last 2024
0000-0003-3101-8118ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 7 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Flexible Systolic Array Platform on Virtual 2-D Multi-FPGA PlaneabstractSystolic arrays are a promising approach to achieving high-performance processing based on highly parallelized designs in various fields, such as AI and bioinformatics. Many previous studies have devoted considerable effort to exploring efficient circuit designs for specific processing. However, the increasing size of systolic arrays forces us to process increasingly large workloads by dividing them into smaller pieces. Therefore, we propose a systolic array platform based on a two-dimensional FPGA plane in which multiple FPGAs are connected by a virtual network. The systolic array realized by this system can be freely customized in shape and size according to the target. Distributed memory access through off-chip memory on each FPGA board and simple stream processing enable scalable performance. This paper presents a preliminary implementation based on the proposed systolic array platform and its performance evaluation. The evaluation results show that the proposed method improves the processing performance in proportion to the number of FPGAs. The results also show that the proposed platform is highly scalable due to the small circuit area required, and that the processing performance depends on the network bandwidth, which means that recent high-bandwidth FPGA boards can be expected to significantly improve the performance. Tomohiro Ueno, Emanuele Del Sozzo, Kentaro Sano |
HPC Asia | 2 |
| 2024 | Exploration of Trade-offs Between General-Purpose and Specialized Processing Elements in HPC-Oriented CGRAabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable accelerators traditionally used in embedded computing. Recently, CGRA-like devices have gained traction for HPC and AI acceleration; however, typical HPC and AI workloads often require operations that current CGRAs cannot implement, such as complex mathematical calculations. In this work, we present a broad architectural study exploring potential heterogeneous computational resources in CGRA architectures for HPC, which are not commonly considered in typical CGRA architecture research. We first improved the general-purpose Processing Element (PE) of a baseline CGRA to optimize computational resources and then developed a new specialized PE for mathematical functions commonly found in HPC applications. Finally, we evaluated multiple CGRA configurations concerning floorplan, size, general-purpose/specialized PE ratio, and Power, Performance, and Area (PPA) results from hardware synthesis. Emanuele Del Sozzo, Xinyuan Wang 0003, Boma Anantasatya Adhi, Carlos Cortes, Jason Helge Anderson, Kentaro Sano |
IPDPS | 1 |
| 2024 | Starlight: A kernel optimizer for GPU processingabstractOver the past few years, GPUs have found widespread adoption in many scientific domains, offering notable performance and energy efficiency advantages compared to CPUs. However, optimizing GPU high-performance kernels poses challenges given the complexities of GPU architectures and programming models. Moreover, current GPU development tools provide few high-level suggestions and overlook the underlying hardware. Here we present Starlight, an open-source, highly flexible tool for enhancing GPU kernel analysis and optimization. Starlight autonomously describes Roofline Models, examines performance metrics, and correlates these insights with GPU architectural bottlenecks. Additionally, Starlight predicts potential performance enhancements before altering the source code. We demonstrate its efficacy by applying it to literature genomics and physics applications, attaining speedups from 1.1× to 2.5× over state-of-the-art baselines. Furthermore, Starlight supports the development of new GPU kernels, which we exemplify through an image processing application, showing speedups of 12.7× and 140× when compared against state-of-the-art FPGA- and GPU-based solutions. Alberto Zeni, Emanuele Del Sozzo, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio |
J. Parallel Distributed Comput. | 2 |
| 2024 | Across Time and Space: Senju's Approach for Scaling Iterative Stencil Loop Accelerators on Single and Multiple FPGAsabstractStencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional point grid multiple times or until convergence. However, due to their iterative and intensive nature, ISLs are highly performance-hungry, demanding specialized solutions. Here, Field Programmable Gate Arrays (FPGAs) represent a valid architectural choice as they enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this article introduces Senju , an automation framework for the design of highly parallel ISL accelerators targeting single-/multi-FPGA systems. Given an input description, Senju automates the entire design process and provides accurate performance estimations. The experimental evaluation shows remarkable and scalable results, outperforming single- and multi-FPGA literature approaches under different metrics. Finally, we present a new analysis of temporal and spatial parallelism trade-offs in a real-case scenario and discuss our performance through a single- and novel specialized multi-FPGA formulation of the Roofline Model. Emanuele Del Sozzo, Davide Conficconi, Kentaro Sano |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | Senju: A Framework for the Design of Highly Parallel FPGA-based Iterative Stencil Loop AcceleratorsabstractStencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving, seismic simulations, and financial option pricing, to name a few. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional system of points until it reaches convergence. However, due to their iterative and computationally intensive nature, these workloads are highly performance-hungry, demanding specialized solutions to boost performance and reduce power consumption. Here, FPGAs represent a valid architectural choice as their peculiar features enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this paper introduces Senju, an automation framework for FPGA-based ISL accelerators. Starting from an input description, Senju builds highly parallel hardware modules and automatizes all their design phases. The experimental evaluation shows remarkable and scalable results, reaching significant performance and energy efficiency improvements compared to the other single-FPGA literature approaches. Emanuele Del Sozzo, Davide Conficconi, Marco D. Santambrogio, Kentaro Sano |
FPGA | 1 |
| 2023 | Faber: A Hardware/SoftWare Toolchain for Image RegistrationabstractImage registration is a well-defined computation paradigm widely applied to align one or more images to a target image. This paradigm, which builds upon three main components, is particularly compute-intensive and represents many image processing pipelines’ bottlenecks. State-of-the-art solutions leverage hardware acceleration to speed up image registration, but they are usually limited to implementing a single component. We present Faber, an open-source HW/SW CAD toolchain tailored to image registration. The Faber toolchain comprises HW/SW highly-tunable registration components, supports users with different expertise in building custom pipelines, and automates the design process. In this direction, Faber provides both default settings for entry-level users and latency and resource models to guide HW experts in customizing the different components. Finally, Faber achieves from 1.5× to 54× in speedup and from 2× to 177× in energy efficiency against state-of-the-art tools on a Xeon Gold. Eleonora D'Arnese, Davide Conficconi, Emanuele Del Sozzo, Luigi Fusco, Donatella Sciuto, Marco D. Santambrogio |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Large Forests and Where to "Partially" Fit ThemabstractThe Artificial Intelligence of Things (AIoT) calls for on-site Machine Learning inference to overcome the instability in latency and availability of networks. Thus, hardware acceleration is paramount for reaching the Cloud's modeling performance within an embedded device's resources. In this paper, we propose Entree, the first automatic design flow for deploying the inference of Decision Tree (DT) ensembles over Field-Programmable Gate Arrays (FPGAs) at the network's edge. It exploits dynamic partial reconfiguration on modern FPGA-enabled Systems-on-a-Chip (SoCs) to accelerate arbitrarily large DT ensembles at a latency a hundred times stabler than software alternatives. Plus, given Entree's suitability for both hardware designers and non-hardware-savvy developers, we believe it has the potential of helping data scientists to develop a non-Cloud-centric AIoT. Andrea Damiani, Emanuele Del Sozzo, Marco D. Santambrogio |
ASP-DAC | 2 |
| 2022 | Surfing the Wavefront of Genome AlignmentabstractPairwise sequence alignment represents a fundamental step in genome and molecular analysis applications, accounting for most of their runtime. Given the quadratic time complexity of alignment algorithms, the community presses for the development of more efficient algorithms. Moreover, current limitations of general-purpose architectures push users to use hardware accelerators to reduce the analysis time. In this context, we present an FPGA implementation of the Wavefront Alignment (WFA) algorithm, a recently introduced solution that exploits homologous regions between the sequences to speed up the alignment process and whose complexity is related to the score of the alignment, rather than to the lengths of the sequences. Our multicore design can achieve up to 8.09 × improvement in speedup and 57.77 × in energy efficiency compared to the multithreaded software implementation run on a Xeon Gold Processor. Moreover, our design highly outperforms the current State-of-the-Art hardware-accelerated solution, reaching up to 2876 Giga Cell Updates Per Second (GCUPS) and 68.47 GCUPS/W on a single FPGA, with an improvement of up to 2.29× and 9.90× in terms of performance and energy efficiency, respectively. Beatrice Branchini, Giulia Gerometta, Luisa Cicolini, Alberto Zeni, Emanuele Del Sozzo, Marco D. Santambrogio |
ISCAS | 5 |
| 2022 | A Comprehensive Methodology to Optimize FPGA Designs via the Roofline ModelabstractWith reconfigurable fabrics delivering increasing performance over the years, Field-Programmable Gate Arrays (FPGAs) are becoming an appealing solution for next-generation High-Performance Computing (HPC) systems. However, in order to gain traction among traditional von Neumann architectures, the optimization process of Field-Programmable Gate Array (FPGA) designs should be further abstracted to a higher level. In fact, while High-Level Synthesis (HLS) already provides a handy way to write FPGA code with common high-level languages, substantial effort and expertise are still required to optimize the resulting FPGA design for the underlying hardware. To overcome this problem, we propose a semi-automated performance optimization methodology based on a Hierarchical Roofline model for FPGAs. System-wide and applications-specific optimizations such as off-chip memory transfer and data locality optimizations are guided by the FPGA Roofline model whereas FPGA-specific optimizations are automatically searched by a Design Space Exploration (DSE) engine. We demonstrate the way this methodology allows to easily analyze and optimize to peak system performance a wide set of applications ranging from particle methods, wavefront algorithms, and sparse arithmetic computations. In addition, we prove that the integrated Design Space Exploration (DSE) engine achieves a 14.36x maximum speedup if compared to previous automated solutions in the literature. Marco Siracusa, Emanuele Del Sozzo, Marco Rabozzi, Lorenzo Di Tucci, Samuel Williams 0001, Donatella Sciuto, Marco D. Santambrogio |
IEEE Trans. Computers | 2 |
| 2022 | On the Automation of Radiomics-Based Identification and Characterization of NSCLCabstractProper detection and accurate characterization of Non-Small Cell Lung Cancer (NSCLC) are an open challenge in the imaging field. Biomedical imaging is fundamental in lung cancer assessment and offers the possibility of calculating predictive biomarkers impacting patients' management. Within this context, radiomics, which consists of extracting quantitative features from digital images, shows encouraging results for clinical applications, but the sub-optimal standardization of the procedure and the lack of definitive results are still a concern in the field. For these reasons, this work proposes the design and development of LuCIFEx, a fully-automated pipeline for non-invasive in-vivo characterization of NSCLC, aiming to speed up the analysis process and enable an early diagnosis of the tumor.LuCIFEx pipeline relies on routinely acquired [18F]FDG-PET/CT images for the automatic segmentation of the cancer lesion, allowing the computation of accurate radiomic features, then employed for cancer characterization through Machine Learning algorithms. The proposed multi-stage segmentation process can identify the lesion with a mean accuracy of 94.2±5.0%. Finally, the proposed data analysis pipeline demonstrates the potential of PET/CT features for the automatic recognition of lung metastases and NSCLC histological subtypes, while highlighting the main current limitations of the radiomic approach. Eleonora D'Arnese, Guido Walter Di Donato, Emanuele Del Sozzo, Martina Sollini, Donatella Sciuto, Marco D. Santambrogio |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | A Framework for Customizable FPGA-based Image Registration AcceleratorsabstractImage Registration is a highly compute-intensive optimization procedure that determines the geometric transformation to align a floating image to a reference one. Generally, the registration targets are images taken from different time instances, acquisition angles, and/or sensor types. Several methodologies are employed in the literature to address the limiting factors of this class of algorithms, among which hardware accelerators seem the most promising solution to boost performance. However, most hardware implementations are either closed-source or tailored to a specific context, limiting their application to different fields. For these reasons, we propose an open-source hardware-software framework to generate a configurable architecture for the most compute-intensive part of registration algorithms, namely the similarity metric computation. This metric is the Mutual Information, a well-known calculus from the Information Theory, used in several optimization procedures. Through different design parameters configurations, we explore several design choices of our highly-customizable architecture and validate it on multiple FPGAs. We evaluated various architectures against an optimized Matlab implementation on an Intel Xeon Gold, reaching a speedup up to 2.86x, and remarkable performance and power efficiency against other state-of-the-art approaches. Davide Conficconi, Eleonora D'Arnese, Emanuele Del Sozzo, Donatella Sciuto, Marco D. Santambrogio |
FPGA | 3 |
| 2021 | CICERO: A Domain-Specific Architecture for Efficient Regular Expression MatchingabstractRegular Expression (RE) matching is a computational kernel used in several applications. Since RE complexity and data volumes are steadily increasing, hardware acceleration is gaining attention also for this problem. Existing approaches have limited flexibility as they require a different implementation for each RE. On the other hand, it is complex to map efficient RE representations like non-deterministic finite-state automata onto software-programmable engines or parallel architectures. In this work, we present CICERO , an end-to-end framework composed of a domain-specific architecture and a companion compilation framework for RE matching. Our solution is suitable for many applications, such as genomics/proteomics and natural language processing. CICERO aims at exploiting the intrinsic parallelism of non-deterministic representations of the REs. CICERO can trade-off accelerators’ efficiency and processors’ flexibility thanks to its programmable architecture and the compilation framework. We implemented CICERO prototypes on embedded FPGA achieving up to 28.6× and 20.8× more energy efficiency than embedded and mainstream processors, respectively. Since it is a programmable architecture, it can be implemented as a custom ASIC that is orders of magnitude more energy-efficient than mainstream processors. Daniele Parravicini, Davide Conficconi, Emanuele Del Sozzo, Christian Pilato, Marco D. Santambrogio |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | Enhancing the Scalability of Multi-FPGA Stencil Computations via Highly Optimized HDL ComponentsabstractStencil-based algorithms are a relevant class of computational kernels in high-performance systems, as they appear in a plethora of fields, from image processing to seismic simulations, from numerical methods to physical modeling. Among the various incarnations of stencil-based computations,Iterative Stencil Loops (ISLs)andConvolutional Neural Networks (CNNs)represent two well-known examples of kernels belonging to the stencil class. Indeed, ISLs apply the same stencil several times until convergence, while CNN layers leverage stencils to extract features from an image. The computationally intensive essence of ISLs, CNNs, and in general stencil-based workloads, requires solutions able to produce efficient implementations in terms of throughput and power efficiency. In this context, FPGAs are ideal candidates for such workloads, as they allow design architectures tailored to the stencil regular computational pattern. Moreover, the ever-growing need for performance enhancement leads FPGA-based architectures to scale to multiple devices to benefit from a distributed acceleration. For this reason, we propose a library of HDL components to effectively compute ISLs and CNNs inference on FPGA, along with a scalable multi-FPGA architecture, based on custom PCB interconnects. Our solution eases the design flow and guarantees both scalability and performance competitive with state-of-the-art works. Enrico Reggiani, Emanuele Del Sozzo, Davide Conficconi, Giuseppe Natale, Carlo Moroni, Marco D. Santambrogio |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2020 | A CAD-based methodology to optimize HLS code via the Roofline modelabstractThe intrinsic complexity of modern computing systems requires structured methods for analyzing and optimizing application performance. In this context, the Roofline model proposes an intuitive and visual method providing performance insight and optimization guidance for a given architecture. Although this methodology successfully models multicore and GPU performance optimizations, the original formulation does not directly apply to FPGA devices. For this reason, we propose a Roofline model analysis for reconfigurable architectures and an associated CAD tool for assisting HLS optimization of C/C++ applications. We firstly model FPGA attainable performance by means of an analytical method. Then, we integrate locality walls and a DSE engine for an enhanced optimization process. Starting from a software version of the N-body algorithm, we firstly illustrate how our methodology helps at quickly achieving performance comparable to a state-of-the-art FPGA bespoke implementation. Then, we illustrate an assisted platform porting of the Smith-Waterman sequence alignment providing a 9x speedup. Finally, we evaluated the single DSE engine on the Poly-Bench test suite and achieved performance improvements up to 14.36x compared to previous automated solutions in the literature. Marco Siracusa, Marco Rabozzi, Emanuele Del Sozzo, Lorenzo Di Tucci, Samuel Williams 0001, Marco D. Santambrogio |
ICCAD | 3 |
| 2019 | Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable CodeabstractThis paper introduces Tiramisu, a polyhedral framework designed to generate high performance code for multiple platforms including multicores, GPUs, and distributed machines. Tiramisu introduces a scheduling language with novel commands to explicitly manage the complexities that arise when targeting these systems. The framework is designed for the areas of image processing, stencils, linear algebra and deep learning. Tiramisu has two main features: it relies on a flexible representation based on the polyhedral model and it has a rich scheduling language allowing fine-grained control of optimizations. Tiramisu uses a four-level intermediate representation that allows full separation between the algorithms, loop transformations, data layouts, and communication. This separation simplifies targeting multiple hardware architectures with the same algorithm. We evaluate Tiramisu by writing a set of image processing, deep learning, and linear algebra benchmarks and compare them with state-of-the-art compilers and hand-tuned libraries. We show that Tiramisu matches or outperforms existing compilers and libraries on different hardware architectures, including multicore CPUs, GPUs, and distributed machines. Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil 0001, Saman P. Amarasinghe |
CGO | 4 |
| 2019 | Automated Acceleration of Dataflow-Oriented C Applications on FPGA-Based SystemsabstractThe acceleration of compute-intensive applications on FPGA-based systems has become an increasingly common trend thanks to their availability as cloud commodities. This trend has also been accompanied by wider support of High-Level Synthesis tools. Despite these solutions reduce the learning curve for hardware development, the programmer still requires specific expertise in order to achieve efficient implementations. In this paper, we propose an automated approach for the acceleration of C applications into dataflow kernels on FPGAs. Francesco Peverelli, Marco Rabozzi, Salvatore Cardamone, Emanuele Del Sozzo, Alex J. W. Thom, Marco D. Santambrogio, Lorenzo Di Tucci |
FCCM | 4 |
| 2019 | Automated Design Space Exploration and Roofline Analysis for FPGA-Based HLS ApplicationsabstractThe growing interest in FPGA-based solutions for accelerating compute demanding algorithms is pushing the need for new tools and methods to improve productivity. In this work, we propose a methodology to support designers in generating optimal FPGA hardware implementations using High-Level Synthesis (HLS). First, we propose an automated roofline model generation that operates directly on a C/C++ description of the algorithm. The approach enables fast evaluation of the operational intensity of the target function and visualizes the main bottlenecks of the current HLS implementation, providing guidance on how to improve it. Second, we integrate it with a Design Space Exploration (DSE) methodology for quickly evaluating different HLS directives to identify an optimal implementation. Marco Siracusa, Marco Rabozzi, Emanuele Del Sozzo, Marco D. Santambrogio, Lorenzo Di Tucci |
FCCM | 3 |
| 2018 | Five-point algorithm: An efficient cloud-based FPGA implementationabstractThe 5-point relative pose problem is to identify the possible relative camera motions given five matching points from two calibrated views. Several algorithms for solving this problem have been presented in the literature providing different tradeoffs in terms of computational complexity and accuracy of the results. Indeed, the research in this field is driven mostly by the need for accurate solutions and high performance to cope with real-time requirements. In this work we propose an implementation to solve the 5-point relative pose problem accelerated on Field Programmable Gate Array (FPGA). The proposed architecture implements the classical Nister's algorithm as a deep pipeline deployed on a AWS F1 instance and outperforms software implementations by a factor ranging from 7.2X to 233X. Furthermore, it achieves a speedup of 64.2X compared to the Nister's software implementation with comparable accuracy. Marco Rabozzi, Emanuele Del Sozzo, Lorenzo Di Tucci, Marco D. Santambrogio |
ASAP | 2 |
| 2018 | FPGA-based PairHMM Forward Algorithm for DNA Variant CallingabstractOne of the main objectives of human genetic research is the identification of DNA variations that may be involved in the development of rare diseases. Thanks to advances in DNA sequencing technologies and to a progressive integration of the available genetic databases, it is now possible to study not only common variants, but also ones occurring at very low frequencies in the population. Despite the presence of consolidated algorithms to perform the analysis of genetic data, the major hurdle is the impossibility to efficiently process the data, and translate them into biologically meaningful information. This prevents the current solutions and architectures from scaling to growing number of individuals, both for the discovery of new variants and the adoption of these methodologies to support diagnosis and treatment of illnesses in a common clinical setting. In this scenario, Field Programmable Gate Arrays (FPGAs) provide a viable alternative to conventional software-based approaches. In particular, they have already proved to efficiently manage huge workloads while reaching outstanding performance over power consumption scores. Therefore, the purpose of this work is exploring novel computing paradigms to tackle the limitations we are facing. In particular, we present an FPGA-based acceleration of the PairHMM Forward Algorithm, the performance bottleneck in the HaplotypeCaller, a variant calling tool in the popular Genome Analysis Toolkit (GATK). Our final architecture is able to achieve 2160x speedup when compared to the Original Java version in GATK, outperforming existing implementation on both CPUs, GPUs and FPGAs. Davide Sampietro, Chiara Crippa, Lorenzo Di Tucci, Emanuele Del Sozzo, Marco D. Santambrogio |
ASAP | 4 |
| 2018 | A Unified Backend for Targeting FPGAs from DSLsabstractThe major flaw of Field Programmable Gate Arrays (FPGAs) is their hard programmability and steep learning curve. Even though High-Level Synthesis (HLS) tools may alleviate this task by providing directives to optimize the hardware design, as well as supporting languages like C/C++ and OpenCL, the development of efficient designs for FPGA is still a challenging and time-consuming task. In this context, Domain Specific Languages (DSLs) represent an emerging solution to generate efficient code to target FPGAs. However, the support for these languages towards FPGA is still limited, and only few DSLs provide FPGA backends. This paper describes FROST, a unified backend for targeting FPGAs from DSLs. FROST takes as input an algorithm described in one of the supported DSLs and generates an optimized design suitable for HLS tools. To this end, FROST exposes a high-level scheduling co-language to drive many aspects of the optimization process, like the resulting architecture, the level of parallelism, and so on. We evaluated FROST on a set of image processing kernels, developed in Halide and TIRAMISU, and compared the results against a hand-tuned FPGA library. The experimental results demonstrate that FROST designs are able to match the performance of such library (exploiting the same level of parallelism), and surpass it by a factor of 10X when combining FROST and the frontends scheduling commands. Emanuele Del Sozzo, Riyadh Baghdadi, Saman P. Amarasinghe, Marco D. Santambrogio |
ASAP | 1 |
| 2018 | A Scalable FPGA Design for Cloud N-Body SimulationabstractThe N-Body simulation process describes the evolution of a system of forces composed of N bodies, which may represent celestial objects, molecules, and so on. The most accurate algorithm for N-Body simulation, the All-Pairs method, is particularly compute intensive and software implementations on CPUs are inefficient in terms of performance and power consumption. An implementation on a hardware accelerator, such as an FPGA, would benefits in both these terms, exploiting a parallel execution at a relative low power profile. Moreover, it would also benefit faster methods with lower computational complexity, since many of them rely on the All-Pairs approach to approximate the calculation of forces. This work proposes a highly scalable, power efficient and high performance hardware architecture for the N-Body All-Pairs simulation problem. Our final implementation is able to scale up to systems with an arbitrary number of bodies thanks to a tiling approach that allows performance in the order of 13,441 MPairs/s, outperforming state of the art implementations on FPGA in terms of both pure performance, as well as performance per watt ratio. Finally, our design results to be more power efficient than Grape-8 ASIC. Emanuele Del Sozzo, Marco Rabozzi, Lorenzo Di Tucci, Donatella Sciuto, Marco D. Santambrogio |
ASAP | 1 |
| 2017 | Heterogeneous exascale supercomputing: The role of CAD in the exaFPGA projectabstractSince the end of Moore's law is limiting the growth of general purpose processors, High Performance Processing (HPC) systems are considering FPGA-based accelerators as a promising solution for several application fields. However, their employment poses challenges the research is still tackling, and existing tools and workflows do not naturally adapt to the scale and complexity of HPC domains. To help researchers and practitioners, this paper proposes CAOS, a platform that implements an FPGA development workflow tailored to HPC systems while being open to external contributions. Indeed, researchers and developers can plug into CAOS to experiment and compare their solutions at each step of the design flow. This paper describes the CAOS workflow and validates it against several case studies to assess its generality and highlight possible research contributions. Marco Rabozzi, Giuseppe Natale, Emanuele Del Sozzo, Alberto Scolari, Luca Stornaiuolo, Marco D. Santambrogio |
DATE | 3 |
| 2017 | A Common Backend for Hardware Acceleration on FPGAabstractField Programmable Gate Arrays (FPGAs) are configurable integrated circuits able to provide a good trade-off in terms of performance, power consumption, and flexibility with respect to other architectures, like CPUs, GPUs and ASICs. The main drawback in using FPGAs, however, is their steep learning curve. An emerging solution to this problem is to write algorithms in a Domain Specific Language (DSL) and to let the DSL compiler generate efficient code targeting FPGAs. This work proposes FROST, a unified backend that enables different DSL compilers to target FPGA architectures. Differently from other code generation frameworks targeting FPGA, FROST exploits a scheduling co-language that enables users to have full control over which optimizations to apply in order to generate efficient code (e.g. loop pipelining, array partitioning, vectorization). At first, FROST analyzes and manipulates the input Abstract Syntax Tree (AST) in order to apply FPGA-oriented transformations and optimizations, then generates a C/C++ implementation suitable for High-Level Synthesis (HLS) tools. Finally, the output of HLS phase is synthesized and implemented on the target FPGA using Xilinx SDAccel toolchain. The experimental results show a speedup up of 15× with respect to O3-optimized implementations of the same algorithms on CPU. Emanuele Del Sozzo, Riyadh Baghdadi, Saman P. Amarasinghe, Marco D. Santambrogio |
ICCD | 1 |
| 2016 | Workload-aware power optimization strategy for asymmetric multiprocessors
Emanuele Del Sozzo, Gianluca Durelli, Ettore M. G. Trainiti, Antonio Miele, Marco D. Santambrogio, Cristiana Bolchini |
DATE | 1 |