Giovanni Agosta

dblp:a/GiovanniAgosta · DBLP profile ↗
← Back
66ranked-venue papers
28as first author
21since 2021 · last 2026
0000-0002-0255-4475ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 48 · 19 first-author · 17 since 2021Software engineering, systems software and programming languages · 11 · 3 first-author · 3 since 2021Security and privacy · 5 · 3 first-author · 1 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quantum Oracle Synthesis from HDL Designs via Multi Level Intermediate Representation
abstract
Quantum computing is increasingly recognized as a promising approach for tackling computationally intractable problems. However, achieving the scalability necessary for real-world applications requires substantial advancements in the quantum software stack. In this work, we introduce a compiler toolchain based on a Multi-Level Intermediate Representation (MLIR) that automatically synthesizes quantum circuits from Hardware Description Language (HDL) specifications of classical functions into quantum assembly languages. Many quantum algorithms rely on combinatorial circuits as subroutines, which traditionally require extensive resources in terms of quantum gates and qubits and are often manually optimized. Our toolchain integrates a sequence of optimization passes that combine classical compiler techniques with quantum-specific improvements, resulting in an average qubit reduction of 30% and an average gate-count reduction of 20% in widely adopted benchmark circuits, including those used in cryptographic applications.
Giacomo Lancellotti, Filippo Buda, Giacomo Carugati, Daniele Gazzola, Alessandro Barenghi, Giovanni Agosta, Gerardo Pelosi
ASP-DAC6
2026 Type Deduction Analysis: Reconstructing Transparent Pointer Types in LLVM-IR
abstract
With version 17, LLVM finalized the transition to opaque pointer types, eliminating explicit pointee‑type information from the Intermediate Representation (IR). Thus, starting from LLVM 17, each pointer type is represented in IR by the unique type ptr. Despite eliminating redundant pointer bitcasts and consequently reducing IR size and compile time, this change disrupts analyses that have reason to rely on pointee-type information, forcing existing compiler projects to depend on outdated LLVM versions. This information can in fact be insightful in fields like approximate computing, where the compiler can apply non-conservative optimizations, or in passes that require it to make analyses and transformations that do not impact the correctness of the program. To address this problem, we present a new Type Deduction Analysis pass that reconstructs transparent pointer types directly from opaque‑pointer IR. Moreover, we illustrate two different case-studies on existing LLVM projects, namely TAFFO and ASPIS, that demonstrate the need for pointee-type information in LLVM compilers.
Niccolò Nicolosi, Gabriele Magnani, Emilio Corigliano, Davide Baroffio, Federico Reghenzani, Giovanni Agosta
CC6
2025 Scrambling Compiler: Automated and Unified Countermeasure for Profiled and Non-profiled Side Channel Attacks
Gabriele Magnani, Isabella Piacentini, Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi
ARES (1)3
2025 Combining MLIR Dialects with Domain-Specific Architecture for Efficient Regular Expression Matching
abstract
Pattern matching based on Regular Expressions (REs) is a pervasive and challenging computational kernel used in several applications to identify critical information in a data stream. Due to the sequential data dependency of REs and the increasing data volume growth, hardware acceleration is gaining attention to address the limitation of general-purpose architectures. RE-oriented Domain-Specific Architectures (DSAs) combine the flexibility of translating REs into binary code with the efficiency of a specialized architecture, filling the gap between frozen hardware accelerators and the versatility of CPUs/GPUs. However, existing DSAs focus mainly on the efficiency execution challenge while missing the optimization opportunities that a structured compilation infrastructure can provide. This paper proposes a RE-tailored multi-level intermediate representation strategy embodied by the MLIR framework at the compiler level to exploit different abstraction optimizations via two domain-specific dialects, one targeting the abstract representation of REs and the other targeting the underlying domain-specific ISA. Moreover, this paper proposes a novel architectural organization of an open-source state-of-the-art DSA to maximize the parallelization capabilities. Overall, the proposed approach significantly improves execution time by up to 2.26×, energy efficiency by up to 2.30×, and resource usage.
Andrea Somaini, Filippo Carloni, Giovanni Agosta, Marco D. Santambrogio, Davide Conficconi
CGO3
2025 Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 Project
abstract
The European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications.
Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini
DSD1
2025 Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and Reliability
abstract
Modern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners.
Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo
DSD1
2025 The CAPSARII Approach to Cyber-Secure Wearable, Ultra-Low-Power Networked Sensors for Soldier Health Monitoring
abstract
The European Defence Agency’s revised Capability Development Plan (CDP) identifies as a priority improving ground combat capabilities by enhancing soldiers’ equipment for better protection. The CAPSARII project proposes an innovative wearable system and Internet of Battlefield Things (IoBT) framework to monitor soldiers’ physiological and psychological status, aiding tactical decisions and medical support. The CAPSARII system will enhance situational awareness and operational effectiveness by monitoring physiological, movement and environmental parameters, providing real-time tactical decision support through AI models deployed on edge nodes and enable data analysis and comparative studies via cloud-based analytics. CAPSARII also aims at improving usability through smart textile integration, longer battery life, reducing energy consumption through software and hardware optimizations, and address security concerns with efficient encryption and strong authentication methods. This innovative approach aims to transform military operations by providing a robust, data-driven decision support tool.
Luciano Bozzi, Christian Celidonio, Umberto Nuzzi, Massimo Biagini, Stefano Cherubin, Asbjørn Djupdal, Tor A. Haugdahl, Andrea Aliverti, Alessandra Angelucci, Giovanni Agosta, Gerardo Pelosi, Paolo Belluco, Samuele Polistina, Riccardo Volpi, Luigi Malagò, Florian Wieczorek, Xabier Eguiluz
DSD10
2025 Enabling Smart Urban Mobility with Edge AI
abstract
Road accidents are a major cause of death in urban areas. Cooperative Intelligent Transport Systems may mitigate or prevent road accidents by providing drivers with relevant and timely warnings, thanks to V2X communications. However, this requires fast detection of dangers, performed at the edge, to minimize transmission latencies and ensure the timeliness of the warnings. The PMDI project extends the STEP platform for intelligent mobility to support real-time operations and ETSI message generation, and provides integration with mobile and embedded systems.
Andrea Tuscano, Paolo Giuseppetti, Pietro Amato, Alessandro Solinas, Federico Saluz, Andrea Tessieri, Mario Pedol, Manuel Pernigotto, Massimo Fioravanti, Giovanni Agosta, William Fornaciari, Paolo Maffezzoni, Fabio Salice, Irene Amerini, Francesco Pro, Paolo Satta, Giovanni Trovini
DSD11
2025 Modern Llvm-Based Compiler Autotuning for Wcet Optimization
abstract
The problem of compiler optimization selection and ordering, known in the literature as compiler autotuning, has been tackled many times for average-case execution time reduction. Optimizing the WCET is becoming a prominent problem for modern hard real-time systems, where the difficulties in accurate WCET estimation hinder the full exploitation of computing platform capabilities. In this article, we propose a novel methodology and a tool based on LLVM for iterative WCET-driven compiler autotuning, which is the first strategy to operate at function-level granularity and to consider not only the selection of optimization passes, but also their ordering. Our findings show that standard optimization levels$\mathrm{O} 0, \mathrm{O} 1, \mathrm{O} 2$, and O 3 are suboptimal when targeting the WCET, and that a per-function selection and ordering of the transformations is necessary. Experimental results show that our approach outperforms the standard optimizations and opens up new directions for future research.
Gabriele Magnani, Davide Baroffio, Federico Reghenzani, Giovanni Agosta, William Fornaciari
RTSS4
2025 Synergistic Memory Optimisations: Precision Tuning in Heterogeneous Memory Hierarchies
abstract
Balancing energy efficiency and high performance in embedded systems requires fine-tuning hardware and software components to co-optimize their interaction. In this work, we address the automated optimization of memory usage through a compiler toolchain that leverages DMA-aware precision tuning and mathematical function memorization. The proposed solution extends the LLVM infrastructure, employing the TAFFO plugins for precision tuning, with the SETHET extension for DMA-aware precision tuning and LUTHET for automated, DMA-aware mathematical function memorization. We performed an experimental assessment on HERO, a heterogeneous platform employing RISC-V cores as a parallel accelerator. Our solution enables speedups ranging from 1.5× to 51.1× on AxBench benchmarks that employ trigonometrical functions and 4.23–48.4× on Polybench benchmarks over the baseline HERO platform.
Gabriele Magnani, Daniele Cattaneo 0002, Lev Denisov, Giuseppe Tagliavini, Giovanni Agosta, Stefano Cherubin
IEEE Trans. Computers5
2024 SeTHet - Sending Tuned numbers over DMA onto Heterogeneous clusters: an automated precision tuning story
abstract
Energy and performance optimization of embedded hardware and software is of critical importance to achieve the overall system goals. In this work, we study the optimization of memory access through a combination of hardware (Direct Memory Access, DMA) and software (Precision Tuning) techniques, and we propose a compiler toolchain for managing both in the context of heterogeneous RISC-Vbased platforms. Our proposed toolchain, SeTHet, enables 3 - - 48 × speedup over the baseline system when employing both DMA and precision tuning, regardless of the availability of floating point units in hardware. SeTHet also achieves up to 16× speedup compared to DMA alone, thus proving that the combination of the two techniques provides a major improvement over either technique employed in isolation.
Gabriele Magnani, Daniele Cattaneo 0002, Lev Denisov, Giuseppe Tagliavini, Giovanni Agosta, Stefano Cherubin
CF5
2024 The TEXTAROSSA Project: Cool all the Way Down to the Hardware
abstract
The TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research.
Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo
DSD2
2024 Optimizing Quantum Circuit Synthesis with Dominator Analysis
abstract
Quantum circuit synthesis translates a classical Boolean function into an equivalent quantum circuit. The syn-thesis is solvable by playing the reversible pebble game on the logic network of the function. However, optimal solutions are im-practical for large networks, affecting the number of qubits and gates of the final quantum circuit. In this work, we improve the solution of the reversible pebble game by leveraging dominance relations on the directed acyclic graph of the classical function, reducing qubits needed for syntheses. The proposed algorithm exposes a tunable tradeoff between the available number of qubits and the circuit size expressed as T-count and T-depth. We experimentally validate our methodology on cryptographic and arithmetic benchmarks, reporting reductions between 37% and 83 % in qubit number with respect to Bennet syntheses algorithms and complete the majority of our syntheses in less than a second, improving on the current running times of state of the art approaches.
Giacomo Lancellotti, Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi
ICCD2
2024 Design-time methodology for optimizing mixed-precision CPU architectures on FPGA
abstract
Approximate computing can significantly reduce the energy consumption of computing systems. Mixed-precision hardware architectures and precision-tuning tools for software provide the ability to introduce approximations, but when applied separately, they do not give complete control over the accuracy-energy trade-off. The co-optimization of approximations in hardware and software is a complex task, but it promises considerable benefits. We present a methodology for the fast design-time selection of mixed-precision hardware-software combinations that minimize the energy consumption and the area of the target FPGA-based softcore CPUs with configurable support for floating-point and fixed-point arithmetic. Our approach can evaluate configurations more than 2000 times faster than the alternative approach of using gate-level simulation. On benchmarks from the PolyBench suite the identified hardware-software configurations showed improvement of the energy-to-solution metric ranging from 20% to 95%.
Lev Denisov, Andrea Galimberti, Daniele Cattaneo 0002, Giovanni Agosta, Davide Zoni
J. Syst. Archit.4
2023 Clever DAE: Compiler Optimizations for Digital Twins at Scale
abstract
Modeling and simulation are fundamental activities in engineering to facilitate prototyping, verification and maintenance. Declarative modeling languages allow to simulate physical phenomena by expressing them in terms of Differential and Algebraic Equations (DAE) systems. In this paper, we focus on the problem of generating code for performing the numerical integration of the model equations, and in particular on the overhead introduced by external numerical solver libraries. We propose a novel methodology for minimizing the amount of equations which require to be solved through an external solver library, together with the number of computations that are required to computed the Jacobian matrix of the system. Through a prototype LLVM-based compiler, we demonstrate how this approach achieves a linear speed-up in simulation time with respect to the baseline.
Michele Scuttari, Nicola Camillucci, Daniele Cattaneo 0002, Giovanni Agosta, Francesco Casella, Stefano Cherubin, Federico Terraneo
CF4
2023 Hardware and Software Support for Mixed Precision Computing: a Roadmap for Embedded and HPC Systems
abstract
Mixed precision is an approximate computing technique that can be used to trade-off computation accuracy for performance and/or energy. It can be applied to many error-tolerant applications, but manual precision tuning is both tedious and error-prone. Furthermore, the effectiveness of the technique heavily depends on hardware characteristics. Therefore, a hardware/software co-design approach is necessary for an effective exploitation of precision tuning opportunities offered by the applications. In this paper, we propose, based on the state of the art of precision tuning software and mixed precision hardware, a roadmap for the evolution of hardware designs and compiler-based precision tuning support, which is ongoing in the context of the European projects TEXTAROSSA and APROPOS.
William Fornaciari, Giovanni Agosta, Daniele Cattaneo 0002, Lev Denisov, Andrea Galimberti, Gabriele Magnani, Davide Zoni
DATE2
2023 Array-Aware Matching: Taming the Complexity of Large-Scale Simulation Models
abstract
Equation-based modelling is a powerful approach to tame the complexity of large-scale simulation problems. Equation-based tools automatically translate models into imperative languages. When confronted with nowadays’ problems, however, well assessed model translation techniques exhibit scalability issues that are particularly severe when models contain very large arrays. In fact, such models can be made very compact by enclosing equations into looping constructs, but reflecting the same compactness into the translated imperative code is nontrivial. In this paper, we face this issue by concentrating on a key step of equations-to-code translation, the equation/variable matching. We first show that an efficient translation of models with (large) arrays needs awareness of their presence, by defining a figure of merit to measure how much the looping constructs are preserved along the translation. We then show that the said figure of merit allows to define an optimal array-aware matching, and as our main result, that the so stated optimal array-aware matching problem is NP-complete. As an additional result, we propose a heuristic algorithm capable of performing array-aware matching in polynomial time. The proposed algorithm can be proficiently used by model translator developers in the implementation of efficient tools for large-scale system simulation.
Massimo Fioravanti, Daniele Cattaneo 0002, Federico Terraneo, Silvano Seva, Stefano Cherubin, Giovanni Agosta, Francesco Casella, Alberto Leva
ACM Trans. Math. Softw.6
2021 The Italian research on HPC key technologies across EuroHPC
abstract
High-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects.
Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati
CF2
2021 Architecture-aware Precision Tuning with Multiple Number Representation Systems
abstract
Precision tuning trades accuracy for speed and energy savings, usually by reducing the data width, or by switching from floating point to fixed point representations. However, comparing the precision across different representations is a difficult task. We present a metric that enables this comparison, and employ it to build a methodology based on Integer Linear Programming for tuning the data type selection. We apply the proposed metric and methodology to a range of processors, demonstrating an improvement in performance (up to $9 \times)$ with a very limited precision loss $(\lt 2.8$% for 90% of the benchmarks) on the PolyBench benchmark suite.
Daniele Cattaneo 0002, Michele Chiari, Nicola Fossati, Stefano Cherubin, Giovanni Agosta
DAC5
2021 TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascale
abstract
To achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research.
Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola
DSD1
2021 Tunable approximations to control time-to-solution in an HPC molecular docking Mini-App
Davide Gadioli, Gianluca Palermo, Stefano Cherubin, Emanuele Vitali, Giovanni Agosta, Candida Manelfi, Andrea Beccari, Carlo Cavazzoni, Nico Sanna, Cristina Silvano
J. Supercomput.5
2020 A Comb for Decompiled C Code
abstract
Decompilers are fundamental tools to perform security assessments of third-party software. The quality of decompiled code can be a game changer in order to reduce the time and effort required for analysis. This paper proposes a novel approach to restructure the control flow graph recovered from binary programs in a semantics-preserving fashion. The algorithm is designed from the ground up with the goal of producing C code that is both goto-free and drastically reducing the mental load required for an analyst to understand it. As a result, the code generated with this technique is well-structured, idiomatic, readable, easy to understand and fully exploits the expressiveness of C language. The algorithm has been implemented on top of the revng static binary analysis framework. The resulting decompiler, revngc, is compared on real-world binaries with state-of-the-art commercial and open source tools. The results show that our decompilation process introduces between 40% and 50% less extra cyclomatic complexity.
Andrea Gussoni, Alessandro Di Federico, Pietro Fezzardi, Giovanni Agosta
AsiaCCS4
2020 Dynamic Precision Autotuning with TAFFO
abstract
Many classes of applications, both in the embedded and high performance domains, can trade off the accuracy of the computed results for computation performance. One way to achieve such a trade-off is precision tuning—that is, to modify the data types used for the computation by reducing the bit width, or by changing the representation from floating point to fixed point. We present a methodology for high-accuracy dynamic precision tuning based on the identification of input classes (i.e., classes of input datasets that benefit from similar optimizations). When a new input region is detected, the application kernels are re-compiled on the fly with the appropriate selection of parameters. In this way, we obtain a continuous optimization approach that enables the exploitation of the reduced precision computation while progressively exploring the solution space, thus reducing the time required by compilation overheads. We provide tools to support the automation of the runtime part of the solution, leaving to the user only the task of identifying the input classes. Our approach provides a significant performance boost (up to 320%) on the typical approximate computing benchmarks, without meaningfully affecting the accuracy of the result, since the error remains always below 3%.
Stefano Cherubin, Daniele Cattaneo 0002, Michele Chiari, Giovanni Agosta
ACM Trans. Archit. Code Optim.4
2020 Compiler-Based Techniques to Secure Cryptographic Embedded Software Against Side-Channel Attacks
abstract
Side-channel attacks are a concrete and practical threat to the security of computing systems, ranging from high performance platforms to embedded devices. In this paper, we will provide a brief systematization of the current existing approaches to analyze the side-channel vulnerability of an implementation, or automatically implement countermeasures, relying on methodologies typical of compiler systems. We will dedicate a spotlight to a significant progress in the countermeasures techniques which is represented by the application of dynamic compilation techniques to prevent a side-channel attacker from devising a model of the attacked application. We conclude the work highlighting promising research directions in this field.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Fixed point exploitation via compiler analyses and transformations: POSTER
abstract
Fixed point computation represents a key feature in the design process of embedded applications. It is also exploited as a mean to data size tuning for HPC tasks [2]. Since the conversion from floating point to fixed point is generally performed manually, it is time-consuming and error-prone. However, the full automation of such task is currently unfeasible, as existing open source tools are not mature enough for industry adoption. To bridge this gap, we introduce our Tuning Assistant for Floating point to Fixed point Optimization (TAFFO). TAFFO is a toolset of LLVM compiler plugins that automatically converts computations from floating point to fixed point. TAFFO leverages programmer hints to understand the characteristics of the input data, and then performs the code conversion using the most appropriate data types. TAFFO allows programmers to equally apply fine-grained precision tuning to a wide range of programming languages, whereas most current competitors are limited to C. Moreover, it is easily applicable to most embedded [1] and high performance applications [10, 11], and it allows easy maintenance and extensions.
Daniele Cattaneo 0002, Antonio Di Bello, Michele Chiari, Stefano Cherubin, Giovanni Agosta
CF5
2019 Challenges in Deeply Heterogeneous High Performance Systems
abstract
RECIPE (REliable power and time-ConstraInts-aware Predictive management of heterogeneous Exascale systems) is a recently started project funded within the H2020 FETHPC programme, which is expressly targeted at exploring new High-Performance Computing (HPC) technologies. RECIPE aims at introducing a hierarchical runtime resource management infrastructure to optimize energy efficiency and minimize the occurrence of thermal hotspots, while enforcing the time constraints imposed by the applications and ensuring reliability for both time-critical and throughput-oriented computation that run on deeply heterogeneous accelerator-based systems. This paper presents a detailed overview of RECIPE, identifying the fundamental challenges as well as the key innovations addressed by the project, which span run-time management, heterogeneous computing architectures, HPC memory/interconnection infrastructures, thermal modelling, reliability, programming models, and timing analysis. For each of these areas, the paper describes the relevant state of the art as well as the specific actions that the project will take to effectively address the identified technological challenges.
Giovanni Agosta, William Fornaciari, David Atienza 0001, Ramon Canal, Alessandro Cilardo, José Flich, Carles Hernández 0001, Michal Kulczewski, Giuseppe Massari, Rafael Tornero, Marina Zapater
DSD1
2019 Supporting the Scale-Up of High Performance Application to Pre-Exascale Systems: The ANTAREX Approach
abstract
The ANTAREX project developed an approach to the performance tuning of High Performance applications based on an Aspect-oriented Domain Specific Language (DSL), with the goal to simplify the enforcement of extra-functional properties in large scale applications. The project aims at demonstrating its tools and techniques on two relevant use cases, one in the domain of computational drug discovery, the other in the domain of online vehicle navigation. In this paper, we present an overview of the project and of its main achievements, as well as of the large scale experiments that have been planned to validate the approach.
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, Loïc Besnard, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Daniele Cesarini, Stefano Cherubin, Federico Ficarelli, Davide Gadioli, Martin Golasowski, Imane Lasri, Antonio Libri, Candida Manelfi, Jan Martinovic, Gianluca Palermo, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová, Emanuele Vitali
PDP2
2018 Autotuning and adaptivity in energy efficient HPC systems: the ANTAREX toolbox
abstract
Designing and optimizing applications for energy-efficient High Performance Computing systems up to the Exascale era is an extremely challenging problem. This paper presents the toolbox developed in the ANTAREX European project for autotuning and adaptivity in energy efficient HPC systems. In particular, the modules of the ANTAREX toolbox are described as well as some preliminary results of the application to two target use cases. 1
Cristina Silvano, Gianluca Palermo, Giovanni Agosta, Amir H. Ashouri, Davide Gadioli, Stefano Cherubin, Emanuele Vitali, Luca Benini, Andrea Bartolini, Daniele Cesarini, João M. P. Cardoso, João Bispo, Pedro Pinto 0002, Ricardo Nobre, Erven Rohou, Loïc Besnard, Imane Lasri, Nico Sanna, Carlo Cavazzoni, Radim Cmar, Jan Martinovic, Katerina Slaninová, Martin Golasowski, Andrea Beccari, Candida Manelfi
CF3
2018 Embedded Operating System Optimization through Floating to Fixed Point Compiler Transformation
abstract
Architectures targeted at embedded systems often have limited floating point computation capabilities, and in many cases do not provide any hardware support. In this work, we propose a self-contained compiler transformation pass implemented within LLVM to perform floating point to fixed point conversion. This pass is used to optimize the scheduler of the MIOSIX embedded real-time operating system. We compare the proposed approach with the original floating point implementation, a hand-tuned fixed point one, and a solution based on a C++ library for fixed-point arithmetic. Our solution achieves speedups with respect to original floating point implementation up to 3.1×.
Daniele Cattaneo 0002, Antonio Di Bello, Stefano Cherubin, Federico Terraneo, Giovanni Agosta
DSD5
2018 ANTAREX: A DSL-Based Approach to Adaptively Optimizing and Enforcing Extra-Functional Properties in High Performance Computing
abstract
The ANTAREX project relies on a Domain Specific Language (DSL) based on Aspect Oriented Programming (AOP) concepts to allow applications to enforce extra functional properties such as energy-efficiency and performance and to optimize Quality of Service (QoS) in an adaptive way. The DSL approach allows the definition of energy-efficiency, performance, and adaptivity strategies as well as their enforcement at runtime through application autotuning and resource and power management. In this paper, we present an overview of the ANTAREX DSL and some of its capabilities through a number of examples, including how the DSL is applied in the context of one of the project use cases.
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, Loïc Besnard, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Stefano Cherubin, Davide Gadioli, Martin Golasowski, Imane Lasri, Jan Martinovic, Gianluca Palermo, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová, Emanuele Vitali
DSD2
2017 rev.ng: a unified binary analysis framework to recover CFGs and function boundaries
Alessandro Di Federico, Mathias Payer, Giovanni Agosta
CC3
2017 MANGO: Exploring Manycore Architectures for Next-GeneratiOn HPC Systems
abstract
The Horizon 2020 MANGO project aims at exploring deeply heterogeneous accelerators for use in High-Performance Computing systems running multiple applications with different Quality of Service (QoS) levels. The main goal of the project is to exploit customization to adapt computing resources to reach the desired QoS. For this purpose, it explores different but interrelated mechanisms across the architecture and system software. In particular, in this paper we focus on the runtime resource management, the thermal management, and support provided for parallel programming, as well as introducing three applications on which the project foreground will be validated.
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Etienne Cappe, Alessandro Cilardo, Leon Dragic, Alexandre Dray, Alen Duspara, William Fornaciari, Gerald Guillaume, Ynse Hoornenborg, Arman Iranfar, Mario Kovac, Simone Libutti, Bruno Maitre, José Maria Martínez, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Tomás Picornell, Igor Piljic, Anna Pupykina, Federico Reghenzani, Isabelle Staub, Rafael Tornero, Marina Zapater, Davide Zoni
DSD2
2016 A jump-target identification method for multi-architecture static binary translation
abstract
Static binary translation is a technique that allows an executable program for a given architecture to be translated into a different one, with a reduced overhead compared to emulators and dynamic binary translators. The main downside of the static approach lies in the absence of runtime information, which is available in other solutions. In particular, one of the key issues consists in the identification of data and code in the program, and, more specifically, in the detection of basic block start addresses (jump targets). The presence of indirect jump instructions whose target is not immediately evident, in particular due to C switch statements, makes the recovery of jump targets a challenging task.
Alessandro Di Federico, Giovanni Agosta
CASES2
2016 Enabling HPC for QoS-sensitive applications: The MANGO approach
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Alessandro Cilardo, William Fornaciari, Ynse Hoornenborg, Mario Kovac, Bruno Maitre, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Fabrice Roudet, Rafael Tornero, Davide Zoni
DATE2
2016 Autotuning and adaptivity approach for energy efficient Exascale HPC systems: The ANTAREX approach
Cristina Silvano, Giovanni Agosta, Andrea Bartolini, Andrea Beccari, Luca Benini, João Bispo, Radim Cmar, João M. P. Cardoso, Carlo Cavazzoni, Jan Martinovic, Gianluca Palermo, Martin Palkovic, Pedro Pinto 0002, Erven Rohou, Nico Sanna, Katerina Slaninová
DATE2
2016 V2I Cooperation for Traffic Management with SafeCop
abstract
The Safe Cooperating Cyber-Physical Systems using Wireless Communication (SafeCop) project addresses safety-related issues in cooperating cyber-physical systems. These systems, characterised by wireless communications, multiple stakeholders, and variable operating environments, are called Cooperative Open Cyber-Physical Systems (CO-CPS). CO-CPSs can successfully address several societal challenges -- cooperative vehicles have been shown to reduce fuel consumption as well as the number of accidents. A vehicle-to-infrastructure (V2I) cooperation for the traffic management scenario is therefore considered as a key use cases of SafeCop. In this paper, we outline the V2I traffic management scenario, assess the research goals that arise from it, and provide an overview of the architecture of the demonstrator, as well as a roadmap for its development and evaluation.
Giovanni Agosta, Alessandro Barenghi, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Stefano Delucchi, Massimo Massa, Maurizio Mongelli, Enrico Ferrari, Leonardo Napoletani, Luciano Bozzi, Carlo Tieri, Dajana Cassioli, Luigi Pomante
DSD1
2016 The M2DC Project: Modular Microserver DataCentre
abstract
The Modular Microserver DataCentre (M2DC) project will investigate, develop and demonstrate a modular, highly-efficient, cost-optimized server architecture composed of heterogeneous microserver computing resources, being able to be tailored to meet requirements from various application domains such as image processing, cloud computing or HPC. M2DC will be built on three main pillars: a flexible server architecture that can be easily customised, maintained and updated, advanced management strategies and system efficiency enhancements (SEE), well-defined interfaces to surrounding software data centre ecosystem.
Mariano Cecowski, Giovanni Agosta, Ariel Oleksiak, Michal Kierzynka, Micha vor dem Berge, Wolfgang Christmann, Stefan Krupop, Mario Porrmann, Jens Hagemeyer, René Griessl, Meysam Peykanu, Lennart Tigges, Sven Rosinger, Daniel Schlitt, Christian Pieper, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Robert Plestenjak, Justin Cinkelj, Loïc Cudennec, Thierry Goubier, Jean-Marc Philippe, Udo Janssen, Chris Adeniyi-Jones
DSD2
2016 Encasing block ciphers to foil key recovery attempts via side channel
abstract
Providing efficient protection against energy consumption based side channel attacks (SCAs) for block ciphers is a relevant topic for the research community, as current overheads are in the 100× range. Unprofiled SCAs exploit information leakage from the outmost rounds of a cipher; we propose a solution encasing it between keyed transformations amenable to an efficient SCA protection. Our solution can be employed as a drop in replacement for an unprotected implementation, or be retrofit to an existing one, while retaining communication capabilities with legacy insecure endpoints. Experiments on a Cortex-M4 µC, show performance improvements in the range of 60×, compared with available solutions.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
ICCAD1
2015 Information leakage chaff: feeding red herrings to side channel attackers
abstract
A prominent threat to embedded systems security is represented by side-channel attacks: they have proven effective in breaching confidentiality, violating trust guarantees and IP protection schemes. State-of-the-art countermeasures reduce the leaked information to prevent the attacker from retrieving the secret key of the cipher. We propose an alternate defense strategy augmenting the regular information leakage with false targets, quite like chaff countermeasures against radars, hiding the correct secret key among a volley of chaff targets. This in turn feeds the attacker with a large amount of invalid keys, which can be used to trigger an alarm whenever the attack attempts a content forgery using them, thus providing a reactive security measure. We realized a LLVM compiler pass able to automatically apply the proposed countermeasure to software implementations of block ciphers. We provide effectiveness and efficiency results on an AES implementation running on an ARM Cortex-M4 showing performance overheads comparable with state-of-the-art countermeasures.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
DAC1
2015 Playful Supervised Smart Spaces (P3S) - A Framework for Designing, Implementing and Deploying Multisensory Play Experiences for Children with Special Needs
abstract
Our research explores novel forms of smart spaces that support full body interaction with smart objects instrumented with audio, light, and motion sensors and actuators, virtual worlds on medium-large displays, and smart lights, and are integrated with cloud services for remote supervision and analysis of user behaviour. In this paper, we present the architecture and coordination infrastructure created in the Playful Supervised Smart Space (P3S) project, as well as initial designs for the Smart Object and Smart Space Gateway Components.
Giovanni Agosta, Luca Borghese, Carlo Brandolese, Francesco Clasadonte, William Fornaciari, Franca Garzotto, Mirko Gelsomini, Matteo Grotto, Cristina Frà, Danny Noferi, Massimo Valla
DSD1
2015 Tailoring instruction-set extensions for an ultra-low power tightly-coupled cluster of OpenRISC cores
abstract
Baseline RISC instruction sets for ultra-low power processors are constantly being tuned to reduce cycle count when executing computation-intensive applications. Performance improvements often come at a non-negligible price in terms of area and critical path length and imply deeper pipelines and complex memory interfaces. This penalizes control-intensive code execution and significantly increases cost and complexity of building multi-core clusters. In addition, some extensions are not easily exploited by compilers and may increase code development effort, especially when considering parallel applications. In this paper we describe our efforts in enhancing a baseline open ISA (OpenRISC) and its LLVM compiler back-end to significantly reduce execution cycles while minimizing the impact on core micro-architecture complexity, number of pipeline stages, area and power. In addition, we improved the core micro-architecture to streamline its integration in a tightly-coupled cluster, sharing instruction cache and data memory, thereby further enhancing parallel execution efficiency. The combined effect of ISA, compiler and micro-architecture evolution gives an average energy efficiency boost of 59% on vector intensive code and 41% otherwise, at an area and power increase of 2.3% and 18% on a four-core processor cluster.
Michael Gautschi, Andreas Traber, Antonio Pullini, Luca Benini, Michele Scandale, Alessandro Di Federico, Michele Beretta 0002, Giovanni Agosta
VLSI-SoC8
2015 OpenCL performance portability for general-purpose computation on graphics processor units: an exploration on cryptographic primitives
abstract
Summary The modern trend toward heterogeneous many‐core architectures has led to high architectural diversity in both high performance and high‐end embedded systems. To effectively exploit the computational resources of such a wide range of architectures, programming languages and APIs such as OpenCL have become increasingly popular. Although OpenCL provides functional code portability and the ability to fine tune the application to the target hardware, providing performance portability is still an open problem. Thus, many research works have investigated the optimization of specific combinations of application and target platform. In this paper, we aim at leveraging the experience obtained in the implementation of algorithms from the cryptography domain to provide a set of guidelines for modern many‐core heterogeneous architecture performance portability and to establish a base on which domain‐specific languages and compiler transformations could be built in the near future. We study algorithmic choices and the effect of compiler transformations on three representative applications in the chosen domain on a set of seven target platforms. To estimate how well the application fits the architecture, we define a metric of computational intensity both for the architecture and the application implementation. Besides being useful to compare either different implementation or algorithmic choices and their fitness to a specific architecture, it can also be useful to the compiler to guide the code optimization process. Copyright © 2014 John Wiley & Sons, Ltd.
Giovanni Agosta, Alessandro Barenghi, Alessandro Di Federico, Gerardo Pelosi
Concurr. Comput. Pract. Exp.1
2015 Trace-based schedulability analysis to enhance passive side-channel attack resilience of embedded software
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
Inf. Process. Lett.1
2015 The MEET Approach: Securing Cryptographic Embedded Software Against Side Channel Attacks
abstract
We propose an efficient and effective methods to secure software implementations of cryptographic primitives on low-end embedded systems, against passive side channel attacks relying on the observation of power consumption or electro-magnetic emissions. The proposed approach exploits a modified LLVM compiler toolchain to automatically generate a secure binary characterized by a randomized execution flow. We improve the current state-of-the-art in dynamic executable code countermeasures removing the requirement of a writable code segment, and reducing the countermeasure overhead. Also, we provide a new method to refresh the random values employed in the share splitting approaches to lookup table protection. Finally, we devise an automated approach to protect spill actions onto the main memory, which are inserted by the compiler backend register allocator when there is a lack of available registers, thus, removing the need for manual assembly inspection. We report a validation of the performances of our approach on all the current ISO-standard block ciphers, employing an ARM Cortex-M4 based microcontroller as the validation platform.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 A Multiple Equivalent Execution Trace Approach to Secure Cryptographic Embedded Software
abstract
We propose an efficient and effective method to secure software implementations of cryptographic primitives on low-end embedded systems, against passive side-channel attacks relying on the observation of power consumption or electro-magnetic emissions. The proposed approach exploits a modified llvm compiler toolchain to automatically generate a secure binary characterized by a randomized execution flow. Also, we provide a new method to refresh the random values employed in the share splitting approaches to lookup table protection, addressing a currently open issue. We improve the current state-of-the-art in dynamic executable code countermeasures removing the requirement of a writeable code segment, and reducing the countermeasure overhead.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
DAC1
2014 OpenCL Application Auto-tuning and Run-Time Resource Management for Multi-core Platforms
abstract
To support adaptivity of data parallel applications on multi-core platforms, we propose a framework based on the combination of OpenCL application auto-tuning and run-time resource management. The framework addresses computationally intensive multimedia OpenCL applications. For these target applications, we show that application auto-tuning, based on design-time analysis, can become synergistic with run-time resource management. In the proposed framework, run-time decisions are taken by each application, autonomously, to achieve system adaptivity. This paper describes the methodology and related toolchain, defined during the 2PARMA European project, based on the integration of independent tools to provide effective compilation of OpenCL code, multi-objective design space exploration, application monitoring and tuning and system-wide run-time resource management. Experimental results are reported for design optimization of an OpenCL stereo-matching application and then for a resource contention scenario where multiple stereo-matching applications are executed on the same platform with different run-time requirements.
Davide Gadioli, Simone Libutti, Giuseppe Massari, Edoardo Paone, Michele Scandale, Patrick Bellasi, Gianluca Palermo, Vittorio Zaccaria, Giovanni Agosta, William Fornaciari, Cristina Silvano
ISPA9
2014 Differential Fault Analysis for Block Ciphers: an Automated Conservative Analysis
abstract
Differential Fault Analysis (DFA) exploits the differences between a correct and a faulty output of a cipher implementation to derive the secret parameters. All the current DFA techniques are tailored to the cipher being attacked and do not provide a general framework. We propose an automated general framework to assess the vulnerability of block ciphers against DFAs, providing a conservative analysis on the attacker capabilities and a practical lower bound on the attacker effort required to extract the secret-key. The proposed technique is based on dataflow analysis of software cipher implementations and has been implemented as a pass of the llvm compiler infrastructure. This work shows how the automated tool we developed is able to detect which and how many faults an attacker can exploit to recover the values of portions of the secret-key material employed by a standard block cipher, validating the effectiveness of our approach. The precise analysis provided by our tool allows to apply the computationally demanding fault attack countermeasures only to the vulnerable portions of the cipher.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
SIN1
2013 Compiler-based side channel vulnerability analysis and optimized countermeasures application
abstract
Modern embedded systems manage sensitive data increasingly often through cryptographic primitives. In this context, side-channel attacks, such as power analysis, represent a concrete threat, regardless of the mathematical strength of a cipher. Evaluating the resistance against power analysis of cryptographic implementations and preventing it, are tasks usually ascribed to the expertise of the system designer. This paper introduces a new security-oriented data-flow analysis assessing the vulnerability level of a cipher with bit-level accuracy. A general and extensible compiler-based tool was implemented to assess the instruction resistance against power-based side-channels. The tool automatically instantiates the essential masking countermeasures, yielding a x2.5 performance speedup w.r.t. protecting the entire code.
Giovanni Agosta, Alessandro Barenghi, Massimo Maggi, Gerardo Pelosi
DAC1
2013 On Task Assignment in Data Intensive Scalable Computing
Giovanni Agosta, Gerardo Pelosi, Ettore Speziale
JSSPP1
2013 Enhancing Passive Side-Channel Attack Resilience through Schedulability Analysis of Data-Dependency Graphs
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi, Michele Scandale
NSS1
2012 A code morphing methodology to automate power analysis countermeasures
abstract
We introduce a general framework to automate the application of countermeasures against Differential Power Attacks aimed at software implementations of cryptographic primitives. The approach enables the generation of multiple versions of the code, to prevent an attacker from recognizing the exact point in time where the observed operation is executed and how such operation is performed. The strategy increases the effort needed to retrieve the secret key through hindering the formulation of a correct hypothetical consumption to be correlated with the power measurements. The experimental evaluation shows how a DPA attack against OpenSSL AES implementation on an industrial grade ARM-based SoC is hindered with limited performance overhead.
Giovanni Agosta, Alessandro Barenghi, Gerardo Pelosi
DAC1
2012 Architecture Optimization of Application-Specific Implicit Instructions
abstract
Dynamic configuration of application-specific implicit instructions has been proposed to better exploit the available parallelism at the instruction level in pipelined processors. The support of such implicit instruction issue-requires the pipeline to be extended with a trigger table that describes the instruction implicitly issued as a response to a value written into a triggering register by a triggering instruction (which may be an add or sub instruction). In this article, we explore the design optimization of the trigger table to maximize the number of instructions that can be implicitly issued while keeping the limited size of the trigger table. The concept of implicitly issued instruction has been formally defined by considering the inter-basic block analysis of control and data dependencies. A compilation tool chain has been developed to automatically identify the optimization opportunities, taking into account the constraints imposed by control and data dependencies as well as by architectural limitations. The proposed solutions have been applied to the case of a baseline scalar MIPS processor where, for the selected set of benchmarks (DSPStone and Mibench/automotive), we obtained an average speedup of 17%.
Andrea Di Biagio, Giovanni Agosta, Martino Sykora, Cristina Silvano
ACM Trans. Embed. Comput. Syst.2
2011 Exploiting Thread-Data Affinity in OpenMP with Data Access Patterns
Andrea Di Biagio, Ettore Speziale, Giovanni Agosta
Euro-Par (1)3
2010 A highly flexible, parallel virtual machine: design and experience of ILDJIT
abstract
Abstract ILDJIT, a new‐generation dynamic compiler and virtual machine designed to support parallel compilation, is introduced here. Our dynamic compiler targets the increasingly popular ECMA‐335 specification. The goal of this project is twofold: on one hand, it aims at exploiting the parallelism exposed by multi‐core architectures to hide the dynamic compilation latencies by pipelining compilation and execution tasks; on the other hand, it provides a flexible, modular and adaptive framework for dynamic code optimization. The ILDJIT organization and the compiler design choices are presented and discussed highlighting how adaptability and extensibility can be achieved. Thanks to the compilation latency masking effect of the pipeline organization, our dynamic compiler is able to mask most of the compilation delay, when the underlying hardware exposes sufficient parallelism. Even when running on a single core, the ILDJIT adaptive optimization framework manages to speedup the computation with respect to other open‐source implementations of ECMA‐335. Copyright © 2010 John Wiley & Sons, Ltd.
Simone Campanoni, Giovanni Agosta, Stefano Crespi-Reghizzi, Andrea Di Biagio
Softw. Pract. Exp.2
2009 Dynamic Look Ahead Compilation: A Technique to Hide JIT Compilation Latencies in Multicore Environment
Simone Campanoni, Martino Sykora, Giovanni Agosta, Stefano Crespi-Reghizzi
CC3
2009 Design of a parallel AES for graphics hardware using the CUDA framework
abstract
Web servers often need to manage encrypted transfers of data. The encryption activity is computationally intensive, and exposes a significant degree of parallelism. At the same time, cheap multicore processors are readily available on graphics hardware, and toolchains for development of general purpose programs are being released by the vendors. In this paper, we propose an effective implementation of the AES-CTR symmetric cryptographic primitive using the CUDA framework. We provide quantitative data for different implementation choices and compare them with the common CPU-based OpenSSL implementation on a performance-cost basis. With respect to previous works, we focus on optimizing the implementation for practical application scenarios, and we provide a throughput improvement of over 14 times. We also provide insights on the programming knowledge required to efficiently exploit the hardware resources by exposing the different kinds of parallelism built in the AES-CTR cryptographic primitive.
Andrea Di Biagio, Alessandro Barenghi, Giovanni Agosta, Gerardo Pelosi
IPDPS3
2009 Fast Disk Encryption through GPGPU Acceleration
abstract
We present the design and performance analysis of a GPU-optimized implementation of a disk encryption application employing the XTS mode of operation applied together with the Twofish algorithm within the well-known TrueCrypt suite. We show how to correctly tune the design parameters, including data allocation, thread packing, and parallelization strategy. Overall, our implementation of TrueCrypt running on a NVidia GTX260 GPU outperforms by 67% the baseline implementation running on a four core CPU.
Giovanni Agosta, Alessandro Barenghi, Fabrizio De Santis, Andrea Di Biagio, Gerardo Pelosi
PDCAT1
2009 A Transform-Parametric Approach to Boolean Matching
abstract
In this paper, we address the problem of P-equivalence Boolean matching. We outline a formal framework that unifies some of the spectral- and canonical-form-based approaches to the problem. As a first major contribution, we show how these approaches are particular cases of a single generic algorithm, parametric with respect to a given linear transformation of the input function. As a second major contribution, we identify a linear transformation that can be used to significantly speed up Boolean matching with respect to the state of the art. Experimental results show that, on average, over a large set of randomly generated Boolean functions, our approach is up to five times faster than the main competitor on 20-variable input and scales better, allowing to match even larger components. Finally, as a representative set of Boolean functions that arise in practice, we considered multiplexers with three, four, and five selectors and functions extracted from the ISCAS85 benchmarks suite with a number of input variables up to 20. The reported performance results show that our approach allows us to halve the canonizing computation time.
Giovanni Agosta, Francesco Bruschi, Gerardo Pelosi, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Static Analysis of Transaction-Level Communication Models
abstract
We propose a methodology for the early estimation of communication implementation choice effects, starting from an abstract transaction-level system model (TLM). The reference version of the TLM considered is the Open SystemC initiative library. The methodology is based on the computation of metrics that abstract useful information from the initial system model. The metrics are precisely defined upon a general formal model of transaction-level system descriptions. A set of design problems of relevant interest, such as shared communication resource assignment, pipelining partitioning, bandwidth, and latency constraint estimation, is considered to show some potential applications of the metrics proposed.
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 A Unified Approach to Canonical Form-based Boolean Matching
abstract
In this paper, we face the problem of P-equivalence Boolean matching. We outline a formal framework that unifies some of the canonical form-based approaches to the problem.
Giovanni Agosta, Francesco Bruschi, Gerardo Pelosi, Donatella Sciuto
DAC1
2007 A Domain Specific Language for Cryptography
Giovanni Agosta, Gerardo Pelosi
FDL1
2007 Countermeasures against Branch Target Buffer Attacks
abstract
Branch Prediction Analysis has been recently proposed as an attack method to extract the key from software implementations of the RSA public key cryptographic algorithm. In this paper, we describe several solutions to protect against such an attack and analyze their impact on the execution time of the cryptographic algorithm. We show that the code transformations required for protection against branch target buffer attacks can be automated and impose only a negligible performance penalty.
Giovanni Agosta, Luca Breveglieri, Gerardo Pelosi, Israel Koren
FDTC1
2007 An efficient cost-based canonical form for Boolean matching
abstract
In this paper, we present new canonical forms for P, NP and NPN equivalence relations on boolean functions. The canonical forms are based on the minimization of a cost function. With respect to previous approaches based on cost minimization, our function allows the minimization algorithm to explore a reduced solution space. This reduction is obtained by partitioning the columns of the boolean function through an equivalence relation. NP and NPN canonical forms are obtained by means of a preprocessing step of negligible computational overhead.
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
ACM Great Lakes Symposium on VLSI1
2006 Adaptive Metrics for System-Level Functional Partitioning
Giovanni Agosta, Marco D. Santambrogio, Seda Ogrenci Memik
FDL1
2005 Aspect Orientation in System Level Design
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
FDL1
2003 Static analysis of transaction-level models
abstract
The introduction of design languages, such as SystemC 2.0, that allow the modelling of digital systems at the transaction level will impose some major changes to the design flows. Since these formalisms allow for a higher level of abstraction in the systems description, new methodological tools will be needed to support all design phases. The goal of this paper is twofold: first we formalize in an abstract way a significant set of features of a Transaction Level Model, according to the SystemC 2.0 formalism. Then, upon this model we define numerical metrics that can provide useful information in the analysis of the system-level specifications. In particular these metrics are useful in the design exploration phase, to define the main characteristics of the hardware and software architectures
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
DAC1