EDBT 2026 Demo / reviewers in the wild / expert
William Fornaciari
dblp:16/4949
· DBLP profile ↗
100ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0001-8294-730XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 87 · 8 first-author · 22 since 2021Software engineering, systems software and programming languages · 17 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Artificial intelligence and machine learning · 2Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantifying Compiler-induced Reliability Loss in Software-Implemented Hardware Fault ToleranceabstractCompiler mechanisms for Software-Implemented Hardware Fault Tolerance (SIHFT) offer a cost-effective solution for reliability, paving the way towards the adoption of Commercial Off-The-Shelf (COTS) components in safety-critical environments. However, default compiler optimizations can remove the SIHFT-induced redundancy and checks. For this reason, the use of compiler optimizations was discouraged in the literature. This article presents a comprehensive study of the reliability degradation introduced by LLVM’s O2 optimization pipeline when using a state-of-the-art SIHFT tool. We quantify, via RTL fault injection, the impact of O2 at different optimization stages, which identified a data corruption rate increase by up to $48 \times$. We also propose a static exploration methodology to identify the LLVM passes that harm the reliability. Then, we remove these harmful passes from the optimization pipeline, demonstrating how to tune optimization pipelines to make SIHFT successful even in the presence of compiler optimizations. Davide Baroffio, Johannes Geier, Federico Reghenzani, Ulf Schlichtmann, William Fornaciari |
ASP-DAC | 5 |
| 2026 | VTS2026 Student Forum
Davide Baroffio, Federico Reghenzani, William Fornaciari, Dipal Halder, Sandip Ray |
VTS | 3 |
| 2026 | Late Breaking Results- Proton Beam Experiments of Compiler-based Hardware Fault Tolerance
Emilio Corigliano, Davide Baroffio, Federico Reghenzani, Tomas Antonio López, William Fornaciari |
VTS | 5 |
| 2025 | Evaluating Compiler-Based Reliability with Radiation Fault InjectionabstractCompiler-based fault tolerance is a cost-effective and flexible family of solutions that transparently improves software reliability. This paper evaluates a compiler tool for fault detection via laser injection and α-particle exposure. A novel memory allocation strategy is proposed to mitigate the effects of multi-bit upsets. We integrated the detection mechanism with a recovery solution based on mixed-criticality scheduling. The results demonstrate the error detection and recovery capabilities in realistic scenarios: reducing undetected errors, enhancing system reliability, and advancing software-implemented fault tolerance. Davide Baroffio, Tomas Antonio López, Federico Reghenzani, William Fornaciari |
DATE | 4 |
| 2025 | Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and ReliabilityabstractModern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners. Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo |
DSD | 8 |
| 2025 | Enabling Smart Urban Mobility with Edge AIabstractRoad accidents are a major cause of death in urban areas. Cooperative Intelligent Transport Systems may mitigate or prevent road accidents by providing drivers with relevant and timely warnings, thanks to V2X communications. However, this requires fast detection of dangers, performed at the edge, to minimize transmission latencies and ensure the timeliness of the warnings. The PMDI project extends the STEP platform for intelligent mobility to support real-time operations and ETSI message generation, and provides integration with mobile and embedded systems. Andrea Tuscano, Paolo Giuseppetti, Pietro Amato, Alessandro Solinas, Federico Saluz, Andrea Tessieri, Mario Pedol, Manuel Pernigotto, Massimo Fioravanti, Giovanni Agosta, William Fornaciari, Paolo Maffezzoni, Fabio Salice, Irene Amerini, Francesco Pro, Paolo Satta, Giovanni Trovini |
DSD | 12 |
| 2025 | Software Techniques for Soft Error Resilience: the ASTRAEUS projectabstractASTRAEUS project aims to improve the use of Commercial-Off-The-Shelf (COTS) devices in space telecommunication applications. The goal is to develop specialized radiation mitigation techniques for both hardware and software components with a special focus on the latter. The aim is to demonstrate the feasibility and reliability of these techniques, enabling their future use in telecommunication payload processing units. The SIHFT (Software Implemented Hardware Fault Tolerance) approach will be enforced, where some proper modification to a conventional compilation toolchain, will make possible the identification of temporary fault and the adoption of fault tolerant solutions. The successful implementation of this project will de-risk the use of software-based radiation mitigation techniques and foster the adoption of high-performance while cost-effective COTS electronics in space applications. Federico Reghenzani, Davide Baroffio, Emilio Corigliano, William Fornaciari, Giancarlo Storti Gajani, Paolo Maffezzoni, Antonino Catanese, Alessandro Balossino, Marco Giuliani |
DSD | 4 |
| 2025 | Laser and Radiation Testing of Compiler-Based Protection for Multi-Bit UpsetsabstractSoftware-Implemented Hardware Fault Tolerance (SIHFT) is advantageous in critical systems where hardware solutions cannot be used due to competing non-functional constraints. Recent works have focused on developing compiler-based protection mechanisms, relying on debuggers and other software mechanisms to introduce single event upsets. In this work, we test a compiler-based technique using physical radiation testing methods, including laser fault injection and alpha-particle exposure. During this evaluation, we identified previously unknown issues that required further development, including a novel memory allocation strategy for improved reliability. Furthermore, we integrated this fault detection solution with a hard real-time recovery mechanism that exploits mixed-criticality scheduling to demonstrate the overall system recovery capabilities. The results show the effectiveness of the proposed approach in detecting faults even under real-world radiation conditions, representing an important step toward the maturity of SIHFT techniques. Davide Baroffio, Tomas Antonio López, Federico Reghenzani, William Fornaciari |
ICCD | 4 |
| 2025 | Modern Llvm-Based Compiler Autotuning for Wcet OptimizationabstractThe problem of compiler optimization selection and ordering, known in the literature as compiler autotuning, has been tackled many times for average-case execution time reduction. Optimizing the WCET is becoming a prominent problem for modern hard real-time systems, where the difficulties in accurate WCET estimation hinder the full exploitation of computing platform capabilities. In this article, we propose a novel methodology and a tool based on LLVM for iterative WCET-driven compiler autotuning, which is the first strategy to operate at function-level granularity and to consider not only the selection of optimization passes, but also their ordering. Our findings show that standard optimization levels$\mathrm{O} 0, \mathrm{O} 1, \mathrm{O} 2$, and O 3 are suboptimal when targeting the WCET, and that a per-function selection and ordering of the transformations is necessary. Experimental results show that our approach outperforms the standard optimizations and opens up new directions for future research. Gabriele Magnani, Davide Baroffio, Federico Reghenzani, Giovanni Agosta, William Fornaciari |
RTSS | 5 |
| 2025 | Calibration-Free Feedforward Temperature Compensation for Wireless Clock SynchronisationabstractWireless systems are entering the arena of time-critical applications, including industrial controls. This makes clock synchronisation vital. A relevant source of synchronisation errors, especially in heavy-duty applications and harsh environments, is given by temperature variations that affect the quartz oscillators. This article takes a control theory-based approach and proposes to augment clock synchronisation schemes—that are typically feedback-based—with a feedforward compensation path using temperature measurements to quickly react to temperature change. Unlike previous approaches relying on per-device data, our solution takes advantage of the combined feedforward-feedback nature of the control scheme to perform compensation using only nominal crystal parameters. When implemented in conjunction with the FLOPSYNC-2 feedback-based clock synchronisation scheme and tested both in simulation and experimentally, the approach resulted in a greatly reduced synchronisation error during temperature transients. In detail, Monte Carlo simulations show that the proposed solution can reduce the peak clock synchronisation error on average by 5.3 \(\times\) and in the worst case by 2.2 \(\times\) , thus proving applicable and effective without the need for a costly per-device calibration. When tested experimentally on a network of sensor nodes, a 3.3 \(\times\) peak clock synchronisation error improvement was observed, confirming the simulation predictions. Federico Terraneo, Zaigham Khalid, William Fornaciari, Alberto Leva |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2024 | The TEXTAROSSA Project: Cool all the Way Down to the HardwareabstractThe TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research. Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo |
DSD | 14 |
| 2024 | Enhanced Compiler Technology for Software-based Hardware Fault DetectionabstractSoftware-Implemented Hardware Fault Tolerance (SIHFT) is a modern approach for tackling random hardware faults of dependable systems employing solely software solutions. This work extends an automatic compiler-based SIHFT hardening tool called ASPIS, enhancing it with novel protection mechanisms and overhead-reduction techniques, also providing an extensive analysis of its compliance with the non-trivial workload of the open-source Real-Time Operating System FreeRTOS. A thorough experimental fault-injection campaign on an STM32 board shows how the system achieves remarkably high tolerance to single-event upsets and a comparison between the SIHFT mechanisms implemented summarises the tradeoff between the overhead introduced and the detection capabilities of the various solutions. Davide Baroffio, Federico Reghenzani, William Fornaciari |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Mixed-Criticality with Integer Multiple WCETs and Dropping Relations: New Scheduling ChallengesabstractScheduling Mixed-Criticality (MC) workload is a challenging problem in real-time computing. Earliest Deadline First Virtual Deadline (EDF-VD) is one of the most famous scheduling algorithm with optimal speedup bound properties. However, when EDF-VD is used to schedule task sets using a model with additional or relaxed constraints, its scheduling properties change. Inspired by an application of MC to the scheduling of fault tolerant tasks, in this article, we propose two models for multiple criticality levels: the first is a specialization of the MC model, and the second is a generalization of it. We then show, via formal proofs and numerical simulations, that the former considerably improves the speedup bound of EDF-VD. Finally, we provide the proofs related to the optimality of the two models, identifying the need of new scheduling algorithms. Federico Reghenzani, William Fornaciari |
ASP-DAC | 2 |
| 2023 | Hardware and Software Support for Mixed Precision Computing: a Roadmap for Embedded and HPC SystemsabstractMixed precision is an approximate computing technique that can be used to trade-off computation accuracy for performance and/or energy. It can be applied to many error-tolerant applications, but manual precision tuning is both tedious and error-prone. Furthermore, the effectiveness of the technique heavily depends on hardware characteristics. Therefore, a hardware/software co-design approach is necessary for an effective exploitation of precision tuning opportunities offered by the applications. In this paper, we propose, based on the state of the art of precision tuning software and mixed precision hardware, a roadmap for the evolution of hardware designs and compiler-based precision tuning support, which is ongoing in the context of the European projects TEXTAROSSA and APROPOS. William Fornaciari, Giovanni Agosta, Daniele Cattaneo 0002, Lev Denisov, Andrea Galimberti, Gabriele Magnani, Davide Zoni |
DATE | 1 |
| 2023 | FCPP+Miosix: Scaling Aggregate Programming to Embedded SystemsabstractAs the density of nodes capable of sensing, computing and actuation increases, it becomes increasingly useful to model an entire network of physical devices as a single, continuous space-time computing machine. The emergent behaviour of the whole software system is then induced by local computations deployed within each node and by the dynamics of the information diffusion. A relevant example of this distribution model is given byaggregate programmingand its minimal set of functional constructs used to manipulate distributed data structures evolving over space and time, and resulting in robustness to changes. In this paper, we propose the first implementation of the aggregate computing paradigm targeting microcontrollers, by integrating FCPP, a C++ implementation of the paradigm, with Miosix, a modern operating system for microcontrollers with full C++ support. To the best of the author's knowledge, we are the first to present results on the effectiveness of FCPP in an embedded operating system setting as opposed to a simulation environment, thus considering tight memory and computational constraints and accounting for packet losses due to nonidealities of the radio channel. We implemented and tested on a network of WandStem nodes two benchmark applications: a network connectivity checker for network planning and preventive maintenance, and a decentralised contact tracing application. Additionally, we show that common problems in sensor networks such as neighbour discovery, construction of a graph of the network topology, coarse grain clock synchronisation as well as network monitoring and the collection of statistics (such as memory occupation data) can be easily performed thanks to the expressive semantics of aggregate programming. Giorgio Audrito, Federico Terraneo, William Fornaciari |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | On the use of hardware accelerators in QC-MDPC code-based cryptographyabstractPublic-key cryptography (PKC) allows exchanging keys over an insecure channel without sharing a secret key. However, quantum computers threaten to break traditional PKC, thus, to mitigate such risk, post-quantum cryptography (PQC) aims to develop cryptosystems that are secure against attacks from quantum and classical computers. BIKE [1] is a key encapsulation mechanism (KEM) based on quasi-cyclic moderate-density parity-check (QC-MDPC) codes that is a candidate within the NIST standardization process to identify a set of PQC algorithms [4]. Figure 1 depicts the key exchange between two client and server nodes, which requires the sequential execution of the key generation, encapsulation, and decapsulation KEM primitives. Key generation and decapsulation are performed on the client side, while encapsulation is carried out by the server. Despite the vast literature targeting efficient hardware support for BIKE, each proposal delivered computing platforms meant either to maximize performance or minimize resource utilization. Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni |
CF | 4 |
| 2022 | FPGA implementation of BIKE for quantum-resistant TLSabstractThe recent advances in quantum computers impose the adoption of post-quantum cryptosystems into secure communication protocols. This work proposes two FPGA-based, client- and server-side hardware architectures to support the integration of the BIKE post-quantum KEM within TLS. Thanks to the parametric hardware design, the paper explores the best option between hardware and software implementations, given a set of available hardware resources and a realistic use-case scenario. The experimental evaluation comparing our client and server designs against the reference AVX2 and hardware implementations of BIKE highlighted two aspects. First, the proposed client and server architectures outperform the reference hardware implementation of BIKE by eight and four times, respectively. Second, the performance comparison between our client and server designs against the reference AVX2 implementation strongly depends on the available resource. Our solution is almost twice as fast as the AVX2 implementation while implemented on the Artix-7 200 FPGA, while it is up to six times slower when targeting smaller FPGAs, thus motivating a careful analysis of the available hardware resources and the optimization of the design's parallelism before opting for hardware support. Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni |
DSD | 4 |
| 2022 | A Mixed-Criticality Approach to Fault Tolerance: Integrating Schedulability and Failure RequirementsabstractMixed-Criticality (MC) systems have been widely studied in the past decade, majorly due to their potential to consolidate applications with different criticality levels onto the same platform. In the original design proposed by Vestal, a target probability of failure per hour specified by certification requirements is assigned to each criticality level. These requirements have been mainly conceived for hardware faults. Software fault tolerance techniques are available to mitigate hardware faults, but their adaptation to real-time systems is challenging due to the introduced overhead. This paper proposes an extension to the traditional MC scheduling theory to implement fault tolerance strategies against transient faults, with the goal of complying with both failure and timing requirements. In particular, we introduce the dropping relationships that generalize the concept of criticality and allow, on the one hand, to improve the schedulability analysis, on the other, to control the dependency between tasks satisfying the certification requirements. The simulation study shows a schedulability ratio improvement of 20-30% compared to classical scheduling while maintaining compliance with failure requirements. Federico Reghenzani, Zhishan Guo, Luca Santinelli, William Fornaciari |
RTAS | 4 |
| 2022 | 3D-ICE 3.0: Efficient Nonlinear MPSoC Thermal Simulation With Pluggable Heat Sink ModelsabstractThe increasing power density in modern high-performance multiprocessor System-on-Chip (MPSoC) is fueling a revolution in thermal management. On the one hand, thermal phenomena are becoming a critical concern, making accurate and efficient simulation a necessity. On the other hand, a variety of physically heterogeneous solutions is coming into play: liquid, evaporative, thermoelectric cooling, and more. A new generation of simulators, with unprecedented flexibility, is thus required. In this article, we present 3D-ICE 3.0, the first thermal simulator to allow for accurate nonlinear descriptions of complex and physically heterogeneous heat dissipation systems, while preserving the efficiency of latest compact modeling frameworks at the silicon die level. 3D-ICE 3.0 allows designers to extend the thermal simulator with new heat sink models while simplifying the time-consuming step of model validation. The support for nonlinear dynamic models is included, for instance, to accurately represent variable coolant flows. Our results present validated models of a commercial water heat sink and an air heat sink plus fan that achieve an average error below 1 °C and simulate, respectively, up to$3\times $and$12\times $faster than the real physical phenomena. Federico Terraneo, Alberto Leva, William Fornaciari, Marina Zapater, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Design of Side-Channel-Resistant Power MonitorsabstractIn modern computing platforms, power monitors (PwrMons) are employed to deliver online power estimates to support different runtime power-performance optimization methodologies. However, the possibility of setting up a successful side-channel attack by analyzing the power estimates imposes the use of a suitable and systematic approach in the design of such PwrMons. This article proposes a design methodology to automatically identify and implement side-channel-resistant PwrMons at the hardware level, for generic computing platforms. The methodology works by designing a PwrMon for which the switching activity of the signals used to compute the power estimates is not a function of both the secret key and the plaintext/ciphertext values processed by the computing platform. According to the most recent standardized methodologies to assess the side-channel security, our experimental validation leverages both correlation power analysis and$t$-test analysis considering a general purpose System on Chip executing different cryptographic primitives and an application-specific accelerator implementing the AES-128 algorithm. Our results confirm the impossibility of retrieving the secret key from the power estimates provided by our side-channel-resistant PwrMon. Considering several temporal resolutions, we highlight an accuracy error of the power estimates limited to less than 2.7%, as well as an average area and power overheads for the protected PwrMons lower than 6% and 5%, respectively. To this end, the proposed methodology is able to deliver a side-channel-resistant PwrMon within state-of-the-art accuracy error and overheads. Davide Zoni, Luca Cremona, William Fornaciari |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 10 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 3 |
| 2021 | Managing the Resource Continuum in a Real Video Surveillance ScenarioabstractOver the last years, the number of IoT devices has grown exponentially. Thus, Fog and Edge computing move part of the computation closer to data sources, exploiting interconnected devices in a computing continuum viewpoint. These devices are heterogeneous in terms of performance, features, and capabilities, requiring proper programming models and run-time management layers. This work presents a first version of the BarMan open-source framework. We developed a run-time task allocation policy to maximize application performance and through an experimental evaluation performed on a real cluster, we evaluated different execution scenarios. The results show an improvement up to 66% on the frame processing latency with respect to a monolithic solution. Filippo Sciamanna, Michele Zanella, Giuseppe Massari, William Fornaciari |
DSD | 4 |
| 2021 | A Multi-Level DPM Approach for Real-Time DAG Tasks in Heterogeneous ProcessorsabstractThe modeling and analysis of real-time applications focus on the worst-case scenario because of their strict timing requirements. However, many real-time embedded systems include critical applications requiring not only timing constraints but also other system limitations, such as energy consumption. In this paper, we study the energy-aware real-time scheduling of Directed Acyclic Graph (DAG) tasks. We integrate the Dynamic Power Management (DPM) policy to reduce the Worst-Case Energy Consumption (WCEC), which is an essential requirement for energy-constrained systems. Besides, we extend our analysis with tasks’ probabilistic information to improve the Average- Case Energy Consumption (ACEC), which is, instead, a common non-functional requirement of embedded systems. To verify the benefits of our approach in terms of reduced energy consumption, we finally conduct an extensive simulation, followed by an experimental study on an Odroid-H2 board. Compared to the state-of-the-art solution, our approach is able to reduce the power consumption up to 32.1%. Federico Reghenzani, Ashikahmed Bhuiyan, William Fornaciari, Zhishan Guo |
RTSS | 3 |
| 2021 | Work-in-Progress: Run-Time pWCET Estimation and Quality MonitoringabstractMeasurement-based WCET methods are reliable as long as we know all the application behaviors at design-time. In particular, we need to observe all the possible execution time behaviors before the actual system run-time to get a reliable estimation. This is, unfortunately, difficult to achieve. In this work, we propose an online monitoring framework with the goal to detect, at run-time, unexpected application behaviors and trigger the probabilistic-WCET re-estimation when needed. The online monitoring and estimation framework allows dealing with unexpected situations, such as never-seen applications, faults, or incorrect offline estimations. The proposed approach is tested with simulated and real data. Federico Reghenzani, Filippo Sciamanna, William Fornaciari |
RTSS | 3 |
| 2020 | Dynamic Thermal Management with Proactive Fan Speed Control Through Reinforcement LearningabstractDynamic Thermal Management (DTM) has become a major challenge since it directly affects Multiprocessors Systems-on-chip (MPSoCs) performance, power consumption, and reliability. In this work, we propose a transient fan model, enabling adaptive fan speed control simulation for efficient DTM. Our model is validated through a thermal test chip achieving less than 2°C error in the worst case. With multiple fan speeds, however, the DTM design space grows significantly, which can ultimately make conventional solutions impractical. We address this challenge through a reinforcement learning-based solution to proactively determine the number of active cores, operating frequency, and fan speed. The proposed solution is able to reduce fan power by up to 40% compared to a DTM with constant fan speed with less than 1% performance degradation. Also, compared to a state-of-the-art DTM technique our solution improves the performance by up to 19% for the same fan power. Arman Iranfar, Federico Terraneo, Gabor Csordas, Marina Zapater, William Fornaciari, David Atienza 0001 |
DATE | 5 |
| 2020 | Predictive Resource Management in Energy-constrained Embedded SystemsabstractThe current trends in Internet of Things (IoT) lead to the deployment of low-power devices covering a wide range of application scenarios. These devices have the goal of executing simple tasks, automatically, usually with strict requirements in terms of space and cost. Typically, these devices have to rely on batteries or by harvesting energy devices (e.g., solar panels), in order to operate. On the other hand, IoT devices may be equipped with powerful multi-core CPUs to achieve performance goals, making the management of the energy budget a challenging task. This requires the development of an effective management system, that takes into account current and future energy budget availability, to dynamically bound the actual allocation of processing resources. Specifically, when exploiting solar panels for power supply, we can leverage on the weather forecast, to estimate the availability of energy in the near future. This paper introduces a predictive energy budget management system, targeting multi-core based embedded platforms. Thanks to both local and large-scale weather information, our solution aims at predicting the future incoming power and, accordingly, tuning the exploitable performance level to keep the system running under any environmental condition. Simone Crippa, Giuseppe Massari, Federico Reghenzani, Michele Zanella, William Fornaciari |
DSD | 5 |
| 2020 | A Game Theory Approach to Heterogeneous Resource Management: Work-in-ProgressabstractHeterogeneous computing is a promising solution to scale the performance of computing systems maintaining energy and power efficiency. Managing such resources is, however, complex and it requires smart resource allocation strategies in both embedded and high-performance systems. In this short paper, we propose a game theory approach to allocate heterogeneous resources to applications, with a focus on performance, power, and energy requirements. The game congestion model has been selected and a cost function designed. The proposed allocation strategy is then evaluated by performing a preliminary experimental evaluation. Lara Premi, Federico Reghenzani, Giuseppe Massari, William Fornaciari |
EMSOFT | 4 |
| 2020 | Tiny Neural Networks for Environmental Predictions: An Integrated Approach with MiosixabstractCollecting vast amount of data and performing complex calculations to feed modern Numerical Weather Prediction (NWP) algorithms require to centralize intelligence into some of the most powerful energy and resource hungry supercomputers in the world. This is due to the chaotic complex nature of the atmosphere which interpretation require virtually unlimited computing and storage resources. With Machine Learning (ML) techniques, a statistical approach can be designed in order to perform weather forecasting activity. Moreover, the recently growing interest in Edge Computing Tiny Intelligent architectures is proposing a shift towards the deployment of ML algorithms on Tiny Embedded Systems (ES). This paper describes how Deep but Tiny Neural Networks (DTNN) can be designed to be parsimonious and can be automatically converted into a STM32 microcontroller-optimized C-library through X-CUBE-AI toolchain; we propose the integration of the obtained library with Miosix, a Real Time Operating System (RTOS) tailored for resource constrained and tiny processors, which is an enabling factor for system scalability and multi tasking. With our experiments we demonstrate that it is possible to deploy a DTNN, with a FLASH and RAM occupation of 45,5 KByte and 480 Byte respectively, for atmospheric pressure forecasting in an affordable cost effective system. We deployed the system in a real context, obtaining the same prediction quality as the same DNN model deployed on the cloud but with the advantage of processing all the necessary data to perform the prediction close to environmental sensors, avoiding raw data traffic to the cloud. Francesco Alongi, Nicolò Ghielmetti, Danilo Pau, Federico Terraneo, William Fornaciari |
SMARTCOMP | 5 |
| 2020 | All-Digital Control-Theoretic Scheme to Optimize Energy Budget and Allocation in Multi-CoresabstractThe Internet-of-Things (IoT) revolution fueled new challenges and opportunities to achieve computational efficiency goals. Embedded devices are required to execute multiple applications for which a suitable distribution of the computing power must be adapted at run-time. Such complex hardware platforms have to sustain the continuous acquisition and processing of data under severe energy budget constraints, since most of them are battery powered. The state-of-the-art offers several ad-hoc contributions to selectively optimize the performance considering aspects like energy, power, thermal, or reliability. However, there is a need for a generic coordinated management strategy able to cope with all of these dimensions, while allowing the Operating System (OS) and the applications to “suggest” or constrain the actuation. This article proposes a unified control-theoretic scheme to coordinate the design of energy-budget and energy allocation solutions for multi-cores. The proposed controller can work with any actuator and it can interact, at run-time, with both the applications and the OS to optimize the actuation signals steering the computing platform. Such control scheme offers the possibility to integrate any performance related policy in the form of an energy-allocation strategy, still ensuring the theoretic exponential stability of the overall controller if the actuation of the policy, coming from the OS and the applications, “is not too fast.” To demonstrate the feasibility of our solution, we have implemented the controller into a RISC multi-core running on the Xilinx Artix 100t FPGA device, available in the the Digilent Nexys4-DDR board. Results considering two actuators and both the quadand the eight-core version of the considered computing platform, highlight the scalability of the proposed solution as well as an area overhead for the -all digital, on chip-controller limited to 0.86 percent (FFs) and 5.3 percent (LUTs) of the FPGA chip. We also considered a dynamic scenario validating the speed of the controller, where our framework has to face with modifications to the energy-allocation control policy carried out by the OS and the applications. The obtained results are collected by executing a huge mix of benchmarks and the statistical significance is accounted by executing each scenario 30 times. Such results are analyzed considering three quality metrics. First, the efficiency in exploiting the imposed budget (EFF9) that is on average 98.27 percent. Second, the overflow of the actual average power consumption with respect to the assigned budget (OνF9), which is limited to 1.43 mW on average. Last, the performance utility loss due to the control scheme that is limited to 1.87 percent on average. Davide Zoni, Luca Cremona, William Fornaciari |
IEEE Trans. Computers | 3 |
| 2020 | Scramble Suit: A Profile Differentiation Countermeasure to Prevent Template AttacksabstractEnsuring protection against side channel attacks (SCAs) is a crucial requirement in the design of modern secure embedded systems. Profiled SCAs, the class to which template attacks and machine learning attacks belong, derive a model of the side channel behavior of a device identical to the target one, and exploit the said model to extract the key from the target, under the hypothesis that the side channel behaviors of the two devices match. We propose an architectural countermeasure against cross-device profiled attacks which differentiates the side channel behavior of different instances of the same hardware design, preventing the reuse of a model derived on a device other than the target one. In particular, we describe an instance of our solution providing a protected hardware implementation of the advanced encryption standard (AES) block cipher and experimentally validate its resistance against both Bayesian templates and machine learning approaches based on support vector machines also considering different state-of-the-art feature reduction techniques to increase the effectiveness of the profiled attacks. Results show that our countermeasure foils the key retrieval attempts via profiled attacks ensuring a key derivation accuracy equivalent to a random guess. Alessandro Barenghi, William Fornaciari, Gerardo Pelosi, Davide Zoni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Optimizing Energy in Non-Preemptive Mixed-Criticality Scheduling by Exploiting Probabilistic InformationabstractThe strict requirements on the timing correctness biased the modeling and analysis of real-time systems toward the worst-case performances. Such focus on the worst-case, however, does not provide enough information to effectively steer the resource/energy optimization. In this article, we integrate a probabilistic-based energy prediction strategy with the precise scheduling of mixed-criticality tasks, where the timing correctness must be met for all tasks at all scenarios. The dynamic voltage and frequency scaling (DVFS) is applied to this precise scheduling policy to enable energy minimization. We propose a probabilistic technique to derive an energy-efficient speed (for the processor) that minimizes the average energy consumption, while guaranteeing the (worst-case) timing correctness for all tasks, including LO-criticality ones, under any execution condition. We present a response time analysis for such systems under the nonpreemptive fixed-priority scheduling policy. Finally, we conduct an extensive simulation campaign based on randomly generated task sets to verify the effectiveness of our algorithm (with respect to energy savings) and it reports up to 46% energy-saving. Ashikahmed Bhuiyan, Federico Reghenzani, William Fornaciari, Zhishan Guo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Dealing with Uncertainty in pWCET EstimationsabstractThe problem of estimating a tight and safe Worst-Case Execution Time (WCET), needed for certification in safety-critical environment, is a challenging problem for modern embedded systems. A possible solution proposed in past years is to exploit statistical tools to obtain a probability distribution of the WCET. These probabilistic real-time analyses for WCET are, however, subject to errors, even when all the applicability hypotheses are satisfied and verified. This is caused by the uncertainties of the probabilistic-WCET distribution estimator. This article aims at improving the measurement-based probabilistic timing analysis approach providing some techniques to analyze and deal with such uncertainties. The so-called region of acceptance model based on state-of-the-art statistical test procedures is defined over the distribution space parameters. From this model, a set of strategies is derived and discussed to provide the methodology to deal with the trade-off safety/tightness of the WCET estimation. These techniques are then tested over real datasets, including industrial safety-critical applications, to show the increased value of using the proposed approach in probabilistic WCET analyses. Federico Reghenzani, Luca Santinelli, William Fornaciari |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Challenges in Deeply Heterogeneous High Performance SystemsabstractRECIPE (REliable power and time-ConstraInts-aware Predictive management of heterogeneous Exascale systems) is a recently started project funded within the H2020 FETHPC programme, which is expressly targeted at exploring new High-Performance Computing (HPC) technologies. RECIPE aims at introducing a hierarchical runtime resource management infrastructure to optimize energy efficiency and minimize the occurrence of thermal hotspots, while enforcing the time constraints imposed by the applications and ensuring reliability for both time-critical and throughput-oriented computation that run on deeply heterogeneous accelerator-based systems. This paper presents a detailed overview of RECIPE, identifying the fundamental challenges as well as the key innovations addressed by the project, which span run-time management, heterogeneous computing architectures, HPC memory/interconnection infrastructures, thermal modelling, reliability, programming models, and timing analysis. For each of these areas, the paper describes the relevant state of the art as well as the specific actions that the project will take to effectively address the identified technological challenges. Giovanni Agosta, William Fornaciari, David Atienza 0001, Ramon Canal, Alessandro Cilardo, José Flich, Carles Hernández 0001, Michal Kulczewski, Giuseppe Massari, Rafael Tornero, Marina Zapater |
DSD | 2 |
| 2019 | A Probabilistic Approach to Energy-Constrained Mixed-Criticality SystemsabstractIn battery-powered embedded systems, the energy budget management is a critical aspect. For systems using unreliable power sources, e.g. solar panels, the continuous system operation is a challenging requirement. In such scenarios, effective management policies must rely on accurate energy estimations. In this paper we propose a measurement-based probabilistic approach to address the worst-case energy consumption (WCEC) estimation, coupled with a job admission algorithm for energy-constrained task scheduling. The overall goal is to demonstrate how the proposed approach can introduce benefits also in mission-critical systems, where unsafe energy budget estimations cannot be tolerated. Federico Reghenzani, Giuseppe Massari, William Fornaciari |
ISLPED | 3 |
| 2019 | Partial Packet Forwarding to Improve Performance in Fully Adaptive Routing for Cache-Coherent NoCsabstractIn the contest of cache-coherent Networks-on-Chip (NoCs), fully adaptive routing algorithms guarantee maximum flexibility to implement power-performance, fault tolerant, thermal and Quality of Service (QoS) management policies. However, to get rid of deadlock at both protocol and network level, their implementation imposes a relevant resource increase. Moreover, their performance are inferior to the one of deterministic and partially adaptive schemes mainly due to the additional constraints imposed to the virtual channel (VC) re-use policy. This work proposes a novel flow control scheme to improve the performance of fully adaptive routing algorithms by allowing an aggressive reuse of VCs in presence of both long and short packets. Our proposal works by splitting long packets in multiple chunks and by reallocating the VCs to the chunks rather that to the entire packet. By carefully sizing each chunk to fit the available space in the reallocated, eventually not empty, VC, we are avoiding deadlocks while increasing the NoC utilization and performance. Experimental results show that our solution offers a 23.8% increase, on average, in the saturation point when compared to the best state of the art flow control scheme for fully adaptive routing algorithms. Moreover, our flow control scheme offers similar or better performance than the XY routing algorithm with the same number of resources, and we also ensure superior flexibility in the definition of the routing function. Tamer Eltaras, William Fornaciari, Davide Zoni |
PDP | 2 |
| 2018 | A constrained extremum-seeking control for CPU thermal managementabstractThe increasing complexity of computing architectures is pushing for novel Dynamic Thermal Management (DTM) techniques. Accordingly, more accurate power and thermal models are required. In this work, we propose a thermal controller based on a constrained extremum-seeking algorithm, enabling resource allocation optimization under specific thermal constraints. This approach comes with many advantages. First, the controller does not require any model of the system, dropping the need for a complex and potentially imprecise estimation phase. Second, it allows the control of derived measurements. We show how this may positively impact on the CPU reliability. Federico Reghenzani, Simone Formentin, Giuseppe Massari, William Fornaciari |
CF | 4 |
| 2018 | PowerProbe: Run-time power modeling through automatic RTL instrumentationabstractOnline power monitoring represents a de-facto solution to enable energyand power-aware run-time optimizations for current and future computing architectures. Traditionally, the performance counters of the target architecture are used to feed a software-based, power model that is continuously updated to deliver the required run-time power estimates. The solution introduces a non-negligible performance and energy overhead. Moreover, itis limited to the availability of such performance counters that, however, are not primarily intended for online power monitoring. This paper introduces PowerProbe, a run-time power monitoring methodology that automatically extracts and implements a power model from the RTL description of the target architecture. The solution does not leverage any performance counter to ensure wide applicability. Moreover, the use of ad-hoc hardware that continuously updates the power estimate minimizes both the performance and the power overheads. We employ a fully compliant OpenRisc 1000 implementation to validate PowerProbe. The results highlight an average prediction error within 9% (standard deviation less than 2%), with a power and area overheads limited to 6.89% and 4.71%, respectively. Davide Zoni, Luca Cremona, William Fornaciari |
DATE | 3 |
| 2018 | TDMH-MAC: Real-Time and Multi-hop in the Same Wireless MACabstractSupporting real-time communications over Wireless networks (WSNs) is a tough challenge, due to packet collisions and the non-determinism of common channel access schemes like CSMA/CA. Real-time WSN communication is even more problematic in the general case of multi-hop mesh networks. For this reason, many real-time WSN solutions are limited to simple topologies, such as star networks. We propose a real-time multi-hop WSN MAC protocol built atop the IEEE 802.15.4 physical layer. By relying on precise clock synchronization and constructive interference-based flooding, the proposed MAC builds a centralized TDMA schedule, supporting multi-hop mesh networks. The real-time multi-hop communication model is connection-oriented, using guaranteed time slots, ad enables point-to-point communications also with redundant paths. The protocol has been implemented in simulation using OMNeT++, and the performance has been verified in a real-world deployment using Wandstem WSN nodes. Federico Terraneo, Paolo Polidori, Alberto Leva, William Fornaciari |
RTSS | 4 |
| 2018 | DarkCache: Energy-Performance Optimization of Tiled Multi-Cores by Adaptively Power-Gating LLC BanksabstractThe Last Level Cache (LLC) is a key element to improve application performance in multi-cores. To handle the worst case, the main design trend employs tiled architectures with a large LLC organized in banks, which goes underutilized in several realistic scenarios. Our proposal, named DarkCache , aims at properly powering off such unused banks to optimize the Energy-Delay Product (EDP) through an adaptive cache reconfiguration, thus aggressively reducing the leakage energy. The implemented solution is general and it can recognize and skip the activation of the DarkCache policy for the few strong memory intensive applications that actually require the use of the entire LLC. The validation has been carried out on 16- and 64-core architectures also accounting for two state-of-the-art methodologies. Compared to the baseline solution, DarkCache exhibits a performance overhead within 2% and an average EDP improvement of 32.58% and 36.41% considering 16 and 64 cores, respectively. Moreover, DarkCache shows an average EDP gain between 16.15% (16 cores) and 21.05% (64 cores) compared to the best state-of-the-art we evaluated, and it confirms a good scalability since the gain improves with the size of the architecture. Davide Zoni, William Fornaciari |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | A Comprehensive Side-Channel Information Leakage Analysis of an In-Order RISC CPU MicroarchitectureabstractSide-channel attacks are a prominent threat to the security of embedded systems. To perform them, an adversary evaluates the goodness of fit of a set of key-dependent power consumption models to a collection of side-channel measurements taken from an actual device, identifying the secret key value as the one yielding the best-fitting model. In this work, we analyze for the first time the microarchitectural components of a 32-bit in-order RISC CPU, showing which one of them is accountable for unexpected side-channel information leakage. We classify the leakage sources, identifying the data serialization points in the microarchitecture and providing a set of hints that can be fruitfully exploited to generate implementations resistant against side-channel attacks, either writing or generating proper assembly code. Davide Zoni, Alessandro Barenghi, Gerardo Pelosi, William Fornaciari |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2017 | HARPA: Tackling physically induced performance variabilityabstractContinuously increasing application demands on both High Performance Computing (HPC) and Embedded Systems (ES) are driving the IC manufacturing industry on an everlasting scaling of devices in silicon. Nevertheless, integration and miniaturization of transistors comes with an important and non-negligible trade-off: time-zero and time-dependent performance variability. Increasing guard-bands to battle variability is not scalable, since worst-case design margins are prohibitive for downscaled technology nodes. This paper discusses the FP7-612069-HARPA project of the European Commission which aims to enable next-generation embedded and high-performance heterogeneous many-cores to cost-effectively confront variations by providing Dependable-Performance: correct functionality and timing guarantees throughout the expected lifetime of a platform under thermal, power, and energy constraints. The HARPA novelty is in seeking synergies in techniques that have been considered virtually exclusively in the ES or HPC domains (worst-case guaranteed partly proactive techniques in embedded, and dynamic best-effort reactive techniques in high-performance). Nikolaos Zompakis, Michail Noltsis, Lorena Ndreu, Zacharias Hadjilambrou, Panayiotis Englezakis, Panagiota Nikolaou, Antoni Portero, Simone Libutti, Giuseppe Massari, Federico Sassi, Alessandro Bacchini, Chrysostomos Nicopoulos, Yiannakis Sazeides, Radim Vavrík, Martin Golasowski, Jiri Sevcík, Vít Vondrák, Francky Catthoor, William Fornaciari, Dimitrios Soudris |
DATE | 19 |
| 2017 | MANGO: Exploring Manycore Architectures for Next-GeneratiOn HPC SystemsabstractThe Horizon 2020 MANGO project aims at exploring deeply heterogeneous accelerators for use in High-Performance Computing systems running multiple applications with different Quality of Service (QoS) levels. The main goal of the project is to exploit customization to adapt computing resources to reach the desired QoS. For this purpose, it explores different but interrelated mechanisms across the architecture and system software. In particular, in this paper we focus on the runtime resource management, the thermal management, and support provided for parallel programming, as well as introducing three applications on which the project foreground will be validated. José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Etienne Cappe, Alessandro Cilardo, Leon Dragic, Alexandre Dray, Alen Duspara, William Fornaciari, Gerald Guillaume, Ynse Hoornenborg, Arman Iranfar, Mario Kovac, Simone Libutti, Bruno Maitre, José Maria Martínez, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Tomás Picornell, Igor Piljic, Anna Pupykina, Federico Reghenzani, Isabelle Staub, Rafael Tornero, Marina Zapater, Davide Zoni |
DSD | 11 |
| 2017 | Mixed Time-Criticality Process Interferences Characterization on a Multicore Linux SystemabstractThe increasing interest in the integration of Mixed Criticality Systems (MCS) in Commercial-Off-The-Shelf (COTS) platforms leads to an increasing number of challenges. The possibility of sharing computing resources among applications with different time criticalities is a key goal for COTS systems, but still hard to achieve. Classical approaches in real-time systems are not feasible when platform and operating system may introduce unpredictability in the task execution. Moreover, if the system must also meet non-functional requirements (e.g., thermal and power management), dynamic approaches of computing resources allocation are more effective than static ones. Unfortunately, this contributes to increasing the complexity of the scenario. In MCS, the overheads and the unpredictability caused by sharing resources like cache memories have been well studied. However, in some cases we could also consider the operating system itself as a potential source of unexpected and unpredictable latencies, if several running tasks perform system calls. This work aims at proposing a model for the intra-core and inter-core interferences and the analysis of the OS-induced latencies in a Linux real-time system, both essential for the creation of smart and effective run-time resource management policies. Federico Reghenzani, Giuseppe Massari, William Fornaciari |
DSD | 3 |
| 2017 | BlackOut: Enabling fine-grained power gating of buffers in Network-on-Chip routers
Davide Zoni, Andrea Canidio, William Fornaciari, Panayiotis Englezakis, Chrysostomos Nicopoulos, Yiannakis Sazeides |
J. Parallel Distributed Comput. | 3 |
| 2016 | Enabling HPC for QoS-sensitive applications: The MANGO approach
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Alessandro Cilardo, William Fornaciari, Ynse Hoornenborg, Mario Kovac, Bruno Maitre, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Fabrice Roudet, Rafael Tornero, Davide Zoni |
DATE | 7 |
| 2016 | V2I Cooperation for Traffic Management with SafeCopabstractThe Safe Cooperating Cyber-Physical Systems using Wireless Communication (SafeCop) project addresses safety-related issues in cooperating cyber-physical systems. These systems, characterised by wireless communications, multiple stakeholders, and variable operating environments, are called Cooperative Open Cyber-Physical Systems (CO-CPS). CO-CPSs can successfully address several societal challenges -- cooperative vehicles have been shown to reduce fuel consumption as well as the number of accidents. A vehicle-to-infrastructure (V2I) cooperation for the traffic management scenario is therefore considered as a key use cases of SafeCop. In this paper, we outline the V2I traffic management scenario, assess the research goals that arise from it, and provide an overview of the architecture of the demonstrator, as well as a roadmap for its development and evaluation. Giovanni Agosta, Alessandro Barenghi, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Stefano Delucchi, Massimo Massa, Maurizio Mongelli, Enrico Ferrari, Leonardo Napoletani, Luciano Bozzi, Carlo Tieri, Dajana Cassioli, Luigi Pomante |
DSD | 4 |
| 2016 | The M2DC Project: Modular Microserver DataCentreabstractThe Modular Microserver DataCentre (M2DC) project will investigate, develop and demonstrate a modular, highly-efficient, cost-optimized server architecture composed of heterogeneous microserver computing resources, being able to be tailored to meet requirements from various application domains such as image processing, cloud computing or HPC. M2DC will be built on three main pillars: a flexible server architecture that can be easily customised, maintained and updated, advanced management strategies and system efficiency enhancements (SEE), well-defined interfaces to surrounding software data centre ecosystem. Mariano Cecowski, Giovanni Agosta, Ariel Oleksiak, Michal Kierzynka, Micha vor dem Berge, Wolfgang Christmann, Stefan Krupop, Mario Porrmann, Jens Hagemeyer, René Griessl, Meysam Peykanu, Lennart Tigges, Sven Rosinger, Daniel Schlitt, Christian Pieper, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Robert Plestenjak, Justin Cinkelj, Loïc Cudennec, Thierry Goubier, Jean-Marc Philippe, Udo Janssen, Chris Adeniyi-Jones |
DSD | 17 |
| 2016 | CONTREX: Design of Embedded Mixed-Criticality CONTRol Systems under Consideration of EXtra-Functional PropertiesabstractThe increasing processing power of today's HW/SW platforms leads to the integration of more and more functions in a single device. Additional design challenges arise when these functions share computing resources and belong to different criticality levels. The paper presents the CONTREX European project and its preliminary results. CONTREX complements current activities in the area of predictable computing platforms and segregation mechanisms with techniques to consider the extra-functional properties, i.e., timing constraints, power, and temperature. CONTREX enables energy efficient and cost aware design through analysis and optimization of these properties with regard to application demands at different criticality levels. Ralph Görgen, Kim Grüttner, Fernando Herrera, Pablo Peñil, Julio L. Medina, Eugenio Villar, Gianluca Palermo, William Fornaciari, Carlo Brandolese, Davide Gadioli, Sara Bocchio, Luca Ceva, Paolo Azzoni, Massimo Poncino, Sara Vinco, Enrico Macii, Salvatore Cusenza, John M. Favaro, Raúl Valencia, Ingo Sander, Kathrin Rosvall, Davide Quaglia |
DSD | 8 |
| 2016 | Demo: A High-Performance, Energy-Efficient Node for a Wide Range of WSN Applications
Federico Terraneo, Alberto Leva, William Fornaciari |
EWSN | 3 |
| 2016 | The MIG Framework: Enabling Transparent Process Migration in Open MPIabstractThis paper introduces the mig framework: an Open MPI extension to transparently support the migration of application processes, over different nodes of a distributed High-Performance Computing (HPC) system. The framework provides mechanism on top of which suitable resource managers can implement policies to react to hardware faults, address performance variability, improve resource utilization, perform a fine-grained load balancing and power thermal management. Federico Reghenzani, Gianmario Pozzi, Giuseppe Massari, Simone Libutti, William Fornaciari |
EuroMPI | 5 |
| 2016 | CUTBUF: Buffer Management and Router Design for Traffic Mixing in VNET-Based NoCsabstractRouter's buffer design and management strongly influence energy, area and performance of on-chip networks, hence it is crucial to encompass all of these aspects in the design process. At the same time, the NoC design cannot disregard preventing network-level and protocol-level deadlocks by devoting ad-hoc buffer resources to that purpose. In chip multiprocessor systems the coherence protocol usually requires different virtual networks (VNETs) to avoid deadlocks. Moreover, VNET utilization is highly unbalanced and there is no way to share buffers between them due to the need to isolate different traffic types. This paper proposes CUTBUF, a novel NoC router architecture to dynamically assign virtual channels (VCs) to VNETs depending on the actual VNETs load to significantly reduce the number of physical buffers in routers, thus saving area and power without decreasing NoC performance. Moreover, CUTBUF allows to reuse the same buffer for different traffic types while ensuring that the optimized NoC is deadlock-free both at network and protocol level. In this perspective, all the VCs are considered spare queues not statically assigned to a specific VNET and the coherence protocol only imposes a minimum number of queues to be implemented. Synthetic applications as well as real benchmarks have been used to validate CUTBUF, considering architectures ranging from 16 up to 48 cores. Moreover, a complete RTL router has been designed to explore area and power overheads. Results highlight how CUTBUF can reduce router buffers up to 33 percent with 2 percent of performance degradation, a 5 percent of operating frequency decrease and area and power saving up to 30.6 and 30.7 percent, respectively. Conversely, the flexibility of the proposed architecture improves by 23.8 percent the performance of the baseline NoC router when the same number of buffers is used. Davide Zoni, José Flich, William Fornaciari |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Playful Supervised Smart Spaces (P3S) - A Framework for Designing, Implementing and Deploying Multisensory Play Experiences for Children with Special NeedsabstractOur research explores novel forms of smart spaces that support full body interaction with smart objects instrumented with audio, light, and motion sensors and actuators, virtual worlds on medium-large displays, and smart lights, and are integrated with cloud services for remote supervision and analysis of user behaviour. In this paper, we present the architecture and coordination infrastructure created in the Playful Supervised Smart Space (P3S) project, as well as initial designs for the Smart Object and Smart Space Gateway Components. Giovanni Agosta, Luca Borghese, Carlo Brandolese, Francesco Clasadonte, William Fornaciari, Franca Garzotto, Mirko Gelsomini, Matteo Grotto, Cristina Frà, Danny Noferi, Massimo Valla |
DSD | 5 |
| 2015 | Harnessing Performance Variability: A HPC-Oriented Application ScenarioabstractThe technology scaling towards the 10nm of the silicon manufacturing, is going to introduce variability challenges, mainly due to the growing susceptibility to thermal hot-spots and time-dependent variations (aging) in the silicon chip. The consequences are two-fold: a) unpredictable performance, b) unreliable computing resources. The goal of the HARPA project is to enable next-generation embedded and high-performance heterogeneous many-core processors to effectively address this issues, through a cross-layer approach, involving several component of the system stack. Each component acts at different levels and time granularity. This paper focus on one of the components of the HARPA stack, the HARPA-OS, showing early results of a first integration step of the HARPA approach in a real High-Performance Computing (HPC) application scenario. Giuseppe Massari, Simone Libutti, Antoni Portero, Radim Vavrík, Stepán Kuchár, Vít Vondrák, Luca Borghese, William Fornaciari |
DSD | 8 |
| 2015 | TEST: Assessing NoC Policies Facing Aging and Leakage PowerabstractThe trend to increase the number of cores integrated on a single die makes Networks-on-Chip (NoCs) a key component from the interconnection viewpoint. Unfortunately, continuous scaling of CMOS technology poses severe concerns regarding failure mechanisms, such as NBTI, that are crucial in achieving a reasonable component lifetime. Furthermore, the leakage power became more and more a critical issues as the technology scales up. Finally, Process Variation (PV) makes harder the scenario, decreasing device lifetime and performance predictability during chip fabrication. Several techniques have been presented in literature facing the NBTI and or the static power consumption. This paper proposes a methodology to analyze such techniques from the feasibility viewpoint. It is explored their effectiveness in contrasting NBTI and saving static power in the NoC as well as the associated overheads and drawbacks. For the two considered policies, it is achieved a NBTI mitigation up to 55% and a power saving up to 51% with performance and area overheads less than 10% and 5%, respectively. Davide Zoni, Luca Borghese, Giuseppe Massari, Simone Libutti, William Fornaciari |
DSD | 5 |
| 2015 | Flood Prediction Model Simulation With Heterogeneous Trade-Offs In High Performance Computing FrameworkabstractIn this paper, we propose a safety-critical system with a run-time resource management that is used to operate an application for flood monitoring and prediction. This application can run with different Quality of Service (QoS) levels depending on the current hydrometeorological situation. The system operation can follow two main scenarios - standard or emergency operation. The standard operation is active when no disaster occurs, but the system still executes short-term prediction simulations and monitors the state of the river discharge and precipitation intensity. Emergency operation is active when some emergency situation is detected or predicted by the simulations. The resource allocation can either be used for decreasing power consumption and minimizing needed resources in standard operation, or for increasing the precision and decreasing response times in emergency operation. This paper shows that it is possible to describe different optimal points at design time and use them to adapt to the current quality of service requirements during run-time. Antoni Portero, Radim Vavrík, Stepán Kuchár, Martin Golasowski, Vít Vondrák, Simone Libutti, Giuseppe Massari, William Fornaciari |
ECMS | 8 |
| 2015 | Modeling DVFS and Power-Gating Actuators for Cycle-Accurate NoC-Based SimulatorsabstractNetworks-on-chip (NoCs) are a widely recognized viable interconnection paradigm to support the multi-core revolution. One of the major design issues of multicore architectures is still the power, which can no longer be considered mainly due to the cores, since the NoC contribution to the overall energy budget is relevant. To face both static and dynamic power while balancing NoC performance, different actuators have been exploited in literature, mainly dynamic voltage frequency scaling (DVFS) and power gating. Typically, simulation-based tools are employed to explore the huge design space by adopting simplified models of the components. As a consequence, the majority of state-of-the-art on NoC power-performance optimization do not accurately consider timing and power overheads of actuators, or (even worse) do not consider them at all, with the risk of overestimating the benefits of the proposed methodologies. This article presents a simulation framework for power-performance analysis of multicore architectures with specific focus on the NoC. It integrates accurate power gating and DVFS models encompassing also their timing and power overheads. The value added of our proposal is manyfold: (i) DVFS and power gating actuators are modeled starting from SPICE-level simulations; (ii) such models have been integrated in the simulation environment; (iii) policy analysis support is plugged into the framework to enable assessment of different policies; (iv) a flexible GALS ( globally asynchronous locally synchronous ) support is provided, covering both handshake and FIFO re-synchronization schemas. To demonstrate both the flexibility and extensibility of our proposal, two simple policies exploiting the modeled actuators are discussed in the article. Davide Zoni, William Fornaciari |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2015 | A control-based methodology for power-performance optimization in NoCs exploiting DVFS
Davide Zoni, Federico Terraneo, William Fornaciari |
J. Syst. Archit. | 3 |
| 2015 | Effective Runtime Resource Management Using Linux Control Groups with the BarbequeRTRM FrameworkabstractThe extremely high technology process reached by silicon manufacturing (smaller than 32nm) has led to production of computational platforms and SoC, featuring a considerable amount of resources. Whereas from one side such multi- and many-core platforms show growing performance capabilities, from the other side they are more and more affected by power, thermal, and reliability issues. Moreover, the increased computational capabilities allows congested usage scenarios with workloads subject to mixed and time-varying requirements. Effective usage of the resources should take into account both the application requirements and resources availability , with an arbiter, namely a resource manager in charge to solve the resource contention among demanding applications. Current operating systems (OS) have only a limited knowledge about application-specific behaviors and their time-varying requirements. Dedicated system interfaces to collect such inputs and forward them to the OS (e.g., its scheduler) are thus an interesting research area that aims at integrating the OS with an ad hoc resource manager. Such a component can exploit efficient low-level OS interfaces and mechanisms to extend its capabilities of controlling tasks and system resources. Because of the specific tasks and timings of a resource manager, this component can be easily and effectively developed as a user-space extension lying in between the OS and the controlled application. This article, which focuses on multicore Linux systems, shows a portable solution to enforce runtime resource management decisions based on the standard control groups framework. A burst and a mixed workload analysis, performed on a multicore-based NUMA platform, have reported some promising results both in terms of performance and power saving. Patrick Bellasi, Giuseppe Massari, William Fornaciari |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | OpenCL Application Auto-tuning and Run-Time Resource Management for Multi-core PlatformsabstractTo support adaptivity of data parallel applications on multi-core platforms, we propose a framework based on the combination of OpenCL application auto-tuning and run-time resource management. The framework addresses computationally intensive multimedia OpenCL applications. For these target applications, we show that application auto-tuning, based on design-time analysis, can become synergistic with run-time resource management. In the proposed framework, run-time decisions are taken by each application, autonomously, to achieve system adaptivity. This paper describes the methodology and related toolchain, defined during the 2PARMA European project, based on the integration of independent tools to provide effective compilation of OpenCL code, multi-objective design space exploration, application monitoring and tuning and system-wide run-time resource management. Experimental results are reported for design optimization of an OpenCL stereo-matching application and then for a resource contention scenario where multiple stereo-matching applications are executed on the same platform with different run-time requirements. Davide Gadioli, Simone Libutti, Giuseppe Massari, Edoardo Paone, Michele Scandale, Patrick Bellasi, Gianluca Palermo, Vittorio Zaccaria, Giovanni Agosta, William Fornaciari, Cristina Silvano |
ISPA | 10 |
| 2014 | An optimal model to partition the evolution of periodic tasks in wireless sensor networksabstractOur research targets an operating paradigm that can be referred to as WSN cloud infrastructure, in which a WSN provider maintains the sensor network over a certain area and allows potential clients to deploy their applications to monitor some parameters of interest. Applications coming from multiple heterogeneous sources may have very inhomogeneous periods and tolerable slacks, inducing a fragmented and inefficient duty cycle on the nodes, strongly accelerating the energy depletion process. In such a paradigm, a way to merge as much as possible the execution moments of the periodic tasks is mandatory, in order to reduce the energy overheads of frequent awakenings of system and devices. The partitioning method presented in this paper subdivides the evolution of the periodic tasks in small self-contained optimization sub-problems, which can be subsequently processed at run-time through a task-merging or a scheduling policy. The model and the related algorithm have been conceived such that the formal guarantees of global optimality are satisfied in the general case of a coverage problem, aimed at finding the sequence of system awakenings that fulfill the deadlines of each periodic task in the system. Moreover, the generality of the model makes it suitable for a variety of task-merging policies and algorithms. Experimental results demonstrate the high computational efficiency of the proposed partitioning approach. Carlo Brandolese, Luigi Rucco, William Fornaciari |
WoWMoM | 3 |
| 2013 | Sensor-wise methodology to face NBTI stress of NoC buffersabstractNetworks-on-Chip (NoCs) are a key component for the new many-core architectures, from the performance and reliability stand-points. Unfortunately, continuous scaling of CMOS technology poses severe concerns regarding failure mechanisms such as NBTI and stress-migration. Process variation makes harder the scenario, decreasing device lifetime and performance predictability during chip fabrication. This paper presents a novel cooperative sensor-wise methodology to reduce the NBTI degradation in the network on-chip (NoC) virtual channel (VC) buffers, considering process variation effects as well. The changes introduced to the reference NoC model exhibit an area overhead below 4%. Experimental validation is obtained using a cycle accurate simulator considering both real and synthetic traffic patterns. We compare our methodology to the best sensor-less round-robin approach used as reference model. The proposed sensor-wise strategy achieves up to 26.6% and 18.9% activity factor improvement over the reference policy on synthetic and real traffic patterns respectively. Moreover a net NBTI Vthsaving up to 54.2% is shown against the baseline NoC that does not account for NBTI. Davide Zoni, William Fornaciari |
DATE | 2 |
| 2013 | A Formal Model for Optimal Autonomous Task Hibernation in Constrained Embedded SystemsabstractThis paper proposes and studies an autonomous hibernation technique and optimal hibernation policies aimed at minimizing the power consumption, while allowing stateful processing in constrained embedded systems with long-lasting lifetime requirements. To this purpose the paper models the energy contributions for hibernating the system-by saving the memory status on an external non-volatile memory and completely powering off the system-rather than maintaining the system in a sleep mode with memory retention-with problems of static leakage power-between two consecutive bursts of processing. Thanks to a simplified yet formal notion of system state, the paper rigorously determines the optimal conditions for deciding whether to hibernate or not the system during idle periods. Hibernation policies have been implemented as a module of the operating system and results demonstrate energy savings up to 50% compared to trivial hibernation approaches. Moreover, the hibernation policy proved to be robust and stable with respect to changes of the application parameters. Carlo Brandolese, William Fornaciari, Luigi Rucco |
DSD | 2 |
| 2013 | Performance/reliability trade-off in superscalar processors for aggressive NBTI restoration of functional unitsabstractNegative-Bias Temperature Instability has become a serious reliability concern in modern processors design, and in the last decade many research effort has been spent in developing circuit-level and architecture-level strategies to mitigate the induced delay variation of nanoscale circuits. At the architecture level, work has been proposed to alleviate this by appropriate dynamic instruction schedulingntechniques. However, their benefit is bounded to the available redundancy, limiting their attractiveness for a cost-effective VLSI solution. This paper presents an in-depth analysis of a performance reliability trade-off FSM design, that is able to attain the desired level of reliability improvement according to the ILP performance constraints. Power-gating is used for aggressive NBTI restoration. Extensive experimental results show several performance reliability trade-off examples in a broad range of application scenarios. Simone Corbetta, William Fornaciari |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Optimal hibernation policies for energy efficient stateful operation in high-end wireless sensor nodesabstractThis paper proposes and studies an hibernation technique and optimal hibernation policies aimed at minimizing the power consumption, while allowing stateful processing and the adoption of more powerful nodes. To this purpose the paper models the energy trade-off for hibernating the system rather than putting it in a memory-retention sleep mode between two consecutive bursts of processing. Thanks to a simplified notion of system state, the paper formally determines the optimal conditions for deciding whether to hibernate or not the system during idle periods. Hibernation policies have been implemented as a module of the operating system and results demonstrate energy savings up to 50% compared to trivial hibernation approaches. Moreover, the hibernation policy proved to be robust and stable with respect to changes of the application parameters. Carlo Brandolese, William Fornaciari, Luigi Rucco |
WOWMOM | 2 |
| 2012 | COMPLEX: COdesign and Power Management in PLatform-Based Design Space EXplorationabstractThe consideration of an embedded device's power consumption and its management is increasingly important nowadays. Currently, it is not easily possible to integrate power information already during the platform exploration phase. In this paper, we discuss the design challenges of today's heterogeneous HW/SW systems regarding power and complexity, both for platform vendors as well as system integrators. As a result, we propose a design flow concept that combines system-level power optimization techniques with platform-based rapid prototyping. Virtual executable prototypes are generated from MARTE/UML and functional C/C++ descriptions, which then allows to study different platforms, mapping alternatives and power management strategies. Our proposed flow combines system-level timing and power estimation techniques available in commercial tools with platform-based rapid prototyping. We propose an efficient code annotation technique for timing and power properties that enables fast host execution as well as adaptive collection of power traces. Combined with a flexible design-space exploration (DSE) approach our flow allows a trade-off between different platforms, mapping alternatives, and optimization techniques, based on domain-specific workload scenarios. The proposed flow is currently under implementation in the COMPLEX FP7 European integrated project. Kim Grüttner, Philipp A. Hartmann, Kai Hylla, Sven Rosinger, Wolfgang Nebel, Fernando Herrera, Eugenio Villar, Carlo Brandolese, William Fornaciari, Gianluca Palermo, Chantal Ykman-Couvreur, Davide Quaglia, Francisco Ferrero 0002, Raúl Valencia |
DSD | 9 |
| 2012 | NBTI mitigation in microprocessor designsabstractNegative-Bias Temperature Instability seriously affects nanoscale circuits reliability and performance. Continuous stress and increasing operating temperatures lead to device degradation and long-term system unavailability. The opportunity to optimize the duty-cycle of the stress/recovery phases to reduce Vth degradation leads to innovative research of reliability-oriented resources allocation at architectural level. This work explores the impact of different allocation strategies on the processor degradation, through a novel estimation methodology. Experimental results show that the proposed NBTI-aware allocation strategy can guarantee from 10% and up to 30% lower degradation compared to classical strategies, under different operating scenarios and under process variability. Simone Corbetta, William Fornaciari |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | HANDS: heterogeneous architectures and networks-on-chip design and simulationabstractIn current multi-core scenario, Networks-on-Chip (NoC) represent a suitable choice to face the increasing communication and performance requirements, however introducing additional design challenges to already complex architectures. In this perspective, there is a need for flexible and configurable virtual platforms for early-stage design exploration. We present the Heterogeneous Architectures and Networks-on-Chip Design and Simulation framework for large-scale high-performance computer simulation, integrating performance, power, thermal and reliability metrics under a unique methodology. Moreover, NoC exploration is possible from a reliability/performance and thermal/performance trade-offs. Davide Zoni, Simone Corbetta, William Fornaciari |
ISLPED | 3 |
| 2011 | Estimation of thermal status in multi-core systemsabstractModern multi-core architectures are prone to a complex dynamic thermal management process: the presence of multiple cores adds challenges to the temperature estimation activity, due to the induced heat-up process from adjacent cores. Run-time estimation can benefit from floorplan information to better estimate the thermal characteristics of each core, and transient information can help the system to predict and avoid thermal alarms, increasing the system reliability and lifetime. In this context, the present work aims at defining a novel on-line measurement methodology based on neighbor nodes and transient information, providing a metric to be employed in thermal-aware designs, either at design-time to characterize application and platform from a thermal view-point, or at run-time in conjunction with the Dynamic Thermal Management subsystem. The proposed methodology intercepts floorplan-induced thermal behavior that would be otherwise unrecognized, and it also shows how a non-floorplan-aware methodology can reveal up to a 30% error in the estimate of thermal status. Simone Corbetta, William Fornaciari |
ISCAS | 2 |
| 2011 | Software energy estimation based on statistical characterization of intermediate compilation code
Carlo Brandolese, Simone Corbetta, William Fornaciari |
ISLPED | 3 |
| 2010 | Constrained Power Management: Application to a multimedia mobile platformabstractIn this paper we provide an overview of CPM, a cross-layer framework for Constrained Power Management, and we present its application on a real use case. This framework involves different layers of a typical embedded system, ranging from device drivers to applications. The main goals of CPM are (i) to aggregate applications' QoS requirements and (ii) to exploit them to support an efficient coordination between different drivers' local optimization policies. This role is supported by a system-wide and multi-objective optimization policy which could be also changed at run-time. In this paper we mostly focus on a real use case to show the very low overhead of CPM both on the management of QoS requirements and on the tracking of hardware crossdependencies, which cannot be directly considered by local optimization policies. Patrick Bellasi, Stefano Bosisio, Matteo Carnevali, William Fornaciari, David Siorpaes |
DATE | 4 |
| 2009 | Predictive models for multimedia applications power consumption based on use-case and OS level analysisabstractPower management at any abstraction level is a key issue for many mobile multimedia and embedded applications. In this paper a design workflow to generate system-level power models will be presented, tailored to support quantitative run-time power optimization policies to be implemented within an operating system. The approach we followed to derive power models is strongly use-case oriented. Starting from a comprehensive general and accurate model of a representative architecture for embedded applications (including a multi core MPSoC, accelerators, interfaces and peripherals), a methodology to derive compact models is presented, based upon the distinctive characteristics of the selected use cases. The methodology to generate such model, whose exploitation is foreseen within a power manager working at the OS level, is the focus of the paper. The value and accuracy of the approach is quantitatively and statistically justified through extensive experiments carried out on a development board designed for multimedia applications. Patrick Bellasi, William Fornaciari, David Siorpaes |
DATE | 2 |
| 2009 | A Framework for Compile-time and Run-time Management of Non-functional Aspects in WSNs NodesabstractThe quality of realistic complex wireless sensor networks requires several non-functional aspects to be accounted for starting from the early phases of the application development. The most relevant aspects that need to be considered for optimization trade-offs are for sure computational efficiency (in a wide sense) and network lifetime. At lower level these aspects are measured as power/energy consumption and execution time. These are though not the only non-functional aspects to be considered: code size, memory requirement, security, reliability and other properties play often an important role. Accounting for and managing all these aspects explicitly and in an ad-hoc manner for each and every application deployed on a WSN is a time consuming and complex task. This paper proposes a portable, flexible and extendable framework for the description and management of non-functional aspects in the wireless sensor network context. Carlo Brandolese, William Fornaciari |
DSD | 2 |
| 2008 | Measurement, Analysis and Modeling of RTOS System Calls TimingabstractThis paper presents a methodology for accurately characterizing the system calls of an operating system for embedded applications. Characterization consists of two phases: measurements and modeling. Measurements allow a coarse-grained quantitative comparison of different operating systems. Models, on the other hand, have been derived to gain a more detailed view of the behavior of a RTOS. Furthermore, they have been used within a source-level execution time estimation framework and their accuracy and usefulness proved through benchmarking. The measurement framework is based on two prototyping boards based on Xilinx and Altera devices and the timing characterization has been performed on two real-time operating systems: VxWorks and RTEMS. Carlo Brandolese, William Fornaciari |
DSD | 2 |
| 2008 | Models and Tradeoffs in WSN System-Level DesignabstractSystem-level design of WSNs includes the selection of the sensing nodes and their dissemination in the environment to be monitored. Many design choices have to be taken during this stage of the development of the application. The goal of this paper is to present a methodology to specify formally the desired behavior of the sensing application and to derive an optimal selection and placement of the network nodes. The approach is flexible and powerful, since it allows the designer to analyze the impact of clustering sensors onto a reduced set of boards (nodes) and to perform sensitivity analysis on parameters such as type of sensor, position of the nodes, observation time, presence of faults, etc. The paper introduces the concepts with some representative examples as well as by considering bigger use cases extracted from international projects. Simone Campanoni, William Fornaciari |
DSD | 2 |
| 2006 | Affinity-Driven System Design Exploration for Heterogeneous Multiprocessor SoCabstractContinuous advances in silicon technology enable the development of complex system-on-chip as cooperation among digital signal processors (DPSs), general purpose processors (GPPs), and specific hardware components. The impact of this choice is not only limited to the target architecture, but also encompasses the overall system specification. It is thus crucial to manage such a complexity using high-level specification languages and a tool chain supporting the designer throughout a set of strategic decisions, such as the identification of a set of possible target architectures, the verification of the correctness of the specification, and the partitioning of the specification onto a set of computational resources. This paper addresses this type of problem by proposing a design flow supporting the system-level design of heterogeneous multiprocessor system-on-chip (MP-SoC), by extracting information from the system description (e.g., SystemC) - statically and in a fast manner - and by providing a set of quantitative measures correlating the type of executor, the functionality, and a timing estimation. Partitioning and architecture selection are built on top of this data and the final analysis of the selected hardware-software solution over the identified candidates is finally submitted to a timing verification via simulation. Note that the possibility of actually performing a comprehensive design space exploration, in general, is tightly influenced by the interaction between partitioning/architecture-selection and timing simulation in the design flow; for this reason, the description of this aspect is particularly emphasized in the presentation of the methodology. To show the applicability of the proposed methodology, two relevant case studies are described in the paper. Carlo Brandolese, William Fornaciari, Luigi Pomante, Fabio Salice, Donatella Sciuto |
IEEE Trans. Computers | 2 |
| 2005 | Multithreaded Extension to Multicluster VLIW Processors for Embedded ApplicationsabstractInstruction level parallelism (ILP) extraction for multicluster VLIW processors is a very hard task. In this paper, we propose a retargetable architecture that can exploit ILP and thread level parallelism jointly, thus allowing an easier parallelism extraction and improving the performance with respect to traditional multicluster VLIW processors. Domenico Barretta, William Fornaciari, Mariagiovanna Sami, Daniele Bagni |
DATE | 2 |
| 2004 | An area estimation methodology for FPGA based designs at systemc-levelabstractThis paper presents a parametric area estimation methodology at SystemC level for FPGA-based designs. The approach is conceived to reduce the effort to adapt the area estimators to the evolutions of the EDA design environments. It consists in identifying the subset of measures that can be derived form the system level description and that are also relevant at VHDL-RT level. Estimators' parameters are then automatically derived from a set of benchmarks. Carlo Brandolese, William Fornaciari, Fabio Salice |
DAC | 2 |
| 2004 | Analysis and Modeling of Energy Reducing Source Code TransformationsabstractThis paper presents a methodology and a set of models supporting energy-driven source-to-source transformations. The most promising code transformation techniques have been isolated and studied leading to accurate analytical and/or statistical models. Experimental results, obtained for some common embedded-system processors over a set of typical benchmarks, are presented, showing the viability of the proposed approach as a support tool for embedded software design. Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto |
DATE | 2 |
| 2003 | Library Functions Timing Characterization for Source-Level Analysis
Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto |
DATE | 2 |
| 2003 | A First Step Towards Hw/Sw Partitioning of UML Specifications
William Fornaciari, P. Micheli, Fabio Salice, L. Zampella |
DATE | 1 |
| 2003 | An Internal Representation Model for System-Level Co-Design of Heterogeneous Multiprocessor Embedded System
Fabio Salice, William Fornaciari, Luigi Pomante, Donatella Sciuto |
FDL | 2 |
| 2002 | SIMD Extension to VLIW Multicluster Processors for Embedded ApplicationsabstractWe propose a retargetable architecture, based on a multicluster VLIW processor that can exploit either instruction level parallelism (ILP) or ILP and data level parallelism (DLP) jointly in a SIMD fashion. Simulation results show that performances may increase significantly when the application is compiled for the proposed architecture. Domenico Barretta, William Fornaciari, Mariagiovanna Sami, Danilo Pau |
ICCD | 2 |
| 2002 | Static power modeling of 32-bit microprocessorsabstractThe paper presents a novel strategy aimed at modeling instruction energy consumption of 32-bit microprocessors. Different from former approaches, the proposed instruction-level power model is founded on a functional decomposition of the activities accomplished by a generic microprocessor. The proposed model has significant generalization capabilities. It allows estimation of the power figures of the entire instruction-set starting from the analysis of a subset, as well as to power characterize new processors by using the model obtained by considering other microprocessors. The model is formally presented and justified and its actual application over five commercial microprocessors is included. This static characterization is the basic information for system-level power modeling of hardware/software architectures. Carlo Brandolese, Fabio Salice, William Fornaciari, Donatella Sciuto |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | An Assembly-Level Execution-Time Model for Pipelined ArchitecturesabstractThe aim of this work is to provide an elegant and accurate static execution timing model for 32-bit microprocessor instruction sets, covering also inter-instruction effects. Such effects depend on the processor state and the pipeline behavior, and are related to the dynamic execution of assembly code. The paper proposes a mathematical model of the delays deriving from instruction dependencies and gives a statistical characterization of such timing overheads. The model has been validated on a commercial architecture, the Intel486, by means of timing analysis of a set of benchmarks, obtaining an error within 5%. This model can be seamlessly integrated with a static energy consumption model in order to obtain precise software power and energy estimations. Giovanni Beltrame, Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto, Vito Trianni |
ICCAD | 3 |
| 2000 | An instruction-level functionally-based energy estimation model for 32-bits microprocessorsabstractThe paper presents a novel strategy aimed at modeling the instruction energy consumption of 32-bits microprocessors. The proposed instruction-level pow er model is founded on afunctional decomposition of the activities accomplished by a generic microprocessor and exhibits significant generalization capabilities. It allo ws estimation of the pow er figures of the en tire instruction-set starting from the analysis of a subset, as w ell as to po w er characterize new processors using the model obtained by considering other microprocessors. Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto |
DAC | 2 |
| 2000 | Determining the optimum extended instruction-set architecture for application specific reconfigurable VLIW CPUs (poster abstract)abstractNo abstract available. Cesare Alippi, William Fornaciari, Laura Pozzi 0001, Mariagiovanna Sami |
FPGA | 2 |
| 2000 | Virtualization of FPGA via segmentation (poster abstract)abstractNo abstract available. William Fornaciari, Vincenzo Piuri, Luigi Ripamonti |
FPGA | 1 |
| 1999 | A DAG-Based Design Approach for Reconfigurable VLIW Processors
Cesare Alippi, William Fornaciari, Laura Pozzi 0001, Mariagiovanna Sami |
DATE | 2 |
| 1999 | Influence of Caching and Encoding on Power Dissipation of System-Level Buses for Embedded SystemsabstractThis paper proposes a methodology to evaluate the effects of encodings on the power consumption of system-level buses in the presence of multi-level cache memories. The proposed model can consider any cache configuration in terms of size, associativity and block. It includes also the most widely adopted power oriented encoding techniques for data and address buses. Experimental results show how the proposed model can be effectively adopted to configure the memory hierarchy and the system bus architecture from the power point of view. William Fornaciari, Donatella Sciuto, Cristina Silvano |
DATE | 1 |
| 1999 | Power Estimation of System-Level Buses for Microprocessor-Based Architectures: A Case StudyabstractThe processor-to-memory communication on system-level buses dissipates a significant amount of the overall power in microprocessor-based architectures. A methodology has been set up to evaluate the effects of both encoding schemes and multi-level cache memories on the power consumption associated with the system-level address and data buses of a high-end computer system based on the PowerPC604e architecture. The main goal is to evaluate how different values of cache parameters (cache size, block size, associativity write strategy, and block replacement policy) and the introduction of bus encoding techniques, at the different levels of the memory hierarchy, affect the system-level power dissipation. William Fornaciari, Donatella Sciuto, Cristina Silvano |
ICCD | 1 |
| 1998 | A Model for System-Level Timed Analysis and ProfilingabstractFast evaluation of functional and timing properties is becoming a key factor to enable cost-effective exploration of mixed hw/sw design alternatives for embedded applications. The goal of this paper is to present a modeling strategy to specify functionality and timing properties of uncommitted mixed hw/sw systems. In addition, the paper proposes a simulation algorithm able to perform fast high-level simulation of the system by taking into account the initial hw vs. sw allocation of system modules. The related CAD simulation environment allows the designer to access profiling information which can be useful to remodel the system to meet the functional/timing goals as well as to drive the following hw vs sw partioning activity. Experimental data obtained by reengineering an industrial design are also included in the paper. Alberto Allara, William Fornaciari, Fabio Salice, Donatella Sciuto |
DATE | 2 |
| 1998 | System-level performance estimation strategy for sw and hwabstractThe design of an embedded system is a process where the tuning of the architecture should take into account both the functionality and the timing performance while considering the heterogeneity of the hw and sw components. The goal of this paper is to present the new model developed during the SEED Esprit project, to estimate the software and hardware characteristics for cosimulation and profiling within the TOSCA codesign framework. The impact on the design space exploration of such an high-level cosimulation strategy has been tested by considering as a benchmark the reengineering of an industrial device. Alberto Allara, Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto |
ICCD | 3 |
| 1998 | Power estimation of embedded systems: a hardware/software codesign approachabstractThe need for low-power embedded systems has become very significant within the microelectronics scenario in the most recent years. A power-driven methodology is mandatory during embedded systems design to meet system-level requirements while fulfilling time-to-market. The aim of this paper is to introduce accurate and efficient power metrics included in a hardware/software (HW/SW) codesign environment to guide the system-level partitioning. Power evaluation metrics have been defined to widely explore the architectural design space at high abstraction level. This is one of the first approaches that considers globally HW and SW contributions to power in a system-level design flow for control dominated embedded systems. William Fornaciari, Paolo Gubian, Donatella Sciuto, Cristina Silvano |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1997 | Improving Design Turnaround Time via Two-Levels Hw/Sw Co-SimulationabstractThe steadily growing demand of fast turnaround time will shift system tuning from physical prototyping to virtual prototyping. The paper proposes a novel approach for mixed HW-SW implementation of embedded systems, allowing high-level simulation of the overall architecture as well as a deeper analysis of timing performance by exploiting commercial VHDL CAD tools. At the higher level, functional debugging and tradeoff analysis is performed on an OCCAM-based system-level model and at the lower level a VHDL-based description for both the HW and SW is built for fine grain verification of the system. The paper introduces the two levels of simulation, showing their impact in terms of design flow management and design time. Alberto Allara, S. Filipponi, William Fornaciari, Fabio Salice, Donatella Sciuto |
ICCD | 3 |
| 1997 | A VHDL-based approach for power estimation of embedded systems
William Fornaciari, Paolo Gubian, Donatella Sciuto, Cristina Silvano |
J. Syst. Archit. | 1 |
| 1995 | A new architecture for the automatic design of custom digital neural networkabstractThis brief presents a novel high-performance architecture for implementation of custom digital feed forward neural networks, without on-line learning capabilities. The proposed methodology covers the entire design flow of a neural application, by addressing the internal neuron's structure, the system level organization of the processing elements, the mapping of the abstract neural topology (obtained through simulation) onto the given digital system and eventually the actual synthesis. Experimental results as well as a brief description of the software environment supporting the proposed methodology are also included. William Fornaciari, Fabio Salice |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1994 | HW/SW Codesign for Embedded Telecom SystemsabstractThe aim of this paper is to define an approach, tailored for control-oriented applications, to manage system cospecification, high-level partitioning, hw/sw tradeoffs and cosynthesis. Our research effort focuses on fulfilling the goal of linking high-level specifications to efficient and cost-effective hw/sw implementations by investigating techniques such as synchronous cospecification styles, direct machine code generation as well as exploiting the capability of commercial VHDL synthesis tools.> Stefano Antoniazzi, Alessandro Balboni, William Fornaciari, Donatella Sciuto |
ICCD | 3 |
| 1993 | X-Nets: A visual formalism for system specification and analysis
Stefano Antoniazzi, Alessandro Balboni, William Fornaciari |
Microprocess. Microprogramming | 3 |
| 1990 | APES - Implementation of a CAD tool for array processor design: Textual definition versus graphic description
Fausto Distante, Vincenzo Piuri, Angelo Aliquo', Nicola Chiari, William Fornaciari, Paolo Rastelli |
Microprocessing and Microprogramming | 5 |